A biomarker combination and its application in predicting the risk of bladder cancer
Through biomarker combination and big data analysis, a high sensitivity and high specificity bladder cancer risk prediction model was constructed, solving the problem of accurately predicting bladder cancer in the early stage and providing technical support for early diagnosis and treatment.
Patent Information
- Application Number
- CN202411548467.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-01
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-11-01
AI Technical Summary
The prior art lacks biomarkers that can accurately predict the risk of bladder cancer in the early stage, and traditional screening methods have problems with high false positive rates and late discovery.
A biomarker combination consists of 76 biomarkers. By detecting the expression level of proteomics, a prediction model is constructed using big data analysis methods, including AGPAT5, ARHGEF10L, ATL3, etc., and a generalized linear regression model of Firmiana software is used to predict.
It has achieved high sensitivity and high specificity of bladder cancer risk prediction, supports early clinical diagnosis and intervention treatment, and has extensive scientific research value.
Smart Images

Figure CN119410773B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of biological medicine technology and diagnosis, and particularly relates to a biomarker combination and its application in predicting the risk of bladder cancer. Background Art
[0002] Bladder cancer refers to a malignant tumor that occurs on the bladder mucosa. It is the most common malignant tumor in the urinary system and one of the top ten common tumors in the body. It ranks first in the incidence of urogenital tumors in China, and in the West, its incidence is second only to prostate cancer, ranking second. In 2012, the incidence of bladder cancer in the national cancer registration areas was 6.61 / 100,000, ranking 9th in the incidence of malignant tumors.
[0003] Therefore, accurate and sensitive early screening methods are very necessary. Traditional cancer tumor screening methods include imaging examinations, pathological examinations, radiological and immunological examinations, etc. Various blood test indicators, X-rays, CTs, etc. in physical examinations are all commonly used methods for screening bladder cancer. However, these methods all have disadvantages such as too high false positive rates and late discovery times.
[0004] Therefore, there is an urgent need for a highly sensitive and accurate diagnostic method to achieve early cancer screening.
[0005] Proteomics has played a major role in revealing the complex molecular events of tumorigenesis, such as tumor occurrence, invasion, metastasis, and treatment resistance. Proteomic tumor diagnosis has the advantages of high sensitivity, strong specificity, and clear background mechanism, and has been increasingly used in tumor detection in recent years. Moreover, the research on these tumor markers is often based on a certain amount of experimental data, and the types of cancers and sample sizes involved are relatively limited. In recent years, with the continuous development of the proteome and the continuous increase of big data on body fluid proteomes. Therefore, by collecting body fluid proteome data and using big data analysis methods to find a tumor risk model with a wide application range and high accuracy is helpful to achieve early diagnosis, which has important clinical significance for early diagnosis and treatment of patients. Summary of the Invention
[0006] The technical problem to be solved by the present invention is that the prior art lacks biomarkers that can accurately predict the risk of bladder cancer at an early stage, and provides a biomarker combination and its application in predicting the risk of bladder cancer. The biomarker combination of the present invention has the advantages of high sensitivity and high specificity in predicting the risk of early bladder cancer, provides favorable technical support for predicting the occurrence and development of bladder cancer, has broad scientific research value, and provides great convenience for early clinical diagnosis, intervention treatment, etc.
[0007] The present invention solves the above technical problems through the following technical solutions.
[0008] The first aspect of the present invention provides an application of a biomarker combination in the preparation of a kit for predicting and / or diagnosing bladder cancer;
[0009] wherein, the biomarker combination is composed of AGPAT5, ARHGEF10L, ATL3, ATP6AP2, B3GAT3, BAG2, BHMT2, BLVRB, C12orf10, C1orf50, CA1, CAST, CAT, CD74, CDC37, CELF1, CLIC1, COPS5, CRP, CTSL, DDT, DENND10, DENND10P1, DNASE2, DUSP3, EIF5A2, EIF5AL1, EML4, FKBP1A, G6PD, GET4, GMPPA, H1-1, H3-4, H3C1, H3C15, HBA1, HBB, HBE1, HCLS1, HNRNPA0, HSPA6, KCNN4, KLHL13, KLHL9, LGALS4, LSS, MAN2B1, MECP2, MRPL17, MSL1, MTPN, NCALD, PCBP3, PDLIM3, PLAA, PLEKHM1, POSTN, PRDX2, PTPRE, PUS1, PXDN, RBM47, RDH10, S100A4, SCGB1A1, SFTPB, SMARCA5, SQSTM1, STK24, TBL3, TEX15, TMPRSS13, TXN, UFC1 and XPNPEP1.
[0010] The second aspect of the present invention provides a reagent for detecting a biomarker combination, which is composed of AGPAT5, ARHGEF10L, ATL3, ATP6AP2, B3GAT3, BAG2, BHMT2, BLVRB, C12orf10, C1orf50, CA1, CAST, CAT, CD74, CDC37, CELF1, CLIC1, COPS5, CRP, CTSL, DDT, DENND10, DENND10P1, DNASE2, DUSP3, EIF5A2, EIF5AL1, EML4, FKBP1A, G6PD, GET4, GMPPA, H1-1, H3-4, H3C1, H3C15, HBA1, HBB, HBE1, HCLS1, HNRNPA0, HSPA6, KCNN4, KLHL13, KLHL9, LGALS4, LSS, MAN2B1, MECP2, MRPL17, MSL1, MTPN, NCALD, PCBP3, PDLIM3, PLAA, PLEKHM1, POSTN, PRDX2, PTPRE, PUS1, PXDN, RBM47, RDH10, S100A4, SCGB1A1, SFTPB, SMARCA5, SQSTM1, STK24, TBL3, TEX15, TMPRSS13, TXN, UFC1 and XPNPEP1.
[0011] In some embodiments of the present invention, the reagent is used to detect the expression level of the biomarker combination; the expression level is the protein expression level and / or the mRNA transcription level.
[0012] In some preferred embodiments of the present invention, the reagent is a biomolecular reagent that specifically binds to the biomarker or specifically hybridizes with the nucleic acid encoding the biomarker.
[0013] In some embodiments of the present invention, the biomolecular reagent is selected from primers, probes and antibodies.
[0014] In some embodiments of the present invention, the reagent is a reagent for genome, transcriptome and / or proteome sequencing.
[0015] The third aspect of the present invention provides the use of a reagent for detecting a biomarker combination in the preparation of a kit for predicting and / or diagnosing bladder cancer;
[0016] Among them, the biomarker combination consists of AGPAT5, ARHGEF10L, ATL3, ATP6AP2, B3GAT3, BAG2, BHMT2, BLVRB, C12orf10, C1orf50, CA1, CAST, CAT, CD74, CDC37, CELF1, CLIC1, COPS5, CRP, CTSL, DDT, DENND10, DENND10P1, DNASE2, DUSP3, EIF5A2, EIF5AL1, EML4, FKBP1A, G6PD, GET4, GMPPA, H1-1, H3-4, H3C1, H3C15, HBA1, HBB, HBE1, HCLS1, HNRNPA0, HSPA6, KCNN4, KLHL13, KLHL9, LGALS4, LSS, MAN2B1, MECP2, MRPL17, MSL1, MTPN, NCALD, PCBP3, PDLIM3, PLAA, PLEKHM1, POSTN, PRDX2, PTPRE, PUS1, PXDN, RBM47, RDH10, S100A4, SCGB1A1, SFTPB, SMARCA5, SQSTM1, STK24, TBL3, TEX15, TMPRSS13, TXN, UFC1 and XPNPEP1.
[0017] In some embodiments of the present invention, the reagent is as described in the second aspect.
[0018] The fourth aspect of the present invention provides a biomarker combination, which is composed of AGPAT5, ARHGEF10L, ATL3, ATP6AP2, B3GAT3, BAG2, BHMT2, BLVRB, C12orf10, C1orf50, CA1, CAST, CAT, CD74, CDC37, CELF1, CLIC1, COPS5, CRP, CTSL, DDT, DENND10, DENND10P1, DNASE2, DUSP3, EIF5A2, EIF5AL1, EML4, FKBP1A, G6PD, GET4, GMPPA, H1-1, H3-4, H3C1, H3C15, HBA1, HBB, HBE1, HCLS1, HNRNPA0, HSPA6, KCNN4, KLHL13, KLHL9, LGALS4, LSS, MAN2B1, MECP2, MRPL17, MSL1, MTPN, NCALD, PCBP3, PDLIM3, PLAA, PLEKHM1, POSTN, PRDX2, PTPRE, PUS1, PXDN, RBM47, RDH10, S100A4, SCGB1A1, SFTPB, SMARCA5, SQSTM1, STK24, TBL3, TEX15, TMPRSS13, TXN, UFC1 and XPNPEP1.
[0019] The fifth aspect of the present invention provides a kit, which contains the reagent as described in the second aspect and the biomarker combination as described in the fourth aspect.
[0020] The sixth aspect of the present invention provides a method for detecting bladder cancer for non-diagnostic purposes, which includes detecting the expression level of the biomarker combination in a sample to be tested;
[0021] Among them, the biomarker combination consists of AGPAT5, ARHGEF10L, ATL3, ATP6AP2, B3GAT3, BAG2, BHMT2, BLVRB, C12orf10, C1orf50, CA1, CAST, CAT, CD74, CDC37, CELF1, CLIC1, COPS5, CRP, CTSL, DDT, DENND10, DENND10P1, DNASE2, DUSP3, EIF5A2, EIF5AL1, EML4, FKBP1A, G6PD, GET4, GMPPA, H1-1, H3-4, H3C1, H3C15, HBA1, HBB, HBE1, HCLS1, HNRNPA0, HSPA6, KCNN4, KLHL13, KLHL9, LGALS4, LSS, MAN2B1, MECP2, MRPL17, MSL1, MTPN, NCALD, PCBP3, PDLIM3, PLAA, PLEKHM1, POSTN, PRDX2, PTPRE, PUS1, PXDN, RBM47, RDH10, S100A4, SCGB1A1, SFTPB, SMARCA5, SQSTM1, STK24, TBL3, TEX15, TMPRSS13, TXN, UFC1, and XPNPEP1;
[0022] The expression level is the protein expression level and / or the mRNA transcription level.
[0023] In the present invention, the "non-diagnostic purpose" means for the purposes of scientific research and pathological data statistics, and the applicable scenarios include verifying whether an animal model is successfully constructed, in vitro drug efficacy experiments, epidemiological statistics of tumors, etc.
[0024] The seventh aspect of the present invention provides a prediction system for bladder cancer risk. The prediction system includes a detection module and an analysis and judgment module; the detection module detects the expression level of the biomarker combination in the sample to be tested and transmits the expression level data to the analysis and judgment module; the analysis and judgment module processes the expression level data through Firmiana software. The expression level data is preferably FOT (Fraction of total, defined as the iBAQ of this protein divided by the total iBAQ of all identified proteins in the sample). A machine learning algorithm based on a generalized linear regression model is preset to construct a prediction model to predict the probability of the sample having bladder cancer and the probability of not having bladder cancer respectively, and to judge whether the expression level data meets the preset judgment conditions to predict the risk of the sample having bladder cancer and output a prediction result; the judgment condition is that the probability of having bladder cancer is greater than or equal to the probability of not having bladder cancer;
[0025] When the expression level data meet the judgment conditions, the predicted result is output as "at risk of bladder cancer"; when the expression level data do not meet the judgment conditions, that is, the probability of having bladder cancer is less than the probability of not having bladder cancer, the predicted result is output as "not at risk of bladder cancer".
[0026] Among them, the biomarker combination consists of AGPAT5, ARHGEF10L, ATL3, ATP6AP2, B3GAT3, BAG2, BHMT2, BLVRB, C12orf10, C1orf50, CA1, CAST, CAT, CD74, CDC37, CELF1, CLIC1, COPS5, CRP, CTSL, DDT, DENND10, DENND10P1, DNASE2, DUSP3, EIF5A2, EIF5AL1, EML4, FKBP1A, G6PD, GET4, GMPPA, H1-1, H3-4, H3C1, H3C15, HBA1, HBB, HBE1, HCLS1, HNRNPA0, HSPA6, KCNN4, KLHL13, KLHL9, LGALS4, LSS, MAN2B1, MECP2, MRPL17, MSL1, MTPN, NCALD, PCBP3, PDLIM3, PLAA, PLEKHM1, POSTN, PRDX2, PTPRE, PUS1, PXDN, RBM47, RDH10, S100A4, SCGB1A1, SFTPB, SMARCA5, SQSTM1, STK24, TBL3, TEX15, TMPRSS13, TXN, UFC1 and XPNPEP1.
[0027] The expression level is the protein expression level and / or the mRNA transcription level.
[0028] In some embodiments of the present invention, the prediction system is used to process the expression level data through Firmiana software after the reception or input is completed. The machine learning algorithm is preset as a generalized linear regression model to construct the prediction system.
[0029] In some embodiments of the present invention, the sample to be tested is a plasma sample.
[0030] In some embodiments of the present invention, the prediction system further includes a data collection module. The data collection module is used to collect the expression level data of the biomarker combination in the sample to be tested. The expression level data is preferably FOT (Fraction of total, defined as the iBAQ of this protein divided by the total iBAQ of all identified proteins in the sample).
[0031] In some embodiments of the present invention, the prediction system is a system for predicting early-stage bladder cancer.
[0032] The eighth aspect of the present invention provides a computer-readable storage medium storing a computer program, which when executed by a processor, can implement the functions of the prediction system as described in the seventh aspect of the present invention, or implement the steps of the method as described in the sixth aspect of the present invention.
[0033] The ninth aspect of the present invention provides an electronic device, which includes a memory and a processor, the memory stores a computer program, and the processor is used to execute the computer program to implement the functions of the prediction system as described in the seventh aspect of the present invention, or implement the steps of the method as described in the sixth aspect of the present invention.
[0034] The present invention establishes a tumor risk model through humoral protein molecules for bladder cancer screening, which helps to achieve early diagnosis of bladder cancer.
[0035] The present invention obtains a group of biomarkers capable of predicting the risk of bladder cancer by screening the humoral proteome, and the screening method includes the following steps:
[0036] (1) Collect humoral samples of healthy people and bladder cancer patients;
[0037] (2) Prepare proteins from the humoral samples of healthy people and bladder cancer patients;
[0038] (3) Detect the expression levels of protein molecules in the humoral samples of healthy people and bladder cancer patients;
[0039] (4) Find the protein group molecules with high specific expression in the body fluids of tumor patients and construct a classifier for discrimination.
[0040] On the basis of conforming to the common knowledge in the art, the above preferred conditions can be combined arbitrarily to obtain various preferred examples of the present invention.
[0041] The reagents and raw materials used in the present invention are all commercially available.
[0042] The positive and progressive effects of the present invention are as follows:
[0043] It is experimentally found that the expression levels of the above-mentioned protein molecule markers provided by the present invention in the humoral samples of tumor patients and healthy people have significant changes. Therefore, the humoral protein molecule markers provided in the present invention can be used for risk prediction and detection of tumor patients, and have the advantages of high sensitivity and high specificity, providing favorable technical support for predicting the occurrence and development of bladder cancer.
[0044] Developing a corresponding prediction device based on protein molecular markers in body fluid samples of healthy people and bladder cancer patients has extensive scientific research value and provides great convenience for early clinical diagnosis, intervention treatment, etc. Description of the Drawings
[0045] Figure 1 It is a schematic diagram of the area under the ROC curve of the biomarker combination in the training set.
[0046] Figure 2 It is a schematic diagram of the area under the ROC curve of the biomarker combination in the test set.
[0047] Figure 3 It is a schematic diagram of the area under the ROC curve of the biomarker combination in the independent validation set.
[0048] Figure 4 It is a schematic structural diagram of a system for predicting the risk of bladder cancer.
[0049] Figure 5 It is a schematic structural diagram of an electronic device. Detailed Implementation Modes
[0050] The present invention will be further described below by way of examples, but the present invention is not limited to the scope of the described examples. For the experimental methods without specific conditions in the following examples, they are carried out according to conventional methods and conditions, or according to the product instructions.
[0051] "Biomarker" refers to a biochemical index that can mark changes or possible changes in the structure or function of a system, organ, tissue, cell, and subcellular structure, and can be used for disease diagnosis, disease staging judgment, or evaluation of the safety and effectiveness of new drugs and new therapies in the target population.
[0052] For those without specific technologies or conditions in the examples, they are carried out according to the technologies or conditions described in the literature in this field, or according to the product instructions. For the reagents or instruments without indicating the manufacturer, they are all conventional products that can be purchased through regular channels.
[0053] The examples include plasma samples of 114 normal people and 121 bladder cancer patients. The design and implementation of this study have passed ethical approval and supervision, and written informed consent has been obtained from all patients.
[0054] Example 1 Screening and Verification of a Combination of Biomarkers for Predicting the Risk of Bladder Cancer
[0055] 1.1 Separation of Plasma
[0056] Collect whole blood samples in EDTA anticoagulant tubes. After inverting and mixing well, use a 4°C low-temperature centrifuge to centrifuge at 1,600×g for 10 min. After centrifugation, collect the supernatant (plasma) into a new EP tube, and centrifuge at 16,000×g for 10 min to remove cell debris. Aliquot the plasma into centrifuge tubes and store at -80°C for later use.
[0057] 1.2 Plasma sample pretreatment
[0058] Add 100 μL of 50 mM ammonium bicarbonate to 2 μL of plasma samples, vortex for 1 min, heat and incubate the samples at 95°C for 4 min to cause thermal denaturation of proteins. After cooling to room temperature, add 2 μg of trypsin to the system, oscillate at 37°C for 18 h, and then add 10 μL of ammonia water to stop the enzymatic digestion. Desalt the peptide samples after enzymatic digestion, dry by evaporation, and store at -80°C until mass spectrometry detection.
[0059] 1.3 Mass spectrometry detection of plasma samples
[0060] Use an Orbitrap Fusion Lumos triple quadrupole high-resolution mass spectrometry system (Thermo Fisher Scientific, Rockford, USA) in tandem with a high-performance liquid chromatography system (EASY-nLC 1200, Thermo Fisher) for detection and obtain the mass spectrometry data of the whole protein corresponding to the peptide samples. The specific operation is as follows:
[0061] Adopt nanoflow liquid chromatography, and the chromatographic column is a self-made C18 chromatographic column (150 μm ID × 8 cm, 1.9 μm / packing). The column oven temperature is 60°C. Reconstitute the dry powder peptides with the loading buffer (aqueous solution of 0.1% formic acid). After loading, separate through the chromatographic column and elute with a linear 6–30% mobile phase B (ACN and 0.1% formic acid) at 600 nL / min, using a 10-min liquid phase gradient combined with data-independent acquisition (DIA) mass spectrometry detection method. The DIA mass spectrometry detection parameters are set as follows: the ion mode is positive ion; the resolution of the first-stage mass spectrometry is 30K, the maximum injection time is 20 ms, the AGC Target is 3e6, and the scanning range is 300 - 1400 m / z; the resolution of the second-stage scanning is 15K, 30 variable isolation windows are obtained, and the collision energy is 27%. The liquid chromatography tandem mass spectrometry system is controlled by Xcalibur software for data acquisition.
[0062] 1.4 Data analysis
[0063] All data were processed using Firmiana (V1.0). Firmiana is a workflow based on the Galaxy system, consisting of multiple functional modules such as a user login interface, raw data, identification and quantification, data analysis, and knowledge mining. DIA data were searched against the UniProt human protein database (updated on December 17, 2019, with 20,406 entries) using DIANN (v12.1). The mass difference of precursor ions was 20 ppm, and the mass difference of product ions was 50 mmu. Up to two missed cleavage sites were allowed. The search engine set carbamidomethylation of cysteine as a fixed modification and N-acetylation and oxidation of methionine as variable modifications. The precursor ion charge range was set to +2, +3, and +4. The False Discovery Rate (FDR) was set to 1%. The results of DIA data were merged into the reference library using SpectraST software. A total of 327 libraries were used as the reference library.
[0064] The quantitative results of the identified peptides were recorded as the average of the chromatographic fragment ion peak areas in all reference spectral libraries. Label-free intensity-based absolute quantification (iBAQ) method was used for protein quantification. We calculated the peak area value as a part of the corresponding protein. The Fraction of Total (FOT) was used to represent the normalized abundance of a specific protein in the sample. FOT was defined as the iBAQ of the protein divided by the total iBAQ of all identified proteins in the sample. Proteins with at least one unique peptide and 1% FDR were selected. The FOT of each protein was calculated, and the FOT of each protein was input into the generalized linear regression model as protein expression data.
[0065] In this example, the selected Firmiana was preset as a machine learning algorithm based on the generalized linear regression model to construct a prediction model, and the probabilities of the sample having bladder cancer and not having bladder cancer were predicted respectively. The code for constructing the prediction model is:
[0066] from sklearn.linear_model import LogisticRegressionCV
[0067] import joblib
[0068] import pandas as pd
[0069] tumor_types = ['***', '***', '***', '***', '***'] # *** refers to the 76 protein molecular markers described in the present invention
[0070] for tumor_type in tumor_types:
[0071] df_train = pd.read_csv('{}_train.csv'.format(tumor_type))
[0072] df_test = pd.read_csv('{}_test.csv'.format(tumor_type))
[0073] df_val = pd.read_csv('{}_validation.csv'.format(tumor_type))
[0074] X_train = df_train.iloc[:, 2:]
[0075] y_train = df_train.iloc[:, 1]
[0076] X_test = df_test.iloc[:, 2:]
[0077] y_test = df_test.iloc[:, 1]
[0078] X_val = df_val.iloc[:, 2:]
[0079] y_val = df_val.iloc[:, 1]
[0080] model = LogisticRegressionCV()
[0081] model.fit(X_train, y_train)
[0082] joblib.dump(model, f'{tumor_type}_model.joblib');
[0083] Experimental findings show that there are significant changes in the expression levels of some proteins in the body fluid samples of cancer patients and healthy individuals. The ROC curve (Receiver Operating Curve) was plotted for the relative expression levels of 76 protein molecular markers (AGPAT5, ARHGEF10L, ATL3, ATP6AP2, B3GAT3, BAG2, BHMT2, BLVRB, C12orf10, C1orf50, CA1, CAST, CAT, CD74, CDC37, CELF1, CLIC1, COPS5, CRP, CTSL, DDT, DENND10, DENND10P1, DNASE2, DUSP3, EIF5A2, EIF5AL1, EML4, FKBP1A, G6PD, GET4, GMPPA, H1-1, H3-4, H3C1, H3C15, HBA1, HBB, HBE1, HCLS1, HNRNPA0, HSPA6, KCNN4, KLHL13, KLHL9, LGALS4, LSS, MAN2B1, MECP2, MRPL17, MSL1, MTPN, NCALD, PCBP3, PDLIM3, PLAA, PLEKHM1, POSTN, PRDX2, PTPRE, PUS1, PXDN, RBM47, RDH10, S100A4, SCGB1A1, SFTPB, SMARCA5, SQSTM1, STK24, TBL3, TEX15, TMPRSS13, TXN, UFC1, and XPNPEP1) in the plasma samples of bladder cancer patients to calculate the AUC (Area Under the ROC Curve). The training set included 58 positive cases and 56 negative cases, with AUC = 1.00, diagnostic sensitivity of 100.00%, and specificity of 100.00% (see Figure 1 ). The test set included 26 positive cases and 24 negative cases, with AUC = 0.91, diagnostic sensitivity of 84.62%, and specificity of 83.33% (see Figure 2 ). The independent validation set included 37 positive cases and 34 negative cases, with AUC = 0.96, diagnostic sensitivity of 91.89%, and specificity of 91.18% (see Figure 3)。For the analysis method, see Karimollah Hajian-Tilaki, Receiver Operating Characteristic (ROC) Curve Analysis for Medical Diagnostic Test Evaluation, Caspian J Intern Med 2013; 4(2): 627-635. For unknown samples, substitute the expression levels of the above biomarkers into the model to obtain the bladder cancer risk prediction for the sample and output the result. When the probability of having bladder cancer is greater than or equal to the probability of not having bladder cancer, the prediction result is output as "at risk of bladder cancer"; when the probability of having bladder cancer is less than the probability of not having bladder cancer, the prediction result is output as "not at risk of bladder cancer". The FOT values of 76 protein biomarkers in the training set, test set, and independent validation set are shown in Tables 1-6, 7-12, and 13-18.
[0084] Table 1 FOT values of 76 protein biomarkers in the training set
[0085]
[0086]
[0087]
[0088] Table 2 FOT values of 76 protein biomarkers in the training set
[0089]
[0090]
[0091]
[0092]
[0093] Table 3 FOT values of 76 protein biomarkers in the training set
[0094]
[0095]
[0096]
[0097] Table 4 FOT values of 76 protein biomarkers in the training set
[0098]
[0099]
[0100]
[0101]
[0102] FOT values of 76 protein markers in the training set
[0103]
[0104]
[0105]
[0106] FOT values of 76 protein markers in the training set
[0107]
[0108]
[0109]
[0110]
[0111] FOT values of 76 protein markers in the test set
[0112]
[0113]
[0114] FOT values of 76 protein markers in the test set
[0115]
[0116]
[0117] FOT values of 76 protein markers in the test set
[0118]
[0119]
[0120] FOT values of 76 protein markers in the test set
[0121]
[0122]
[0123] FOT values of 76 protein markers in the test set
[0124]
[0125]
[0126] FOT values of 76 protein markers in the test set, Table 12
[0127]
[0128]
[0129] FOT values of 76 protein markers in the independent validation set, Table 13
[0130]
[0131]
[0132] FOT values of 76 protein markers in the independent validation set, Table 14
[0133]
[0134]
[0135]
[0136] FOT values of 76 protein markers in the independent validation set, Table 15
[0137]
[0138]
[0139] FOT values of 76 protein markers in the independent validation set, Table 16
[0140]
[0141]
[0142]
[0143] FOT values of 76 protein markers in the independent validation set, Table 17
[0144]
[0145]
[0146] FOT values of 76 protein markers in the independent validation set, Table 18
[0147]
[0148]
[0149]
[0150] As can be seen from the above results, the combined use of 76 protein molecular markers (AGPAT5, ARHGEF10L, ATL3, ATP6AP2, B3GAT3, BAG2, BHMT2, BLVRB, C12orf10, C1orf50, CA1, CAST, CAT, CD74, CDC37, CELF1, CLIC1, COPS5, CRP, CTSL, DDT, DENND10, DENND10P1, DNASE2, DUSP3, EIF5A2, EIF5AL1, EML4, FKBP1A, G6PD, GET4, GMPPA, H1-1, H3-4, H3C1, H3C15, HBA1, HBB, HBE1, HCLS1, HNRNPA0, HSPA6, KCNN4, KLHL13, KLHL9, LGALS4, LSS, MAN2B1, MECP2, MRPL17, MSL1, MTPN, NCALD, PCBP3, PDLIM3, PLAA, PLEKHM1, POSTN, PRDX2, PTPRE, PUS1, PXDN, RBM47, RDH10, S100A4, SCGB1A1, SFTPB, SMARCA5, SQSTM1, STK24, TBL3, TEX15, TMPRSS13, TXN, UFC1, and XPNPEP1) in the plasma of tumor patients can be used to predict tumor risk.
[0151] Example 2 System for Predicting the Risk of Bladder Cancer
[0152] The system 61 for predicting the risk of bladder cancer: a data processing module 52 and a judgment and output module 53, further including a data collection module 51 ( Figure 4 ).
[0153] The data collection module 51 is used to collect the expression level data of the biomarker combination in the bladder cancer tissue sample of the patient and transmit it to the data processing module.
[0154] The data processing module 52 is used to analyze the expression level data of the biomarker combination received or input according to the data analysis method described in Example 4 to obtain a calculation result. Among them, the expression level data of the biomarker combination can be collected by the data collection module 51, or the expression level data of the biomarker combination can be obtained from other sources.
[0155] The judgment and output module 53 is used to judge whether the calculated result meets a preset judgment condition, that is, the risk probability of suffering from bladder cancer is greater than or equal to the risk prediction probability of not suffering from bladder cancer, so as to predict the risk of bladder cancer and output a prediction result; wherein, in the judgment and output module, when the expression level data meets the judgment condition that the risk probability of suffering from bladder cancer is greater than or equal to the risk prediction probability of not suffering from bladder cancer, the output prediction result is "at risk of bladder cancer"; when the expression level data does not meet the judgment condition that the risk probability of suffering from bladder cancer is less than the risk prediction probability of not suffering from bladder cancer, the output prediction result is "not at risk of bladder cancer".
[0156] Embodiment 3 Electronic device
[0157] This embodiment provides an electronic device, which can be presented in the form of a computing device (for example, it can be a server device), including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for predicting the risk of bladder cancer in Embodiment 1 of the present invention can be implemented.
[0158] Figure 5 The hardware structure diagram of this embodiment is shown. The electronic device 4 specifically includes:
[0159] At least one processor 91, at least one memory 92, and a bus 93 for connecting different system components (including the processor 91 and the memory 92), wherein:
[0160] The bus 93 includes a data bus, an address bus, and a control bus.
[0161] The memory 92 includes volatile memory, such as a random access memory (RAM) 921 and / or a cache memory 922, and may further include a read-only memory (ROM) 923.
[0162] The memory 92 further includes a program / utility 925 having a set (at least one) of program modules 924. Such program modules 924 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0163] The processor 91 executes various functional applications and data processing by running the computer program stored in the memory 92, such as the data analysis method in Embodiment 1 of the present invention.
[0164] The electronic device 9 can further communicate with one or more external devices 94 (such as a keyboard, a pointing device, etc.). Such communication can be carried out through an input / output (I / O) interface 95. Moreover, the electronic device 9 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 96. The network adapter 96 communicates with other modules of the electronic device 9 through a bus 93. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 9, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (redundant array of independent disks) systems, tape drives, and data backup storage systems, etc.
[0165] It should be noted that, although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more units / modules described above can be embodied in one unit / modules. Conversely, the features and functions of one unit / modules described above can be further divided and embodied by multiple units / modules.
[0166] Embodiment 4 Computer-readable storage medium
[0167] The embodiments of the present invention provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method for predicting the risk of bladder cancer in Embodiment 1 of the present invention are implemented.
[0168] Among them, the more specific forms that the readable storage medium can adopt can include but are not limited to: portable disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories, optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0169] In a possible implementation manner, the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps of the method for predicting the risk of bladder cancer in Embodiment 1 of the present invention.
[0170] Among them, the program code for executing the present invention can be written in any combination of one or more programming languages. The program code can be executed completely on the user device, partially on the user device, executed as an independent software package, partially on the user device and partially on a remote device, or executed completely on a remote device.
[0171] Finally, the above specific implementation methods are only used to clarify the technical solutions of the present invention, rather than to limit them.
[0172] Biomarker name: (Reference can be made to the NCBI or genecards database)
[0173] AGPAT5: 1-acylglycerol-3-phosphate O-acyltransferase 5, Gene ID: 55326
[0174] ARHGEF10L: Rho guanine nucleotide exchange factor 10like, Gene ID: 55160
[0175] ATL3: atlastin GTPase 3, Gene ID: 25923
[0176] ATP6AP2: ATPase H+ transporting accessory protein 2, Gene ID: 10159
[0177] B3GAT3: beta-1,3-glucuronyltransferase 3, Gene ID: 26229
[0178] BAG2: BAG cochaperone 2, Gene ID: 9532
[0179] BHMT2: betaine--homocysteine S-methyltransferase 2, Gene ID: 23743
[0180] BLVRB: biliverdin reductase B, Gene ID: 645
[0181] C12orf10: Gene ID: 60314
[0182] C1orf50: chromosome 1 open reading frame 50, Gene ID: 79078
[0183] CA1: carbonic anhydrase 1, Gene ID: 759
[0184] CAST: calpastatin, Gene ID: 831
[0185] CAT: catalase, Gene ID: 847
[0186] CD74: CD74 molecule, Gene ID: 972
[0187] CDC37: cell division cycle 37, HSP90 cochaperone, Gene ID: 11140
[0188] CELF1: CUGBP Elav-like family member 1, Gene ID: 10658
[0189] CLIC1: chloride intracellular channel 1, Gene ID: 1192
[0190] COPS5: COP9 signalosome subunit 5, Gene ID: 10987
[0191] CRP: C-reactive protein, Gene ID: 1401
[0192] CTSL: cathepsin L, Gene ID: 1514
[0193] DDT: D-dopachrome tautomerase, Gene ID: 1652
[0194] DENND10: DENN domain containing 10, Gene ID: 404636
[0195] DENND10P1: DENND10 pseudogene 1, Gene ID: 55855
[0196] DNASE2: deoxyribonuclease 2, lysosomal, Gene ID: 1777
[0197] DUSP3: dual specificity phosphatase 3, Gene ID: 1845
[0198] EIF5A2: eukaryotic translation initiation factor 5A2, Gene ID: 56648 EIF5AL1: eukaryotic translation initiation factor 5A like 1, Gene ID: 143244 EML4: EMAP like 4, Gene ID: 27436
[0199] FKBP1A: FKBP prolyl isomerase 1A, Gene ID: 2280
[0200] G6PD: glucose-6-phosphate dehydrogenase, Gene ID: 2539 GET4: guided entry of tail-anchored proteins factor 4, Gene ID: 51608
[0201] GMPPA: GDP-mannose pyrophosphorylase A, Gene ID: 29926 H1-1: H1.1 linker histone, cluster member, Gene ID: 3024
[0202] H3-4: H3.4 histone, cluster member, Gene ID: 8290
[0203] H3C1: H3 clustered histone 1, Gene ID: 8350
[0204] H3C15: H3 clustered histone 15, Gene ID: 333932
[0205] HBA1: hemoglobin subunit alpha 1, Gene ID: 3039
[0206] HBB: hemoglobin subunit beta, Gene ID: 3043
[0207] HBE1: hemoglobin subunit epsilon 1, Gene ID: 3046
[0208] HCLS1: hematopoietic cell-specific Lyn substrate 1, Gene ID: 3059; HNRNPA0: heterogeneous nuclear ribonucleoprotein A0, Gene ID: 10949
[0209] HSPA6: heat shock protein family A (Hsp70) member 6, Gene ID: 3310
[0210] KCNN4: potassium calcium-activated channel subfamily N member 4, Gene ID: 3783; KLHL13: kelch like family member 13, Gene ID: 90293
[0211] KLHL9: kelch like family member 9, Gene ID: 55958
[0212] LGALS4: galectin 4, Gene ID: 3960
[0213] LSS: lanosterol synthase, Gene ID: 4047
[0214] MAN2B1: mannosidase alpha class 2B member 1, Gene ID: 4125; MECP2: methyl-CpG binding protein 2, Gene ID: 4204
[0215] MRPL17: mitochondrial ribosomal protein L17, Gene ID: 63875; MSL1: MSL complex subunit 1, Gene ID: 339287
[0216] MTPN: myotrophin, Gene ID: 136319
[0217] NCALD: neurocalcin delta, Gene ID: 83988
[0218] PCBP3: poly(rC) binding protein 3, Gene ID: 54039
[0219] PDLIM3: PDZ and LIM domain 3, Gene ID: 27295
[0220] PLAA: phospholipase A2 activating protein, Gene ID: 9373 PLEKHM1: pleckstrin homology and RUN domain containing M1, Gene ID: 9842
[0221] POSTN: periostin, Gene ID: 10631
[0222] PRDX2: peroxiredoxin 2, Gene ID: 7001
[0223] PTPRE: protein tyrosine phosphatase receptor type E, Gene ID: 5791
[0224] PUS1: pseudouridine synthase 1, Gene ID: 80324
[0225] PXDN: peroxidasin, Gene ID: 7837
[0226] RBM47: RNA binding motif protein 47, Gene ID: 54502
[0227] RDH10: retinol dehydrogenase 10, Gene ID: 157506
[0228] S100A4: S100 calcium binding protein A4, Gene ID: 6275
[0229] SCGB1A1: secretoglobin family 1A member 1, Gene ID: 7356
[0230] SFTPB: surfactant protein B, Gene ID: 6439
[0231] SMARCA5: SWI / SNF related, matrix associated, actin dependent regulator of chromatin, subfamily a, member 5, Gene ID: 8467
[0232] SQSTM1: sequestosome 1, Gene ID: 8878
[0233] STK24: serine / threonine kinase 24, Gene ID: 8428
[0234] TBL3: transducin beta like 3, Gene ID: 10607
[0235] TEX15: testis expressed 15, meiosis and synapsis associated, Gene ID: 56154
[0236] TMPRSS13: transmembrane serine protease 13, Gene ID: 84000
[0237] TXN: thioredoxin, Gene ID: 7295
[0238] UFC1: ubiquitin-fold modifier conjugating enzyme 1, Gene ID: 51506
[0239] XPNPEP1: X-prolyl aminopeptidase 1, Gene ID: 7511.
Claims
1. Use of a reagent for detecting a biomarker combination in the preparation of a product for predicting and / or diagnosing bladder cancer, characterized in that, The biomarker combination consists of AGPAT5, ARHGEF10L, ATL3, ATP6AP2, B3GAT3, BAG2, BHMT2, BLVRB, C12orf10, C1orf50, CA1, CAST, CAT, CD74, CDC37, CELF1, CLIC1, COPS5, CRP, CTSL, DDT, DENND10, DENND10P1, DNASE2, DUSP3, EIF5A2, EIF5AL1, EML4, FKBP1A, G6PD, GET4, GMPPA, H1-1, H3-4, H3C1, H3C15, HBA1, HBB, HBE1, HCLS1, HNRNPA0, HSPA6, KCNN4, KLHL13, KLHL9, LGALS4, LSS, MAN2B1, MECP2, MRPL17, MSL1, MTPN, NCALD, PCBP3, PDLIM3, PLAA, PLEKHM1, POSTN, PRDX2, PTPRE, PUS1, PXDN, RBM47, RDH10, S100A4, SCGB1A1, SFTPB, SMARCA5, SQSTM1, STK24, TBL3, TEX15, TMPRSS13, TXN, UFC1 and XPNPEP1.
2. The application according to claim 1, wherein The reagent is used to detect the expression level of the biomarker combination, and the expression level is the protein expression level and / or the mRNA transcription level.
3. The application according to claim 2, characterized in that, The reagent is a biomolecular reagent that specifically binds to the biomarker or specifically hybridizes with the nucleic acid encoding the biomarker; and / or, the reagent is a reagent for genome, transcriptome and / or proteome sequencing.
4. The application according to claim 3, wherein The biomolecular reagent is selected from primers, probes and antibodies.
5. A prediction system for bladder cancer risk, characterized in that, The prediction system includes a detection module and an analysis and judgment module; the detection module detects the expression level of the biomarker combination in a sample to be tested and transmits the expression level data to the analysis and judgment module; the analysis and judgment module processes the expression level data through Firmiana software, which is preset as a machine learning algorithm based on a generalized linear regression model, constructs a prediction model, predicts the probability of the sample having bladder cancer and the probability of not having bladder cancer respectively, determines whether the expression level data meets the preset judgment conditions, so as to predict the risk of the sample having bladder cancer, and outputs a prediction result; The judgment condition is that the probability of having bladder cancer is greater than or equal to the probability of not having bladder cancer; When the expression level data meets the judgment condition, the prediction result output is "at risk of bladder cancer"; when the expression level data does not meet the judgment condition, that is, the probability of having bladder cancer is less than the probability of not having bladder cancer, the prediction result output is "not at risk of bladder cancer"; Among them, the biomarker combination consists of AGPAT5, ARHGEF10L, ATL3, ATP6AP2, B3GAT3, BAG2, BHMT2, BLVRB, C12orf10, C1orf50, CA1, CAST, CAT, CD74, CDC37, CELF1, CLIC1, COPS5, CRP, CTSL, DDT, DENND10, DENND10P1, DNASE2, DUSP3, EIF5A2, EIF5AL1, EML4, FKBP1A, G6PD, GET4, GMPPA, H1-1, H3-4, H3C1, H3C15, HBA1, HBB, HBE1, HCLS1, HNRNPA0, HSPA6, KCNN4, KLHL13, KLHL9, LGALS4, LSS, MAN2B1, MECP2, MRPL17, MSL1, MTPN, NCALD, PCBP3, PDLIM3, PLAA, PLEKHM1, POSTN, PRDX2, PTPRE, PUS1, PXDN, RBM47, RDH10, S100A4, SCGB1A1, SFTPB, SMARCA5, SQSTM1, STK24, TBL3, TEX15, TMPRSS13, TXN, UFC1, and XPNPEP1, and the expression level is the protein expression level and / or the mRNA transcription level.
6. The prediction system according to claim 5, wherein The sample to be tested is a human plasma sample; and / or, the prediction system further includes a data collection module, and the data collection module is used to collect the expression level data of the biomarker combination in the sample to be tested.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it can realize the functions of the prediction system as described in claim 5 or 6.
8. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor is used to execute the computer program to realize the functions of the prediction system as described in claim 5 or 6.
Citation Information
Patent Citations
Non-invasive diagnostic method for diagnosing bladder cancer
CN105229169A
Set of genes used for bladder cancer detection and application thereof
CN108866194A