A biomarker combination and its application in predicting gastric cancer risk
Through the biomarker combination of ABLIM1 and ABRACL expression level detection, combined with a generalized linear regression model, the problem of high false positive rate and late detection time in the prior art diagnosis of gastric cancer is solved, and high sensitivity and high specificity prediction of early gastric cancer risk is achieved.
Patent Information
- Application Number
- CN202411548477.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-01
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2044-11-01
AI Technical Summary
The prior art lacks biomarkers that can accurately predict gastric cancer risk early, resulting in high false positive rates and late detection time for gastric cancer diagnosis.
A biomarker combination is provided, including 145 proteins such as ABLIM1, ABRACL, ACTA1, etc., and by detecting their expression levels, a machine learning algorithm of generalized linear regression model is used to construct a predictive model to achieve high sensitivity and high specific diagnosis of early gastric cancer risk.
It achieves 100% diagnostic sensitivity and 100% specificity, and can accurately predict the risk of gastric cancer in the early stage, providing convenience for early clinical diagnosis and intervention treatment.
Smart Images

Figure CN119464494B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biomedical technology and diagnosis, and specifically relates to a biomarker combination and its application in predicting gastric cancer risk. Background Art
[0002] East Asian countries have a high incidence of gastric cancer and monopolize the top few countries in terms of gastric cancer incidence. According to the latest data, the top three countries in terms of gastric cancer incidence are South Korea, Mongolia and Japan.
[0003] Clinically, early detection of lesions can be achieved through ultrasound, CT scans, and biopsy. Regular self-examination can also help. However, these methods suffer from high false-positive rates and delayed detection. Therefore, a highly sensitive and accurate diagnostic method is urgently needed for early cancer screening.
[0004] Proteomics has played a significant role in revealing the complex molecular events of tumorigenesis, such as tumorigenesis, invasion, metastasis, and resistance to treatment. Proteomic tumor diagnosis, with its advantages of high sensitivity, strong specificity, and clear underlying mechanisms, has been increasingly used for tumor detection in recent years. Furthermore, the research on these tumor markers is often based on a limited amount of experimental data, involving relatively limited cancer types and sample sizes. In recent years, with the continuous development of the proteome, big data on bodily fluid proteomes has continued to increase. Therefore, it is urgent to utilize big data analysis methods to identify a tumor risk model with broad applicability and high accuracy, which can facilitate early diagnosis and provide patients with important clinical significance for early diagnosis and treatment. Summary of the Invention
[0005] The present invention addresses the technical problem of the lack of biomarkers capable of accurately predicting gastric cancer risk at an early stage. The present invention provides a biomarker combination and its application in predicting gastric cancer risk. The biomarker combination of the present invention has the advantages of high sensitivity and specificity in predicting early gastric cancer risk, providing advantageous technical support for predicting the development and progression of gastric cancer. It has broad scientific research value and greatly facilitates early clinical diagnosis, interventional treatment, and other aspects.
[0006] The present invention solves the above technical problems through the following technical solutions.
[0007] A first aspect of the present invention provides a use of a biomarker combination in preparing a kit for predicting and / or diagnosing gastric cancer;
[0008] Among them, the biomarker combination consists of ABLIM1, ABRACL, ACTA1, ACTA2, ACTB, ACTBL2, ACTC1, ACTG1, ACTG2, ACTR3B, ACTR3C, ALDH9A1, ANKHD1, ANKRD17, ARHGDIB, ARHGEF10L, ARMC8, ARPC5, BHMT2, BIN2, CALM1, CALM2, CALM3, CALML3, CAPN1, CAVIN2, CFL1, CHGB, CLIC1, CNN2, CORO1C, CRP, CSRP1, CYFIP1, DLG1, DSC3, DTNB, DUSP3, EIF5A2, EIF5AL1, EIF5B, ELMOD2, EML4, ENO1, ENO2, ENO3, FHL1, FKBP1A, FSCN1, GAPDH, GIT1, GP1BB, GP6, GSTO1, HBE1, HNRNPA0, HSPA6, HSPA7, IGF1, ILK, IMPDH1, ISLR, ITGA2B, ITIH3, ITIH5, LGALSL, LRG1, MAP1A, MAPK8IP3, MCM7, MECP2, MIF, MMP10, MRPL37, MSL1, MTHFD2, MTPN, MYL12B, MYL3, MYL6B, NAXD, NCALD, NIT2, NUP155, ORM1, PARVB, PCBP1, PCBP3, PDCD10, PDLIM1, PEX5, PFN1, PGAM2, PGK2, PKLR, PKM, PMVK, POSTN, POTEE, POTEF, POTEI, POTEJ, POTEKP, PPCS, PPP2R2D, PRDX5, PRKAG1, PSMB8, PTGIS, PTMA, PUDP, RAB1B, RARRES2, RCN1, RDH10, RECK, RPL5, RPS27A, RPS7, S100A9, SAA1, SAA2, SF3A3, SFTPB, SKP1, SMC1A, SMC2, SPG21, SYNE1, SYNE2, SYNM, TAGLN2, TBXAS1, TMSB4X, TNFAIP2, TPM4, TUBB8B, TXN, UBA52, UBB, UBC, VPS13A, VWF, XPNPEP1, and YKT6.
[0009] The second aspect of the present invention provides a reagent for detecting a biomarker combination, which is composed of ABLIM1, ABRACL, ACTA1, ACTA2, ACTB, ACTBL2, ACTC1, ACTG1, ACTG2, ACTR3B, ACTR3C, ALDH9A1, ANKHD1, ANKRD17, ARHGDIB, ARHGEF10L, ARMC8, ARPC5, BHMT2, BIN2, CALM1, CALM2, CALM3, CALML3, CAPN1, CAVIN2, CFL1, CHGB, CLIC1, CNN2, CORO1C, CRP, CSRP1, CYFIP1, DLG1, DSC3, DTNB, DUSP3, EIF5A2, EIF5AL1, EIF5B, ELMOD2, EML4, ENO1, ENO2, ENO3, FHL1, FKBP1A, FSCN1, GAPDH, GIT1, GP1BB, GP6, GSTO1, HBE1, HNRNPA0, HSPA6, HSPA7, IGF1, ILK, IMPDH1, ISLR, ITGA2B, ITIH3, ITIH5, LGALSL, LRG1, MAP1A, MAPK8IP3, MCM7, MECP2, MIF, MMP10, MRPL37, MSL1, MTHFD2, MTPN, MYL12B, MYL3, MYL6B, NAXD, NCALD, NIT2, NUP155, ORM1, PARVB, PCBP1, PCBP3, PDCD10, PDLIM1, PEX5, PFN1, PGAM2, PGK2, PKLR, PKM, PMVK, POSTN, POTEE, POTEF, POTEI, POTEJ, POTEKP, PPCS, PPP2R2D, PRDX5, PRKAG1, PSMB8, PTGIS, PTMA, PUDP, RAB1B, RARRES2, RCN1, RDH10, RECK, RPL5, RPS27A, RPS7, S100A9, SAA1, SAA2, SF3A3, SFTPB, SKP1, SMC1A, SMC2, SPG21, SYNE1, SYNE2, SYNM, TAGLN2, TBXAS1, TMSB4X, TNFAIP2, TPM4, TUBB8B, TXN, UBA52, UBB, UBC, VPS13A, VWF, XPNPEP1 and YKT6.
[0010] In some embodiments of the present invention, the reagent is used to detect the expression level of the biomarker combination; the expression level is the protein expression level and / or the mRNA transcription level.
[0011] In some preferred embodiments of the present invention, the reagent is a biomolecular reagent that specifically binds to the biomarker or specifically hybridizes with the nucleic acid encoding the biomarker.
[0012] In some embodiments of the present invention, the biomolecular reagent is selected from primers, probes, and antibodies.
[0013] In some embodiments of the present invention, the reagent is a reagent for genome, transcriptome, and / or proteome sequencing.
[0014] The third aspect of the present invention provides the use of a reagent for detecting a combination of biomarkers in the preparation of a kit for predicting and / or diagnosing gastric cancer;
[0015] Among them, the biomarker combination consists of ABLIM1, ABRACL, ACTA1, ACTA2, ACTB, ACTBL2, ACTC1, ACTG1, ACTG2, ACTR3B, ACTR3C, ALDH9A1, ANKHD1, ANKRD17, ARHGDIB, ARHGEF10L, ARMC8, ARPC5, BHMT2, BIN2, CALM1, CALM2, CALM3, CALML3, CAPN1, CAVIN2, CFL1, CHGB, CLIC1, CNN2, CORO1C, CRP, CSRP1, CYFIP1, DLG1, DSC3, DTNB, DUSP3, EIF5A2, EIF5AL1, EIF5B, ELMOD2, EML4, ENO1, ENO2, ENO3, FHL1, FKBP1A, FSCN1, GAPDH, GIT1, GP1BB, GP6, GSTO1, HBE1, HNRNPA0, HSPA6, HSPA7, IGF1, ILK, IMPDH1, ISLR, ITGA2B, ITIH3, ITIH5, LGALSL, LRG1, MAP1A, MAPK8IP3, MCM7, MECP2, MIF, MMP10, MRPL37, MSL1, MTHFD2, MTPN, MYL12B, MYL3, MYL6B, NAXD, NCALD, NIT2, NUP155, ORM1, PARVB, PCBP1, PCBP3, PDCD10, PDLIM1, PEX5, PFN1, PGAM2, PGK2, PKLR, PKM, PMVK, POSTN, POTEE, POTEF, POTEI, POTEJ, POTEKP, PPCS, PPP2R2D, PRDX5, PRKAG1, PSMB8, PTGIS, PTMA, PUDP, RAB1B, RARRES2, RCN1, RDH10, RECK, RPL5, RPS27A, RPS7, S100A9, SAA1, SAA2, SF3A3, SFTPB, SKP1, SMC1A, SMC2, SPG21, SYNE1, SYNE2, SYNM, TAGLN2, TBXAS1, TMSB4X, TNFAIP2, TPM4, TUBB8B, TXN, UBA52, UBB, UBC, VPS13A, VWF, XPNPEP1 and YKT6.
[0016] In some embodiments of the present invention, the reagent is as described in the second aspect.
[0017] The fourth aspect of the present invention provides a biomarker combination, which is composed of ABLIM1, ABRACL, ACTA1, ACTA2, ACTB, ACTBL2, ACTC1, ACTG1, ACTG2, ACTR3B, ACTR3C, ALDH9A1, ANKHD1, ANKRD17, ARHGDIB, ARHGEF10L, ARMC8, ARPC5, BHMT2, BIN2, CALM1, CALM2, CALM3, CALML3, CAPN1, CAVIN2, CFL1, CHGB, CLIC1, CNN2, CORO1C, CRP, CSRP1, CYFIP1, DLG1, DSC3, DTNB, DUSP3, EIF5A2, EIF5AL1, EIF5B, ELMOD2, EML4, ENO1, ENO2, ENO3, FHL1, FKBP1A, FSCN1, GAPDH, GIT1, GP1BB, GP6, GSTO1, HBE1, HNRNPA0, HSPA6, HSPA7, IGF1, ILK, IMPDH1, ISLR, ITGA2B, ITIH3, ITIH5, LGALSL, LRG1, MAP1A, MAPK8IP3, MCM7, MECP2, MIF, MMP10, MRPL37, MSL1, MTHFD2, MTPN, MYL12B, MYL3, MYL6B, NAXD, NCALD, NIT2, NUP155, ORM1, PARVB, PCBP1, PCBP3, PDCD10, PDLIM1, PEX5, PFN1, PGAM2, PGK2, PKLR, PKM, PMVK, POSTN, POTEE, POTEF, POTEI, POTEJ, POTEKP, PPCS, PPP2R2D, PRDX5, PRKAG1, PSMB8, PTGIS, PTMA, PUDP, RAB1B, RARRES2, RCN1, RDH10, RECK, RPL5, RPS27A, RPS7, S100A9, SAA1, SAA2, SF3A3, SFTPB, SKP1, SMC1A, SMC2, SPG21, SYNE1, SYNE2, SYNM, TAGLN2, TBXAS1, TMSB4X, TNFAIP2, TPM4, TUBB8B, TXN, UBA52, UBB, UBC, VPS13A, VWF, XPNPEP1 and YKT6.
[0018] The fifth aspect of the present invention provides a kit, which contains the reagent as described in the second aspect and the biomarker combination as described in the fourth aspect.
[0019] The sixth aspect of the present invention provides a method for detecting gastric cancer for non-diagnostic purposes, the method comprising detecting the expression levels of a biomarker combination in a sample to be tested;
[0020] wherein the biomarker combination consists of ABLIM1, ABRACL, ACTA1, ACTA2, ACTB, ACTBL2, ACTC1, ACTG1, ACTG2, ACTR3B, ACTR3C, ALDH9A1, ANKHD1, ANKRD17, ARHGDIB, ARHGEF10L, ARMC8, ARPC5, BHMT2, BIN2, CALM1, CALM2, CALM3, CALML3, CAPN1, CAVIN2, CFL1, CHGB, CLIC1, CNN2, CORO1C, CRP, CSRP1, CYFIP1, DLG1, DSC3, DTNB, DUSP3, EIF5A2, EIF5AL1, EIF5B, ELMOD2, EML4, ENO1, ENO2, ENO3, FHL1, FKBP1A, FSCN1, GAPDH, GIT1, GP1BB, GP6, GSTO1, HBE1, HNRNPA0, HSPA6, HSPA7, IGF1, ILK, IMPDH1, ISLR, ITGA2B, ITIH3, ITIH5, LGALSL, LRG1, MAP1A, MAPK8IP3, MCM7, MECP2, MIF, MMP10, MRPL37, MSL1, MTHFD2, MTPN, MYL12B, MYL3, MYL6B, NAXD, NCALD, NIT2, NUP155, ORM1, PARVB, PCBP1, PCBP3, PDCD10, PDLIM1, PEX5, PFN1, PGAM2, PGK2, PKLR, PKM, PMVK, POSTN, POTEE, POTEF, POTEI, POTEJ, POTEKP, PPCS, PPP2R2D, PRDX5, PRKAG1, PSMB8, PTGIS, PTMA, PUDP, RAB1B, RARRES2, RCN1, RDH10, RECK, RPL5, RPS27A, RPS7, S100A9, SAA1, SAA2, SF3A3, SFTPB, SKP1, SMC1A, SMC2, SPG21, SYNE1, SYNE2, SYNM, TAGLN2, TBXAS1, TMSB4X, TNFAIP2, TPM4, TUBB8B, TXN, UBA52, UBB, UBC, VPS13A, VWF, XPNPEP1 and YKT6;
[0021] The expression level is the protein expression level and / or the mRNA transcription level.
[0022] In the present invention, the "non-diagnostic purpose" means for the purposes of scientific research and pathological data statistics, and the applicable scenarios include verifying whether an animal model is successfully constructed, in vitro drug efficacy experiments, epidemiological statistics of tumors, etc.
[0023] The seventh aspect of the present invention provides a prediction system for gastric cancer risk. The prediction system includes a detection module and an analysis and judgment module. The detection module detects the expression level of a biomarker combination in a sample to be tested and transmits the expression level data to the analysis and judgment module. The analysis and judgment module processes the expression level data through Firmiana software. The expression level data is preferably FOT (Fraction of total, defined as the iBAQ of this protein divided by the total iBAQ of all identified proteins in the sample). A machine learning algorithm based on a generalized linear regression model is preset to construct a prediction model, respectively predict the probability of the sample having gastric cancer and the probability of not having gastric cancer, judge whether the expression level data meets the preset judgment conditions, so as to predict the risk of the sample having gastric cancer, and output a prediction result. The judgment condition is that the probability of having gastric cancer is greater than or equal to the probability of not having gastric cancer.
[0024] When the expression level data meets the judgment conditions, the output prediction result is "at risk of gastric cancer"; when the expression level data does not meet the judgment conditions, that is, the probability of having gastric cancer is less than the probability of not having gastric cancer, the output prediction result is "not at risk of gastric cancer".
[0025] Among them, the biomarker combination consists of ABLIM1, ABRACL, ACTA1, ACTA2, ACTB, ACTBL2, ACTC1, ACTG1, ACTG2, ACTR3B, ACTR3C, ALDH9A1, ANKHD1, ANKRD17, ARHGDIB, ARHGEF10L, ARMC8, ARPC5, BHMT2, BIN2, CALM1, CALM2, CALM3, CALML3, CAPN1, CAVIN2, CFL1, CHGB, CLIC1, CNN2, CORO1C, CRP, CSRP1, CYFIP1, DLG1, DSC3, DTNB, DUSP3, EIF5A2, EIF5AL1, EIF5B, ELMOD2, EML4, ENO1, ENO2, ENO3, FHL1, FKBP1A, FSCN1, GAPDH, GIT1, GP1BB, GP6, GSTO1, HBE1, HNRNPA0, HSPA6, HSPA7, IGF1, ILK, IMPDH1, ISLR, ITGA2B, ITIH3, ITIH5, LGALSL, LRG1, MAP1A, MAPK8IP3, MCM7, MECP2, MIF, MMP10, MRPL37, MSL1, MTHFD2, MTPN, MYL12B, MYL3, MYL6B, NAXD, NCALD, NIT2, NUP155, ORM1, PARVB, PCBP1, PCBP3, PDCD10, PDLIM1, PEX5, PFN1, PGAM2, PGK2, PKLR, PKM, PMVK, POSTN, POTEE, POTEF, POTEI, POTEJ, POTEKP, PPCS, PPP2R2D, PRDX5, PRKAG1, PSMB8, PTGIS, PTMA, PUDP, RAB1B, RARRES2, RCN1, RDH10, RECK, RPL5, RPS27A, RPS7, S100A9, SAA1, SAA2, SF3A3, SFTPB, SKP1, SMC1A, SMC2, SPG21, SYNE1, SYNE2, SYNM, TAGLN2, TBXAS1, TMSB4X, TNFAIP2, TPM4, TUBB8B, TXN, UBA52, UBB, UBC, VPS13A, VWF, XPNPEP1, and YKT6;
[0026] The expression level is the protein expression level and / or the mRNA transcription level.
[0027] In some embodiments of the present invention, the prediction system is used to process the expression level data through Firmiana software after the reception or input is completed. It is preset with a machine learning algorithm based on a generalized linear regression model to construct the prediction system.
[0028] In some embodiments of the present invention, the sample to be tested is a plasma sample.
[0029] In some embodiments of the present invention, the prediction system further includes a data collection module. The data collection module is used to collect the expression level data of the biomarker combination in the sample to be tested. The expression level data is preferably FOT (Fraction of total, defined as the iBAQ of this protein divided by the total iBAQ of all identified proteins in the sample).
[0030] In some embodiments of the present invention, the prediction system is a system for predicting early gastric cancer.
[0031] The eighth aspect of the present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement the functions of the prediction system as described in the seventh aspect of the present invention, or implement the steps of the method as described in the sixth aspect of the present invention.
[0032] The ninth aspect of the present invention provides an electronic device including a memory and a processor. The memory stores a computer program, and the processor is used to execute the computer program to implement the functions of the prediction system as described in the seventh aspect of the present invention, or implement the steps of the method as described in the sixth aspect of the present invention.
[0033] The present invention establishes a tumor risk model through humoral protein molecules for gastric cancer screening, which helps to achieve early diagnosis of gastric cancer.
[0034] The present invention obtains a set of biomarkers capable of predicting gastric cancer risk by screening the humoral proteome. The screening method includes the following steps:
[0035] (1) Collect humoral samples of healthy people and gastric cancer patients;
[0036] (2) Prepare proteins from the humoral samples of healthy people and gastric cancer patients;
[0037] (3) Detect the expression levels of protein molecules in the humoral samples of healthy people and gastric cancer patients;
[0038] (4) Find the protein group molecules highly expressed specifically in the humors of tumor patients and construct a classifier for discrimination.
[0039] On the basis of conforming to the common knowledge in the art, the above preferred conditions can be combined arbitrarily to obtain various preferred examples of the present invention.
[0040] The reagents and raw materials used in the present invention are all commercially available.
[0041] The positive and progressive effects of the present invention are as follows:
[0042] The above-mentioned protein molecular markers provided by the present invention are found through experiments to have significant changes in the expression levels in the body fluid samples of tumor patients and healthy people. Therefore, the body fluid protein molecular markers provided in the present invention can be used for the risk prediction and detection of tumor patients, and have the advantages of high sensitivity and high specificity, providing favorable technical support for predicting the occurrence and development of gastric cancer.
[0043] Developing a corresponding prediction device based on the protein molecular markers in the body fluid samples of healthy people and gastric cancer patients has broad scientific research value and provides great convenience for early clinical diagnosis, intervention treatment, etc. The sensitivity achieved by the combination of 145 markers provided by the present invention reaches 100%, and the specificity reaches 100%. Description of the Drawings
[0044] Figure 1 It is a schematic diagram of the area under the ROC curve of the marker combination in the training set. I
[0045] Figure 2 It is a schematic diagram of the area under the ROC curve of the marker combination in the test set.
[0046] Figure 3 It is a schematic diagram of the area under the ROC curve of the marker combination in the external validation set.
[0047] Figure 4 It is a schematic diagram of the structure of the system for predicting the risk of gastric cancer.
[0048] Figure 5 It is a schematic diagram of the structure of the electronic device. Detailed Embodiments
[0049] The present invention will be further described below by way of examples, but the present invention is not limited to the scope of the described examples. The experimental methods without specific conditions in the following examples are carried out according to conventional methods and conditions, or selected according to the product instructions.
[0050] The plasma samples of 114 normal people and 49 gastric cancer patients are included in the examples. The design and implementation of this study have been approved and supervised by ethics, and written informed consent has been obtained from all patients.
[0051] Example 1 Screening and Validation of the Combination of Biomarkers for Predicting the Risk of Gastric Cancer
[0052] 1.1 Isolation of Plasma
[0053] Collect whole blood samples in EDTA anticoagulant tubes. After inverting and mixing well, use a 4°C low-temperature centrifuge to centrifuge at 1,600×g for 10 min. After centrifugation, collect the supernatant (plasma) into a new EP tube and centrifuge at 16,000×g for 10 min to remove cell debris. Aliquot the plasma into centrifuge tubes and store at -80°C for later use.
[0054] 1.2 Plasma sample pretreatment
[0055] Add 100 μL of 50 mM ammonium bicarbonate to 2 μL of plasma sample, vortex for 1 min, heat and incubate the sample at 95°C for 4 min to denature the protein. After cooling to room temperature, add 2 μg of trypsin to the system and oscillate at 37°C for 18 h. Then add 10 μL of ammonia water to stop the enzymatic digestion. Desalt the peptide sample after enzymatic digestion, dry it by evaporation, and store at -80°C until mass spectrometry detection.
[0056] 1.3 Mass spectrometry detection of plasma samples
[0057] Use an Orbitrap Fusion Lumos triple quadrupole high-resolution mass spectrometry system (Thermo Fisher Scientific, Rockford, USA) in tandem with a high-performance liquid chromatography system (EASY-nLC 1200, Thermo Fisher) for detection and obtain the mass spectrometry data of the whole protein corresponding to the peptide sample. The specific operation is as follows:
[0058] Adopt nano-liquid chromatography with a self-made C18 chromatographic column (150 μm ID × 8 cm, 1.9 μm / packing material). The column oven temperature is 60°C. Reconstitute the dry powder peptide with the loading buffer (aqueous solution of 0.1% formic acid). After loading, separate it through the chromatographic column and elute with a linear 6–30% mobile phase B (ACN and 0.1% formic acid) at 600 nL / min, using a 10-min liquid phase gradient combined with data-independent acquisition (DIA) mass spectrometry detection method. The DIA mass spectrometry detection parameters are set as follows: the ion mode is positive ion; the resolution of the first-stage mass spectrometry is 30K, the maximum injection time is 20 ms, the AGC Target is 3e6, and the scanning range is 300 - 1400 m / z; the resolution of the second-stage scanning is 15K, 30 variable isolation windows are obtained, and the collision energy is 27%. The liquid chromatography tandem mass spectrometry system is controlled by Xcalibur software for data acquisition.
[0059] 1.4 Data analysis
[0060] All data was processed using Firmiana (V1.0). Firmiana is a workflow based on the Galaxy system, consisting of multiple functional modules such as a user login interface, raw data, identification and quantification, data analysis, and knowledge mining. DIA data was searched against the UniProt human protein database (updated on December 17, 2019, with 20,406 entries) using DIANN (v12.1). The mass difference of precursor ions was 20 ppm, and the mass difference of product ions was 50 mmu. Up to two missed cleavage sites were allowed. The search engine set carbamidomethylation of cysteine as a fixed modification and N-acetylation and oxidation of methionine as variable modifications. The precursor ion charge range was set to +2, +3, and +4. The False Discovery Rate (FDR) was set to 1%. The results of DIA data were merged into the reference library using SpectraST software. A total of 327 libraries were used as the reference library.
[0061] The quantitative results of the identified peptides were recorded as the average of the chromatographic fragment ion peak areas in all reference spectral libraries. Protein quantification was performed using label-free intensity-based absolute quantification (iBAQ) method. The peak area values were calculated as part of the corresponding protein. The total score (FOT) was used to represent the normalized abundance of a specific protein in the sample. FOT was defined as the iBAQ of the protein divided by the total iBAQ of all identified proteins in the sample. Proteins with at least one unique peptide and 1% FDR were selected. The FOT of each protein was calculated, and the FOT of each protein was used as the protein expression data and input into the generalized linear regression model.
[0062] In this example, the selected Firmiana was preset with a machine learning algorithm based on the generalized linear regression model to construct a prediction model (the FOT values of the training set are shown in Table 1-12, the FOT values of the test set are shown in Table 13-24, and the FOT values of the external validation set are shown in Table 25-36), and the probabilities of the sample having gastric cancer and not having gastric cancer were predicted respectively. The code for constructing the prediction model is as follows:
[0063] from sklearn.linear_model import LogisticRegressionCV
[0064] import joblib
[0065] import pandas as pd
[0066] tumor_types = ['***', '***', '***', '***', '***'] # *** represents the 145 protein molecular markers described in the present invention
[0067] for tumor_type in tumor_types:
[0068] df_train = pd.read_csv('{}_train.csv'.format(tumor_type))
[0069] df_test = pd.read_csv('{}_test.csv'.format(tumor_type))
[0070] df_val = pd.read_csv('{}_validation.csv'.format(tumor_type))
[0071] X_train = df_train.iloc[:, 2:]
[0072] y_train = df_train.iloc[:, 1]
[0073] X_test = df_test.iloc[:, 2:]
[0074] y_test = df_test.iloc[:, 1]
[0075] X_val = df_val.iloc[:, 2:]
[0076] y_val = df_val.iloc[:, 1]
[0077] model = LogisticRegressionCV()
[0078] model.fit(X_train, y_train)
[0079] joblib.dump(model, f'{tumor_type}_model.joblib');
[0080] Experimental findings show that there are significant changes in the expression levels of some proteins in the body fluid samples of cancer patients and healthy individuals. The ROC curve (Receiver Operating Curve) was plotted for the relative expression levels of 145 protein molecular markers (ABLIM1, ABRACL, ACTA1, ACTA2, ACTB, ACTBL2, ACTC1, ACTG1, ACTG2, ACTR3B, ACTR3C, ALDH9A1, ANKHD1, ANKRD17, ARHGDIB, ARHGEF10L, ARMC8, ARPC5, BHMT2, BIN2, CALM1, CALM2, CALM3, CALML3, CAPN1, CAVIN2, CFL1, CHGB, CLIC1, CNN2, CORO1C, CRP, CSRP1, CYFIP1, DLG1, DSC3, DTNB, DUSP3, EIF5A2, EIF5AL1, EIF5B, ELMOD2, EML4, ENO1, ENO2, ENO3, FHL1, FKBP1A, FSCN1, GAPDH, GIT1, GP1BB, GP6, GSTO1, HBE1, HNRNPA0, HSPA6, HSPA7, IGF1, ILK, IMPDH1, ISLR, ITGA2B, ITIH3, ITIH5, LGALSL, LRG1, MAP1A, MAPK8IP3, MCM7, MECP2, MIF, MMP10, MRPLAmong them, the training set included 24 positive cases and 55 negative cases, with AUC = 0.99, diagnostic sensitivity of 83.33%, and specificity of 100.00% (see, Figure 1 ); the test set included the remaining 10 positive cases and 25 negative cases, with AUC = 0.99, diagnostic sensitivity of 90.00%, and specificity of 92.00% (see Figure 2 ); the external validation set included 15 positive cases and 34 negative cases, with AUC = 0.96, diagnostic sensitivity of 80.00%, and specificity of 94.12% ( Figure 3 ). The analysis method refers to Karimollah Hajian-Tilaki, Receiver Operating Characteristic (ROC) Curve Analysis for Medical Diagnostic Test Evaluation, Caspian J Intern Med 2013; 4(2): 627-635. For unknown samples, substitute the expression levels of the above biomarkers into the model to obtain the gastric cancer risk prediction for the sample and output the result. When the probability of having gastric cancer is greater than or equal to the probability of not having gastric cancer, the predicted result is output as "at risk of gastric cancer"; when the probability of having gastric cancer is less than the probability of not having gastric cancer, the predicted result is output as "not at risk of gastric cancer".
[0081] As can be seen from the above results, the combination of 145 protein molecular markers (ABLIM1, ABRACL, ACTA1, ACTA2, ACTB, ACTBL2, ACTC1, ACTG1, ACTG2, ACTR3B, ACTR3C, ALDH9A1, ANKHD1, ANKRD17, ARHGDIB, ARHGEF10L, ARMC8, ARPC5, BHMT2, BIN2, CALM1, CALM2, CALM3, CALML3, CAPN1, CAVIN2, CFL1, CHGB, CLIC1, CNN2, CORO1C, CRP, CSRP1, CYFIP1, DLG1, DSC3, DTNB, DUSP3, EIF5A2, EIF5AL1, EIF5B, ELMOD2, EML4, ENO1, ENO2, ENO3, FHL1, FKBP1A, FSCN1, GAPDH, GIT1, GP1BB, GP6, GSTO1, HBE1, HNRNPA0, HSPA6, HSPA7, IGF1, ILK, IMPDH1, ISLR, ITGA2B, ITIH3, ITIH5, LGALSL, LRG1, MAP1A, MAPK8IP3, MCM7, MECP2, MIF, MMP10, MRPL37, MSL1, MTHFD2, MTPN, MYL12B, MYL3, MYL6B, NAXD, NCALD, NIT2, NUP155, ORM1, PARVB, PCBP1, PCBP3, PDCD10, PDLIM1, PEX5, PFN1, PGAM2, PGK2, PKLR, PKM, PMVK, POSTN, POTEE, POTEF, POTEI, POTEJ, POTEKP, PPCS, PPP2R2D, PRDX5, PRKAG1, PSMB8, PTGIS, PTMA, PUDP, RAB1B, RARRES2, RCN1, RDH10, RECK, RPL5, RPS27A, RPS7, S100A9, SAA1, SAA2, SF3A3, SFTPB, SKP1, SMC1A, SMC2, SPG21, SYNE1, SYNE2, SYNM, TAGLN2, TBXAS1, TMSB4X, TNFAIP2, TPM4, TUBB8B, TXN, UBA52, UBB, UBC, VPS13A, VWF, XPNPEP1 and YKT6) in the plasma of tumor patients can be used to predict tumor risk.
[0082] FOT values of 145 protein markers in the training set
[0083]
[0084]
[0085]
[0086] FOT values of 145 protein markers in the training set in Table 2
[0087]
[0088]
[0089]
[0090] FOT values of 145 protein markers in the training set in Table 3
[0091]
[0092]
[0093]
[0094] FOT values of 145 protein markers in the training set in Table 4
[0095]
[0096]
[0097] FOT values of 145 protein markers in the training set in Table 5
[0098]
[0099]
[0100]
[0101] FOT values of 145 protein markers in the training set in Table 6
[0102]
[0103]
[0104]
[0105] FOT values of 145 protein markers in the training set in Table 7
[0106]
[0107]
[0108] FOT values of 145 protein markers in the training set
[0109]
[0110]
[0111] FOT values of 145 protein markers in the training set
[0112]
[0113]
[0114]
[0115] FOT values of 145 protein markers in the training set
[0116]
[0117]
[0118]
[0119] FOT values of 145 protein markers in the training set
[0120]
[0121]
[0122] FOT values of 145 protein markers in the training set
[0123]
[0124]
[0125] FOT values of 145 protein markers in the test set
[0126]
[0127]
[0128] FOT values of 145 protein markers in the test set
[0129]
[0130]
[0131] FOT values of 145 protein markers in the test set
[0132]
[0133]
[0134] FOT values of 145 protein markers in the test set, Table 16
[0135]
[0136] FOT values of 145 protein markers in the test set, Table 17
[0137]
[0138]
[0139] FOT values of 145 protein markers in the test set, Table 18
[0140]
[0141]
[0142] FOT values of 145 protein markers in the test set, Table 19
[0143]
[0144]
[0145] FOT values of 145 protein markers in the test set, Table 20
[0146]
[0147]
[0148] FOT values of 145 protein markers in the test set, Table 21
[0149]
[0150] FOT values of 145 protein markers in the test set, Table 22
[0151]
[0152]
[0153] FOT values of 145 protein markers in the test set, Table 23
[0154]
[0155]
[0156] FOT values of 145 protein markers in the test set, Table 24
[0157]
[0158]
[0159] FOT values of 145 protein markers in the external validation set, Table 25
[0160]
[0161]
[0162] FOT values of 145 protein markers in the external validation set, Table 26
[0163]
[0164]
[0165] FOT values of 145 protein markers in the external validation set, Table 27
[0166]
[0167]
[0168] FOT values of 145 protein markers in the external validation set, Table 28
[0169]
[0170]
[0171] FOT values of 145 protein markers in the external validation set, Table 29
[0172]
[0173]
[0174] FOT values of 145 protein markers in the external validation set, Table 30
[0175]
[0176]
[0177] FOT values of 145 protein markers in the external validation set, Table 31
[0178]
[0179]
[0180] Table 32 FOT values of 145 protein markers in the external validation set
[0181]
[0182]
[0183] Table 33 FOT values of 145 protein markers in the external validation set
[0184]
[0185]
[0186] Table 34 FOT values of 145 protein markers in the external validation set
[0187]
[0188]
[0189] Table 35 FOT values of 145 protein markers in the external validation set
[0190]
[0191]
[0192] Table 36 FOT values of 145 protein markers in the external validation set
[0193]
[0194]
[0195] Example 2 System for predicting gastric cancer risk
[0196] System 61 for predicting gastric cancer risk: data processing module 52 and judgment and output module 53, further including data collection module 51( Figure 4 ).
[0197] The data collection module 51 is used to collect the expression level data of the biomarker combination in the gastric cancer tissue sample of the patient and transmit it to the data processing module.
[0198] The data processing module 52 is used to analyze the expression level data of the received or input biomarker combination according to the data analysis method described in Embodiment 4 to obtain a calculation result. Among them, the expression level data of the biomarker combination can be collected by the data collection module 51, or the expression level data of the biomarker combination can be obtained from other sources.
[0199] The judgment and output module 53 is used to judge whether the calculation result meets a preset judgment condition, that is, the risk probability of suffering from gastric cancer is greater than or equal to the risk prediction probability of not suffering from gastric cancer, so as to predict the risk of gastric cancer and output a prediction result; among them, in the judgment and output module, when the expression level data meets the judgment condition that the risk probability of suffering from gastric cancer is greater than or equal to the risk prediction probability of not suffering from gastric cancer, the output prediction result is "at risk of gastric cancer"; when the expression level data does not meet the judgment condition that the risk probability of suffering from gastric cancer is less than the risk prediction probability of not suffering from gastric cancer, the output prediction result is "not at risk of gastric cancer".
[0200] Embodiment 3 Electronic device
[0201] This embodiment provides an electronic device, which can be presented in the form of a computing device (for example, it can be a server device), including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for predicting the risk of gastric cancer in Embodiment 1 of the present invention can be implemented.
[0202] Figure 5 The schematic diagram of the hardware structure of this embodiment is shown. The electronic device 4 specifically includes:
[0203] At least one processor 91, at least one memory 92, and a bus 93 for connecting different system components (including the processor 91 and the memory 92), where:
[0204] The bus 93 includes a data bus, an address bus, and a control bus.
[0205] The memory 92 includes a volatile memory, such as a random access memory (RAM) 921 and / or a cache memory 922, and may further include a read-only memory (ROM) 923.
[0206] The memory 92 further includes a program / utilities 925 having a set (at least one) of program modules 924. Such program modules 924 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0207] The processor 91 executes various functional applications and data processing by running the computer program stored in the memory 92, such as the data analysis method in Embodiment 1 of the present invention.
[0208] The electronic device 9 can further communicate with one or more external devices 94 (such as a keyboard, a pointing device, etc.). Such communication can be carried out through the input / output (I / O) interface 95. Moreover, the electronic device 9 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 96. The network adapter 96 communicates with other modules of the electronic device 9 through the bus 93. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 9, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (redundant array of independent disks) systems, tape drives, and data backup storage systems, etc.
[0209] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-described units / modules can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.
[0210] Embodiment 4 Computer-readable storage medium
[0211] The embodiments of the present invention provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method for predicting the risk of gastric cancer in Embodiment 1 of the present invention are implemented.
[0212] Among them, the more specific forms that the readable storage medium can adopt can include but are not limited to: portable disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories, optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0213] In a possible implementation manner, the present invention can also be implemented in the form of a program product, which includes program code, and when the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps of the method for predicting the risk of gastric cancer in Embodiment 1 of the present invention.
[0214] Among them, the program code for implementing the present invention can be written in any combination of one or more programming languages, and the program code can be executed entirely on the user device, partially on the user device, executed as an independent software package, partially on the user device and partially on a remote device, or entirely on a remote device.
[0215] Finally, the above specific implementation methods are only used to illustrate the technical solutions of the present invention, rather than limiting it. Biomarker names: (reference can be made to the NCBI or genecards databases)
[0216] ABLIM1: actin binding LIM protein 1, Gene ID: 3983; ABRACL: ABRA C-terminallike, Gene ID: 58527
[0217] ACTA1: actin alpha 1, Gene ID: 58
[0218] ACTA2: actin alpha 2, Gene ID: 59
[0219] ACTB: actin beta, Gene ID: 60
[0220] ACTBL2: actin beta like 2, Gene ID: 345651
[0221] ACTC1: Actin Alpha Cardiac Muscle 1, Gene ID: 70
[0222] ACTG1: Actin Gamma 1, Gene ID: 71
[0223] ACTG2: Actin Gamma 2, Gene ID: 72
[0224] ACTR3B: actin related protein 3B, Gene ID: 57180
[0225] ACTR3C: actin related protein 3C, Gene ID: 653857
[0226] ALDH9A1: aldehyde dehydrogenase 9 family member A1, Gene ID: 223; ANKHD1: ankyrin repeat and KH domain containing 1, Gene ID: 54882; ANKRD17: ankyrin repeat domain 17, Gene ID: 26057; ARHGDIB: Rho GDP dissociation inhibitor beta, Gene ID: 397
[0227] ARHGEF10L: Rho guanine nucleotide exchange factor 10 like, Gene ID: 55160; ARMC8: Armadillo Repeat Containing 8, Gene ID: 25852; ARPC5: actin related protein 2 / 3 complex subunit 5, Gene ID: 10092; BHMT2: betaine--homocysteine S-methyltransferase 2, Gene ID: 23743; BIN2: bridging integrator 2, Gene ID: 51411
[0228] CALM1: calmodulin 1, Gene ID: 801
[0229] CALM2: calmodulin 2, Gene ID: 805
[0230] CALM3: calmodulin 3, Gene ID: 808
[0231] CALML3: calmodulin like 3, Gene ID: 810
[0232] CAPN1: calpain 1, Gene ID: 823
[0233] CAVIN2: caveolae associated protein 2, Gene ID: 8436; CFL1: cofilin 1, Gene ID: 1072
[0234] CHGB: chromogranin B, Gene ID: 1114
[0235] CLIC1: chloride intracellular channel 1, Gene ID: 1192; CNN2: calponin 2, Gene ID: 1265
[0236] CORO1C: coronin 1C, Gene ID: 23603
[0237] CRP: C-reactive protein, Gene ID: 1401
[0238] CSRP1: cysteine and glycine rich protein 1, Gene ID: 1465
[0239] CYFIP1: cytoplasmic FMR1 interacting protein 1, Gene ID: 23191; DLG1: discslarge MAGUK scaffold protein 1, Gene ID: 1739; DSC3: desmocollin 3, Gene ID: 1825
[0240] DTNB: dystrobrevin beta, Gene ID: 1838
[0241] DUSP3: dual specificity phosphatase 3, Gene ID: 1845
[0242] EIF5A2: eukaryotic translation initiation factor 5A2, Gene ID: 56648
[0243] EIF5AL1: eukaryotic translation initiation factor 5A like 1, Gene ID: 143244; EIF5B: eukaryotic translation initiation factor 5B, Gene ID: 9669; ELMOD2: ELMO domain containing 2, Gene ID: 255520
[0244] EML4: EMAP like 4, Gene ID: 27436
[0245] ENO1: enolase 1, Gene ID: 2023
[0246] ENO2: enolase 2, Gene ID: 2026
[0247] ENO3: enolase 3, Gene ID: 2027
[0248] FHL1: four and a half LIM domains 1, Gene ID: 2273
[0249] FKBP1A: FKBP prolyl isomerase 1A, Gene ID: 2280
[0250] FSCN1: fascin actin - bundling protein 1, Gene ID: 6624
[0251] GAPDH: glyceraldehyde - 3 - phosphate dehydrogenase, Gene ID: 2597
[0252] GIT1: GIT ArfGAP 1, Gene ID: 28964
[0253] GP1BB: glycoprotein Ib platelet subunit beta, Gene ID: 2812 GP6: glycoprotein VI platelet, Gene ID: 51206
[0254] GSTO1: glutathione S - transferase omega 1, Gene ID: 9446
[0255] HBE1: hemoglobin subunit epsilon 1, Gene ID: 3046
[0256] HNRNPA0: heterogeneous nuclear ribonucleoprotein A0, Gene ID: 10949
[0257] HSPA6: heat shock protein family A(Hsp70)member 6, Gene ID: 3310
[0258] HSPA7: heat shock protein family A (Hsp70) member 7 (pseudogene), Gene ID: 3311; IGF1: insulin like growth factor 1, Gene ID: 3479
[0259] ILK: integrin linked kinase, Gene ID: 3611
[0260] IMPDH1: inosine monophosphate dehydrogenase 1, Gene ID: 3614; ISLR: immunoglobulin superfamily containing leucine rich repeat, Gene ID: 3671; ITGA2B: integrin subunit alpha 2b, Gene ID: 3674
[0261] ITIH3: inter-alpha-trypsin inhibitor heavy chain 3, Gene ID: 3699; ITIH5: inter-alpha-trypsin inhibitor heavy chain 5, Gene ID: 80760; LGALSL: galectinlike, Gene ID: 29094
[0262] LRG1: leucine rich alpha-2-glycoprotein 1, Gene ID: 116844; MAP1A: microtubule associated protein 1A, Gene ID: 4130
[0263] MAPK8IP3: mitogen-activated protein kinase 8 interacting protein 3, Gene ID: 23162; MCM7: minichromosome maintenance complex component 7, Gene ID: 4176
[0264] MECP2: methyl-CpG binding protein 2, Gene ID: 4204
[0265] MIF: macrophage migration inhibitory factor, Gene ID: 4282
[0266] MMP10: matrix metallopeptidase 10, Gene ID: 4319
[0267] MRPL37: mitochondrial ribosomal protein L37, Gene ID: 51253
[0268] MSL1: MSL complex subunit 1, Gene ID: 339287
[0269] MTHFD2: methylenetetrahydrofolate dehydrogenase (NADP+ dependent) 2, GeneID: 10797 MTPN: myotrophin, Gene ID: 136319
[0270] MYL12B: myosin light chain 12B, Gene ID: 103910
[0271] MYL3: myosin light chain 3, Gene ID: 4634
[0272] MYL6B: myosin light chain 6B, Gene ID: 140465
[0273] NAXD: NAD(P)HX dehydratase, Gene ID: 55739
[0274] NCALD: neurocalcin delta, Gene ID: 83988
[0275] NIT2: nitrilase family member 2, Gene ID: 56954
[0276] NUP155: nucleoporin 155, Gene ID: 9631
[0277] ORM1: orosomucoid 1, Gene ID: 5004
[0278] PARVB: parvin beta, Gene ID: 29780
[0279] PCBP1: poly(rC) binding protein 1, Gene ID: 5093
[0280] PCBP3: poly(rC) binding protein 3, Gene ID: 54039
[0281] PDCD10: programmed cell death 10, Gene ID: 11235
[0282] PDLIM1: PDZ and LIM domain 1, Gene ID: 9124
[0283] PEX5: peroxisomal biogenesis factor 5, Gene ID: 5830
[0284] PFN1: profilin 1, Gene ID: 5216
[0285] PGAM2: phosphoglycerate mutase 2, Gene ID: 5224
[0286] PGK2: phosphoglycerate kinase 2, Gene ID: 5232
[0287] PKLR: pyruvate kinase L / R, Gene ID: 5313
[0288] PKM: pyruvate kinase M1 / 2, Gene ID: 5315
[0289] PMVK: phosphomevalonate kinase, Gene ID: 10654
[0290] POSTN: periostin, Gene ID: 10631
[0291] POTEE: POTE ankyrin domain family member E, Gene ID: 445582 POTEF: POTE ankyrin domain family member F, Gene ID: 728378 POTEI: POTE ankyrin domain family member I, Gene ID: 653269 POTEJ: POTE ankyrin domain family member J, Gene ID: 653781 POTEKP: POTE ankyrin domain family member K, pseudogene, Gene ID: 440915 PPCS: phosphopantothenoylcysteine synthetase, Gene ID: 79717 PPP2R2D: protein phosphatase 2 regulatory subunit B delta, Gene ID: 55844
[0292] PRDX5: peroxiredoxin 5, Gene ID: 25824
[0293] PRKAG1: protein kinase AMP-activated non-catalytic subunit gamma 1, Gene ID: 5571 PSMB8: proteasome 20S subunit beta 8, Gene ID: 5696
[0294] PTGIS: prostaglandin I2 synthase, Gene ID: 5740
[0295] PTMA: prothymosin alpha, Gene ID: 5757
[0296] PUDP: pseudouridine 5'-phosphatase, Gene ID: 8226
[0297] RAB1B: RAB1B, member RAS oncogene family, Gene ID: 81876 RARRES2: retinoic acid receptor responder 2, Gene ID: 5919 RCN1: reticulocalbin 1, Gene ID: 5954
[0298] RDH10: retinol dehydrogenase 10, Gene ID: 157506
[0299] RECK: reversion inducing cysteine rich protein with kazal motifs, GeneID: 8434RPL5: ribosomal protein L5, Gene ID: 6125
[0300] RPS27A: ribosomal protein S27a, Gene ID: 6233
[0301] RPS7: ribosomal protein S7, Gene ID: 6201
[0302] S100A9: S100 calcium binding protein A9, Gene ID: 6280
[0303] SAA1: serum amyloid A1, Gene ID: 6288
[0304] SAA2: serum amyloid A2, Gene ID: 6289
[0305] SF3A3: splicing factor 3a subunit 3, Gene ID: 10946
[0306] SFTPB: surfactant protein B, Gene ID: 6439
[0307] SKP1: S-phase kinase associated protein 1, Gene ID: 6500
[0308] SMC1A: structural maintenance of chromosomes 1A, Gene ID: 8243SMC2: structural maintenance of chromosomes 2, Gene ID: 10592SPG21: SPG21 abhydrolasedomain containing, maspardin, Gene ID: 51324
[0309] SYNE1: spectrin repeat containing nuclear envelope protein 1, Gene ID: 23345 SYNE2: spectrin repeat containing nuclear envelope protein 2, Gene ID: 23224 SYNM: synemin, Gene ID: 23336
[0310] TAGLN2: transgelin 2, Gene ID: 8407
[0311] TBXAS1: thromboxane A synthase 1, Gene ID: 6916
[0312] TMSB4X: thymosin beta 4 X-linked, Gene ID: 7114
[0313] TNFAIP2: TNF alpha induced protein 2, Gene ID: 7127
[0314] TPM4: tropomyosin 4, Gene ID: 7171
[0315] TUBB8B: tubulin beta 8B, Gene ID: 260334
[0316] TXN: thioredoxin, Gene ID: 7295
[0317] UBA52: ubiquitin A-52 residue ribosomal protein fusion product 1, GeneID: 7311 UBB: ubiquitin B, Gene ID: 7314
[0318] UBC: ubiquitin C, Gene ID: 7316
[0319] VPS13A: vacuolar protein sorting 13 homolog A, Gene ID: 23230
[0320] VWF: von Willebrand factor, Gene ID: 7450
[0321] XPNPEP1: X-prolyl aminopeptidase 1, Gene ID: 7511
[0322] YKT6: YKT6 v-SNARE homolog, Gene ID: 10652.
Claims
1. A reagent for detecting a biomarker combination, characterized in that, The biomarker combination consists of ABLIM1, ABRACL, ACTA1, ACTA2, ACTB, ACTBL2, ACTC1, ACTG1, ACTG2, ACTR3B, ACTR3C, ALDH9A1, ANKHD1, ANKRD17, ARHGDIB, ARHGEF10L, ARMC8, ARPC5, BHMT2, BIN2, CALM1, CALM2, CALM3, CALML3, CAPN1, CAVIN2, CFL1, CHGB, CLIC1, CNN2, CORO1C, CRP, CSRP1, CYFIP1, DLG1, DSC3, DTNB, DUSP3, EIF5A2, EIF5AL1, EIF5B, ELMOD2, EML4, ENO1, ENO2, ENO3, FHL1, FKBP1A, FSCN1, GAPDH, GIT1, GP1BB, GP6, GSTO1, HBE1, HNRNPA0, HSPA6, HSPA7, IGF1, ILK, IMPDH1, ISLR, ITGA2B, ITIH3, ITIH5, LGALSL, LRG1, MAP1A, MAPK8IP3, MCM7, MECP2, MIF, MMP10, MRPL37, MSL1, MTHFD2, MTPN, MYL12B, MYL3, MYL6B, NAXD, NCALD, NIT2, NUP155, ORM1, PARVB, PCBP1, PCBP3, PDCD10, PDLIM1, PEX5, PFN1, PGAM2, PGK2, PKLR, PKM, PMVK, POSTN, POTEE, POTEF, POTEI, POTEJ, POTEKP, PPCS, PPP2R2D, PRDX5, PRKAG1, PSMB8, PTGIS, PTMA, PUDP, RAB1B, RARRES2, RCN1, RDH10, RECK, RPL5, RPS27A, RPS7, S100A9, SAA1, SAA2, SF3A3, SFTPB, SKP1, SMC1A, SMC2, SPG21, SYNE1, SYNE2, SYNM, TAGLN2, TBXAS1, TMSB4X, TNFAIP2, TPM4, TUBB8B, TXN, UBA52, UBB, UBC, VPS13A, VWF, XPNPEP1, and YKT6.
2. The reagent according to claim 1, wherein The reagent is used to detect the expression level of the biomarker combination, and the expression level is the protein expression level and / or the mRNA transcription level.
3. The reagent according to claim 2, wherein The reagent is a biomolecular reagent that specifically binds to the biomarker or specifically hybridizes with the nucleic acid encoding the biomarker; or, the reagent is a reagent for genome, transcriptome, and / or proteome sequencing.
4. The reagent according to claim 3, wherein The biomolecular reagent is selected from primers, probes and antibodies.
5. A biomarker combination, characterized in that, The biomarker combination consists of ABLIM1, ABRACL, ACTA1, ACTA2, ACTB, ACTBL2, ACTC1, ACTG1, ACTG2, ACTR3B, ACTR3C, ALDH9A1, ANKHD1, ANKRD17, ARHGDIB, ARHGEF10L, ARMC8, ARPC5, BHMT2, BIN2, CALM1, CALM2, CALM3, CALML3, CAPN1, CAVIN2, CFL1, CHGB, CLIC1, CNN2, CORO1C, CRP, CSRP1, CYFIP1, DLG1, DSC3, DTNB, DUSP3, EIF5A2, EIF5AL1, EIF5B, ELMOD2, EML4, ENO1, ENO2, ENO3, FHL1, FKBP1A, FSCN1, GAPDH, GIT1, GP1BB, GP6, GSTO1, HBE1, HNRNPA0, HSPA6, HSPA7, IGF1, ILK, IMPDH1, ISLR, ITGA2B, ITIH3, ITIH5, LGALSL, LRG1, MAP1A, MAPK8IP3, MCM7, MECP2, MIF, MMP10, MRPL37, MSL1, MTHFD2, MTPN, MYL12B, MYL3, MYL6B, NAXD, NCALD, NIT2, NUP155, ORM1, PARVB, PCBP1, PCBP3, PDCD10, PDLIM1, PEX5, PFN1, PGAM2, PGK2, PKLR, PKM, PMVK, POSTN, POTEE, POTEF, POTEI, POTEJ, POTEKP, PPCS, PPP2R2D, PRDX5, PRKAG1, PSMB8, PTGIS, PTMA, PUDP, RAB1B, RARRES2, RCN1, RDH10, RECK, RPL5, RPS27A, RPS7, S100A9, SAA1, SAA2, SF3A3, SFTPB, SKP1, SMC1A, SMC2, SPG21, SYNE1, SYNE2, SYNM, TAGLN2, TBXAS1, TMSB4X, TNFAIP2, TPM4, TUBB8B, TXN, UBA52, UBB, UBC, VPS13A, VWF, XPNPEP1 and YKT6.
6. A kit, characterized in that, The kit contains the reagent as described in claim 1 and the biomarker combination as described in claim 5.
7. Use of the reagent as described in any one of claims 1-4, the biomarker combination as described in claim 5, or the kit as described in claim 6 in the preparation of a product for predicting and / or diagnosing gastric cancer.
8. A prediction system for gastric cancer risk, characterized in that, The prediction system includes a detection module and an analysis and judgment module; the detection module detects the expression levels of biomarker combinations in a sample to be tested and transmits the expression level data to the analysis and judgment module; the analysis and judgment module processes the expression level data through Firmiana software, which is preset as a machine learning algorithm based on a generalized linear regression model, constructs a prediction model, predicts the probabilities of the sample having gastric cancer and not having gastric cancer respectively, determines whether the expression level data meets the preset judgment conditions, so as to predict the risk of the sample having gastric cancer, and outputs a prediction result; The judgment condition is that the probability of having gastric cancer is greater than or equal to the probability of not having gastric cancer; When the expression level data meets the judgment condition, the output prediction result is "at risk of gastric cancer"; when the expression level data does not meet the judgment condition, that is, the probability of having gastric cancer is less than the probability of not having gastric cancer, the output prediction result is "not at risk of gastric cancer"; Among them, the biomarker combination consists of ABLIM1, ABRACL, ACTA1, ACTA2, ACTB, ACTBL2, ACTC1, ACTG1, ACTG2, ACTR3B, ACTR3C, ALDH9A1, ANKHD1, ANKRD17, ARHGDIB, ARHGEF10L, ARMC8, ARPC5, BHMT2, BIN2, CALM1, CALM2, CALM3, CALML3, CAPN1, CAVIN2, CFL1, CHGB, CLIC1, CNN2, CORO1C, CRP, CSRP1, CYFIP1, DLG1, DSC3, DTNB, DUSP3, EIF5A2, EIF5AL1, EIF5B, ELMOD2, EML4, ENO1, ENO2, ENO3, FHL1, FKBP1A, FSCN1, GAPDH, GIT1, GP1BB, GP6, GSTO1, HBE1, HNRNPA0, HSPA6, HSPA7, IGF1, ILK, IMPDH1, ISLR, ITGA2B, ITIH3, ITIH5, LGALSL, LRG1, MAP1A, MAPK8IP3, MCM7, MECP2, MIF, MMP10, MRPL37, MSL1, MTHFD2, MTPN, MYL12B, MYL3, MYL6B, NAXD, NCALD, NIT2, NUP155, ORM1, PARVB, PCBP1, PCBP3, PDCD10, PDLIM1, PEX5, PFN1, PGAM2, PGK2, PKLR, PKM, PMVK, POSTN, POTEE, POTEF, POTEI, POTEJ, POTEKP, PPCS, PPP2R2D, PRDX5, PRKAG1, PSMB8, PTGIS, PTMA, PUDP, RAB1B, RARRES2, RCN1, RDH10, RECK, RPL5, RPS27A, RPS7, S100A9, SAA1, SAA2, SF3A3, SFTPB, SKP1, SMC1A, SMC2, SPG21, SYNE1, SYNE2, SYNM, TAGLN2, TBXAS1, TMSB4X, TNFAIP2, TPM4, TUBB8B, TXN, UBA52, UBB, UBC, VPS13A, VWF, XPNPEP1 and YKT6, and the expression level is the protein expression level and / or the mRNA transcription level.
9. The prediction system according to claim 8, wherein The sample to be tested is a human plasma sample; and / or, the prediction system further includes a data collection module, and the data collection module is used to collect the expression level data of the biomarker combination in the sample to be tested.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it can implement the functions of the prediction system as described in claim 8 or 9.
11. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor is used to execute the computer program to implement the functions of the prediction system as described in claim 8 or 9.
Citation Information
Patent Citations
Biomarker(s) for early detection / diagnosis / prognosis of gastric cancer
US20140287939A1
Compositions and methods for screening, monitoring and treating gastrointestinal diseases
US20190154692A1