Application of the biomarker UGT2B15 in the diagnosis of colorectal adenoma
By using UGT2B15 biomarkers and machine learning methods, a diagnostic model of colorectal adenoma was constructed, which solved the problems of strong invasiveness, low detection rate and high false positives in the prior art, and achieved high sensitivity and specificity of colorectal adenoma screening.
Patent Information
- Application Number
- CN202411530046.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2024-06-27
- Filing Date
- 2024-10-30
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2044-10-30
AI Technical Summary
The existing colorectal adenoma screening methods are highly invasive, have low detection rate, high false positive rate, and lack high sensitivity and specific diagnostic methods.
UGT2B15 biomarkers and their detection reagents are used, combined with machine learning methods, to construct kits and electronic devices for diagnosing or predicting colorectal adenomas, and to construct predictive models by detecting the expression of genes or proteins in the sample.
It improves the diagnostic sensitivity and specificity of colorectal adenomas, provides non-invasive and efficient screening methods, and reduces the false positive rate.
Smart Images

Figure CN119193841B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biotechnology and relates to the application of the biomarker UGT2B15 in the diagnosis of colorectal adenomas. Background Art
[0002] Colorectal adenoma is considered to be the main precancerous lesion of colorectal cancer. It is generally believed that it takes 7 to 10 years for adenoma to develop into cancer, which leaves enough time and space for precise intervention and treatment of colorectal cancer. Therefore, early detection of colorectal adenoma and timely and effective intervention are crucial.
[0003] At present, the screening methods for colorectal adenomas mainly include the following: colonoscopy is an invasive procedure with low patient compliance. At the same time, colonoscopy resources are limited and affected by endoscopic examination technology, resulting in a low detection rate of colorectal adenomas and a high missed diagnosis rate; stool testing is easily affected by food, drugs or other intestinal diseases, resulting in a high false positive rate; tumor marker testing is often used for colorectal cancer screening, but has low sensitivity and specificity for colorectal adenomas; CT colonography technology has radiation hazards and there is currently no standardized method, resulting in low performance in diagnosing adenomas. Summary of the Invention
[0004] In order to solve the technical problems existing in the prior art, the present invention provides the following technical solutions:
[0005] The present invention provides the use of a UGT2B15 biomarker for diagnosing or predicting colorectal adenoma or / and a detection reagent thereof in preparing a product for diagnosing or predicting colorectal adenoma.
[0006] Furthermore, the products include reagents, kits, primers, chips, antibodies, and test strips.
[0007] As used herein, the term "UGT2B15" is intended to encompass fragments, variants (e.g., allelic variants), and derivatives thereof. Representative human UGT2B15 cDNAs and human UGT2B15 protein sequences are known in the art and are publicly available on the NCBI website. For example, at least one human UGT2B15 isoform is known: human UGT2B15, Gene ID: 7366. Nucleic acid and polypeptide sequences of human UGT2B15 orthologs in organisms other than humans are also known.
[0008] Furthermore, the sample types detected by the detection reagent include tissues, blood, plasma, serum, and blood cells.
[0009] Furthermore, the sample types include tissue, blood, and plasma.
[0010] Furthermore, the sample types are tissues and blood.
[0011] Furthermore, the sample type is tissue.
[0012] As used herein, the term "sample" refers to any composition containing nucleic acids or proteins isolated from a subject. In some embodiments, the sample can be selected from at least one of blood, tissue, blood cells, bone marrow, ascites, fine needle biopsy specimens, body fluids containing cells, free-floating nucleic acids, sputum, saliva, urine, semen, cerebrospinal fluid, peritoneal fluid, pleural fluid, stool, lymph, skin swabs, oral swabs, nasal swabs, or lavage materials. In some embodiments, the sample can be blood or tissue.
[0013] Furthermore, the detection reagent is used to detect the content of the biomarker.
[0014] Furthermore, the detection reagent includes a reagent for detecting the gene expression amount or protein expression level of the biomarker.
[0015] The present invention provides a kit for diagnosing or predicting colorectal adenoma in a subject, wherein the kit comprises a detection reagent for detecting the content of a biomarker, wherein the biomarker comprises UGT2B15.
[0016] The present invention provides an electronic device for diagnosing or predicting colorectal adenoma, the electronic device comprising:
[0017] an acquisition and detection module, configured to acquire a sample, perform detection on the sample, and obtain the content of a biomarker in the sample;
[0018] The diagnosis or prediction module uses a machine learning method and the content of the biomarker to establish a regression equation, construct a model, and then output a diagnosis or prediction result based on the model.
[0019] As used herein, the term "electronic device" refers to any suitable computing or processing device, or other device constructed or adapted to store data or information. Examples of electronic devices suitable for use with embodiments of the present invention include stand-alone computing devices; networks, including local area networks (LANs), wide area networks (WANs), the Internet, intranets, and extranets; electronic appliances such as personal digital assistants (PDAs), cell phones, web browsers, and the like; and local and distributed processing systems.
[0020] Furthermore, the machine learning method includes a linear regression algorithm model, a support vector machine algorithm model, a nearest neighbor / k-nearest neighbor algorithm model, a logistic regression algorithm model, a decision tree algorithm model, a k-means algorithm model, a random forest algorithm model, a naive Bayes algorithm model, a dimensionality reduction algorithm model, and a gradient enhancement algorithm model.
[0021] The term "machine learning" as used in the present invention refers to an algorithm that gives a computer the ability to learn without being explicitly programmed, including an algorithm that learns from data and makes predictions about the data. The machine learning algorithms used in the embodiments disclosed herein may include (but are not limited to) random forest (RF), least absolute shrinkage and selection operator (LASSO) logistic regression, regularized logistic regression, XGBoost, decision tree learning, artificial neural network (ANN), deep neural network (DNN), support vector machine, rule-based machine learning, etc. Algorithms such as linear regression or logistic regression can be used as part of the machine learning process.
[0022] Furthermore, the sample includes tissue, blood, plasma, serum, and blood cells.
[0023] Furthermore, the sample includes tissue, blood, and plasma.
[0024] Furthermore, the sample is tissue or blood.
[0025] Furthermore, the sample is tissue.
[0026] Furthermore, the biomarker includes UGT2B15.
[0027] In some embodiments, the method for detecting the expression level of a gene or protein in a sample is a sequencing method, which includes but is not limited to a second-generation sequencing method or a third-generation sequencing method. The means for performing sequencing is not particularly limited, and sequencing by a second-generation or third-generation sequencing method can achieve rapid and efficient sequencing. The high-throughput and deep sequencing characteristics of sequencing can be utilized, thereby facilitating the analysis of subsequent data, especially the accuracy of statistical analysis.
[0028] The present invention provides a method for screening biomarkers for diagnosing or predicting colorectal adenoma, the screening method comprising:
[0029] Obtaining colorectal adenoma prevalence and samples from a sample population, testing the samples to obtain gene or protein expression levels in the samples, wherein the sample population includes healthy individuals and colorectal adenoma patients;
[0030] According to the test results of the sample, screening genes or proteins with significant differences between the healthy individuals and the colorectal adenoma patients by statistical methods;
[0031] Using the levels of genes or proteins with significant differences screened out from the sample population and the prevalence of colorectal adenoma in the sample population, a prediction model is constructed through a machine learning training method;
[0032] The gene or protein that appears in the prediction model more than a threshold number of times is the biomarker.
[0033] Furthermore, the selection of the threshold in the screening method needs to be randomized, the number of randomizations is multiple, the number of prediction models constructed is multiple, and the prediction model needs to be cross-validated multiple times.
[0034] Furthermore, the statistical method includes any one or more of T-test, generalized linear model, rank test, logistic regression, multiple difference, multiple hypothesis testing correction or Kruskal-Wallis test.
[0035] The statistical methods used to determine the significant difference test results include, but are not limited to, T-test (T test), Wilcoxon test (rank test), logistic regression, KW test (Kruskal-Wallis test), FoldChange, Bonferroni correction, or one or more of other existing statistical methods. When constructing a prediction model, the colorectal adenoma prevalence and gene or protein expression levels in the significant difference test results of each sample are summarized to obtain the aggregated data of different individuals. Machine learning training is then performed, and cross-validation is used to construct a prediction model.
[0036] Furthermore, the sample includes tissue, blood, plasma, serum, and blood cells.
[0037] Furthermore, the sample includes tissue, blood, and plasma.
[0038] Furthermore, the sample is tissue or blood.
[0039] Furthermore, the sample is tissue.
[0040] Furthermore, the process of constructing the prediction model includes dividing the sample group into a training set and a test set, establishing a mapping relationship with the prevalence of colorectal adenoma, obtaining a training data set or a test data set respectively, and obtaining the prediction model through machine learning training methods and cross-validation.
[0041] In the above-mentioned biomarker screening method, samples including healthy individuals and colorectal adenoma patients and the prevalence of colorectal adenoma are first obtained. The gene or protein expression levels of the samples are detected, and a certain number of genes or proteins are selected as the targets to be detected. After the detection, the content of the genes or proteins to be detected in each sample is obtained. Statistical methods are used to test the genes or proteins that are different between colorectal adenoma patients and healthy individuals, which are significant difference detection results. The significant difference detection results indicate that the expression levels of the genes or proteins of these test results represent the number present in the sample population, which is related to the prevalence of colorectal adenoma. In the healthy group and colorectal adenoma patients, the expression levels of the genes or proteins of the significant difference detection results are significantly different, but it is impossible to confirm what effect the significant difference detection results have on diagnosis or prediction, and it is unclear how to perform diagnosis / prediction.
[0042] In some embodiments, in order to improve and ensure the effectiveness of diagnosing / predicting colorectal adenomas, machine learning methods are used to model the content of significant difference detection results of sample groups and the prevalence of colorectal adenomas, and construct a predictive model capable of diagnosing / predicting colorectal adenomas. The genes or proteins obtained from the predictive model are all genes or proteins that have been screened out more than a threshold number of times and have good effects, that is, the final biomarkers.
[0043] In some embodiments, in the process of establishing a prediction model, one or more analysis methods including but not limited to LASSO regression, BIC (Bayesian Information Criterion), and stepwise regression are used to analyze the established prediction model and select a prediction model with high effectiveness.
[0044] The present invention provides a method for constructing a diagnostic model for colorectal adenoma, wherein the method uses the aforementioned screening method to obtain biomarkers and applies a machine learning training method to construct a diagnostic model.
[0045] Furthermore, the biomarker includes UGT2B15.
[0046] The present invention provides a processor, which is used to run a program, and when the program is run, it executes the above-mentioned screening method, the above-mentioned construction method, or the above-mentioned diagnostic model.
[0047] The present invention provides a computer medium, which includes a storage medium, and the storage medium is used to store a program. When the program is run, the device connected to the storage medium is controlled to execute the screening method, the construction method, or the diagnostic model described above.
[0048] In some embodiments, the present invention can be used in a variety of general-purpose or special-purpose computing system environments, such as personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above, and the like.
[0049] In some embodiments, the above modules or steps can achieve the effect through a computing device. In some embodiments, the modules or steps can be concentrated on a single computing device. In some embodiments, the modules or steps can be distributed on multiple computing devices, and the multiple computing devices are linked by a network or other linking methods. In some embodiments, the computing device stores program code for executing the above modules or steps, and the program code is stored in a storage device. In some embodiments, the modules or steps can be prepared into integrated circuits separately, or combined into a single integrated circuit. In the specific embodiments of the present invention, no implementation method is limited, whether it is hardware or software.
[0050] The present invention provides the use of biomarkers in constructing a system for automated diagnosis or prediction of colorectal adenoma, wherein the biomarkers include UGT2B15.
[0051] The present invention provides a device for screening biomarkers for diagnosing or predicting colorectal adenoma, the device comprising:
[0052] an acquisition module configured to acquire colorectal adenoma disease conditions and samples from a sample population, test the samples, and obtain gene or protein expression levels of the samples, wherein the sample population includes healthy individuals and colorectal adenoma patients;
[0053] a statistical module configured to use the test results of the sample to screen genes or proteins that are significantly different between the healthy individuals and the colorectal adenoma patients through a statistical method;
[0054] A model building module is configured to build a prediction model by using the levels of genes or proteins with significant differences screened out from the sample population and the colorectal adenoma prevalence in the sample population through a machine learning training method;
[0055] The biomarker determination module is configured to define a gene or protein that appears in the prediction model more than a threshold value as the biomarker.
[0056] The term "module" used in the present invention refers to a combination of software and / or hardware that can implement a predetermined function. Implementation using hardware, software, or a combination of hardware and software is conceivable.
[0057] As used herein, the term "diagnosis" refers to determining a subject's health status, encompassing aspects such as detecting the presence or absence of a disease, response to treatment, recurrence risk assessment, assessment of the risk and extent of cancer, and prognosis. In some cases, the term "diagnosis" refers to a single factor used to determine, verify, or confirm a patient's clinical status.
[0058] In some specific embodiments, the device for screening biomarkers for diagnosing colorectal adenoma or predicting colorectal adenoma can screen for biomarkers for diagnosing colorectal adenoma or predicting colorectal adenoma by detecting gene or protein levels in samples from a sample population. In some embodiments, the device for screening biomarkers for diagnosing colorectal adenoma or predicting colorectal adenoma can perform new screening for biomarkers or partially update biomarkers when factors such as the subject's region and environment affect the predictive effect of the biomarkers.
[0059] Furthermore, the gene or protein expression levels detected in the sample include UGT2B15.
[0060] The term "ROC curve" used in the present invention is a receiver operating characteristic curve, which is a coordinate diagram composed of the false positive probability (1-specificity) as the horizontal axis and the true positive probability (sensitivity) as the vertical axis, and a curve drawn by the different results obtained by the subjects under specific stimulation conditions due to the use of different judgment criteria. Select the best diagnostic limit value. The closer the ROC curve is to the upper left corner, the higher the accuracy of the test. The point of the ROC curve closest to the upper left corner is the best threshold with the fewest errors, and its total number of false positives and false negatives is the least. Comparison of the ability of two or more different diagnostic tests to identify diseases. When comparing two or more diagnostic methods for the same disease, the ROC curves of each test can be plotted on the same coordinate to intuitively distinguish the pros and cons. The ROC curve represented by the ROC curve close to the upper left corner is the most accurate. It can also be compared by calculating the area under the ROC curve (AUC) of each test respectively. The test with the largest AUC has the best diagnostic value.
[0061] Advantages and beneficial effects of the present invention:
[0062] This study uses whole-transcriptome analysis to explore a novel biomarker, UGT2B15, associated with colorectal adenoma. The biomarker, UGT2B15, can screen patients for colorectal adenoma with high sensitivity and specificity. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 This is the HE staining image of the test set of colorectal adenoma tissue and adjacent normal tissue; CRA is colorectal adenoma tissue, and NOR is adjacent normal tissue;
[0064] Figure 2 This is the differential expression result diagram of UGT2B15 in colorectal adenoma tissue and adjacent normal tissue in the test set;
[0065] Figure 3 This is the ROC curve of UGT2B15 in the test set for colorectal adenoma diagnosis;
[0066] Figure 4 This is the HE staining image of the colorectal adenoma tissue and adjacent normal tissue. CRA is the colorectal adenoma tissue, and NOR is the adjacent normal tissue.
[0067] Figure 5 This is the immunohistochemical staining of UGT2B15; CRA represents adenoma tissue, and NOR represents normal adjacent tissue. Positive expression is indicated by yellow-brown cytoplasm and blue nuclei.
[0068] Figure 6 It is a bar graph of the average optical density values of UGT2B15 in the validation set colorectal adenoma tissue and adjacent normal tissue;
[0069] Figure 7 This is the ROC curve of UGT2B15 for the diagnosis of colorectal adenoma in the validation set. DETAILED DESCRIPTION
[0070] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present invention will be described in detail below with reference to the embodiments.
[0071] Example 1: Marker UGT2B15 is associated with the diagnosis of colorectal adenoma
[0072] 1. Clinical samples
[0073] The whole transcriptome sequencing (RNA sequencing, RNA-seq) test set used 12 cases of colorectal adenoma and 12 cases of adjacent normal tissue, for a total of 24 cases. Detailed information is shown in Table 1. All samples used were formaldehyde-fixed paraffin-embedded adenoma tissues that were surgically removed by colonoscopy at Wangjing Hospital between 2019 and 2023. The study samples were again confirmed as colorectal adenoma and adjacent normal tissue by two senior pathologists ( Figure 1 ).
[0074] Table 1 Clinical information of patients in the RNA-seq test set
[0075] serial number gender age Pathological type Part Grading level 1 female 75 Tubular adenoma hepatic flexure of the colon Mild to moderate dysplasia 2 male 58 Tubular adenoma sigmoid colon Mild dysplasia 3 female 75 Tubular adenoma hepatic flexure of the colon Mild to moderate dysplasia 4 female 73 Tubular adenoma transverse colon Mild dysplasia 5 male 83 Tubular adenoma sigmoid colon Mild dysplasia 6 male 63 Tubular adenoma B-level Moderate to severe dysplasia 7 female 52 Tubular adenoma sigmoid colon Moderate to severe dysplasia 8 female 50 Tubular adenoma transverse colon Mild dysplasia 9 male 54 Tubular adenoma sigmoid colon Mild to moderate dysplasia 10 male 60 Tubular adenoma sigmoid colon Mild to moderate dysplasia 11 male 57 Tubular adenoma rectum Moderate to severe dysplasia 12 male 59 Tubulovillous adenoma The intersection of Zhi and Yi Moderate to severe dysplasia
[0076] 2. Experimental equipment
[0077] Agilent 2100 Bioanalyzer (Agilent), Qubit 2.0 Fluorometer (Invitrogen), Veriti 96-well Multiplex Temperature Gradient Amplifier (MJ Research, USA), Peltier Thermal Cycler PTC-225 (MJ Research, USA), DNA Quantification Kit (KAPA Biosystem), Multi-Sample Adapter Primer Kit (NEBNext), High-Sensitivity DNA Detection Kit (Agilent), RNA 6000 Pico chip (Agilent).
[0078] 3. Experimental methods
[0079] Whole transcriptome sequencing was performed on 12 cases of colorectal adenoma tissues and 12 cases of adjacent normal tissues. The sequencing process mainly included RNA extraction, library construction, sequencing, data quality control, and data analysis.
[0080] 3.1 Library construction and sequencing
[0081] Total RNA from colorectal adenomas and adjacent normal tissues was extracted using Trizol. RNA integrity was assessed using an Agilent 2100 Bioanalyzer, with quality requirements of RIN ≥7 and 28S / 18S ratio ≥1.5:1. The QUBIT RNAASSAY KIT was used to accurately quantify the starting total RNA, with a minimum input of 0.1–1 μg. mRNA from the total RNA was purified using beads containing Oligod(T). After fragmentation to the target length using NEBNext Reaction Buffer, the first complementary DNA (cDNA) strand was synthesized using random primers. Second-strand cDNA was then synthesized using DNA polymerase I. The purified double-stranded cDNA was end-repaired, A-tailed, and ligated with sequencing adapters before PCR amplification. PCR products were quality-controlled using a 2100 Bioanalyzer chip. The final library was then sequenced on an Illumina Novaseq 6000 sequencing platform.
[0082] 3.2 Data filtering and quality control
[0083] Base calling was performed on the raw files of IILumina sequencing and converted into raw sequenced reads. The collection of raw sequenced reads was called Raw Data. Raw Data was filtered to remove reads that did not meet the analysis criteria to obtain Clean Reads for subsequent analysis. The filtering conditions were as follows: (1) Remove contaminated samples; (2) If the reads contained adapter sequences, the adapter sequences were truncated first; (3) Using a sliding window of 4 bases, bases with an average base quality of less than 15 at both ends of the read were truncated; (4) Read pairs with high N content were removed (if the N content in a read was greater than 5%, the entire pair of reads was removed); (5) Read pairs with high content of low-quality bases (bases with a quality value ≤ 20) were removed (if the low-quality bases in a read were greater than 30%, the entire pair of reads was removed); (6) If the length of the truncated reads was less than 100 bp, the entire pair of reads was removed; (7) Unpaired reads were removed.
[0084] 3.3 Sequence alignment and gene quantification
[0085] The reference genome index file was downloaded from the official website. Clean reads were aligned to the reference genome sequence using Hisat2 software. The number of clean reads aligned to the reference genome was required to be ≥70% of the total number of clean reads, and the number of reads aligned to multiple positions in the reference genome was ≤10% of the total number of reads aligned to the reference genome. Gene expression read counts were calculated using featureCounts and StringTie software. The number of fragments per kilobase of exon model per million mapped fragments (FPKM) was calculated taking into account the length and depth of the RNA sequence. The FPKM values of gene expression were filtered: (1) genes with a length of less than 100 bp were removed; (2) genes with an FPKM < 0.0001 were considered to have an expression of 0, and genes with an expression of 0 in all samples were removed.
[0086] 3.4 Analysis of differentially expressed genes
[0087] Statistical analysis was performed on the quantified gene expression data. The R language limma package and DESeq package were used to screen out differentially expressed genes (DEGs) between colorectal tissues and adjacent normal tissues. The screening criteria were: |log2FC|>1 and P<0.05, where FC represents the fold change (FC) and the P value represents the statistical significance of the difference (P-Value). At the same time, the FPKM value of gene expression was used to further screen out genes that were expressed in colorectal adenoma tissues but almost not expressed in normal tissues.
[0088] 3.5 Statistical analysis
[0089] Statistical analysis was performed using IBM SPSS Statistics 26.0 software. Measurement data that conformed to a normal distribution were expressed as mean ± standard deviation (x ± s). Comparisons between the two groups were performed with the independent sample t-test; those that did not conform to a normal distribution were compared with nonparametric tests. P < 0.05 indicated statistical significance.
[0090] 4. Experimental results
[0091] According to the results of whole transcriptome sequencing, the expression level of UGT2B15 in colorectal adenoma tissue was significantly higher than that in normal tissue, with extremely significant statistical differences (P < 0.001). See Table 2 and Table 3. In the evaluation of its diagnostic efficacy, the ROC curve showed that the area under the curve (AUC) was 0.9514, at which point its sensitivity and specificity could reach 91.67%. Figure 2 、 Figure 3 .
[0092] Table 2 Differential expression of UGT2B15 in adenoma tissue and adjacent tissue
[0093]
[0094]
[0095] Note: CRA represents colorectal adenoma tissue, and NOR represents normal adjacent tumor tissue.
[0096] Table 3 Quantitative expression values (FPKM) of UGT2B15 in each sample of the test set
[0097] Sample No. FPKM value of adenoma tissue FPKM value of peritumoral tissue 1 3.28283 0 2 1.27065 0 3 2.99115 0 4 0.85987 0 5 1.23642 0 6 0.24646 0 7 2.41035 0 8 1.72367 0.7081 9 2.75271 0.40397 10 1.92791 0 11 1.81842 1.40223 12 1.21561 0
[0098] Example 2 Immunohistochemical Validation of UGT2B15 for Diagnostic Performance of Colorectal Adenomas
[0099] 1. Clinical samples
[0100] The immunohistochemistry (IHC) validation set used 45 cases of colorectal adenomas and 45 cases of adjacent normal tissues, for a total of 90 cases (see Table 4). All samples used were formaldehyde-fixed paraffin-embedded adenoma tissues that were surgically removed by colonoscopy at Wangjing Hospital between 2019 and 2023. The study samples were again confirmed as colorectal adenomas and adjacent normal tissues by two senior pathologists (see Figure 4 ).
[0101] Table 4 Clinical information of patients in the IHC validation set
[0102]
[0103]
[0104] 2. Experimental equipment
[0105] 4℃ / -20℃ refrigerator (BCD-196F) (Qingdao Haier Co., Ltd.), RM2235 microtome (Leica, Germany), HI1210 slide spreader (Leica, Germany), HI1210 slide baking machine (Leica, Germany), high pressure steam sterilizer (YX280) (Hefei Huatai Medical Co., Ltd.), electric blower drying oven (101-0AB) (Linmao Technology Beijing Co., Ltd.), three-purpose constant temperature water bath (SH.W21.600) (Shanghai Shuli Instrument Co., Ltd.), BX51 microscope (OLYMPUS, Japan) Co., Ltd.), hematoxylin blue solution (Beijing Shiji Heli Biotechnology Co., Ltd.), DAB color development kit (Beijing Zhongshan Biotechnology Co., Ltd.), hydrogen peroxide solution (Beijing Haiderun Pharmaceutical Group Co., Ltd.), UGT2B15 antibody reagent (A16657) (Wuhan Aibotek Biotechnology Co., Ltd.), citric acid repair solution (pH 6.0) (Beijing Zhongshan Jinqiao Biotechnology Co., Ltd.), EDTA antigen repair solution (pH 8.0) (Beijing Zhongshan Jinqiao Biotechnology Co., Ltd.), PBS buffer powder (pH 7.3) (Beijing Zhongshan Jinqiao Biotechnology Co., Ltd.).
[0106] 3. Experimental methods
[0107] The UGT2B15 gene identified through whole-transcriptome sequencing was validated by immunohistochemistry. Image ProPlus 6.0 software was used to calculate the mean optical density of UGT2B15, perform semi-quantitative analysis, and plot a receiver operating characteristic (ROC) curve to assess diagnostic value. The immunohistochemical validation process is as follows:
[0108] (1) Tissue sections:
[0109] 1) CRA paraffin samples were sectioned at 2.5 μm thickness;
[0110] 2) Spread in a 48°C spreader;
[0111] 3) Bake the slices in a 63°C oven for 1 hour.
[0112] (2) Dewaxing and hydration:
[0113] 1) Soak the prepared sections in dewaxing solutions I and II for 15 min respectively;
[0114] 2) Soak in anhydrous ethanol I and II for 5 minutes respectively;
[0115] 3) Soak in 95% ethanol for 2 minutes and then in 80% ethanol for 1 minute;
[0116] 4) Rinse with distilled water three times and then soak in PBS for 5 minutes three times.
[0117] (3) Antigen retrieval:
[0118] 1) Heat EDTA or citric acid repair solution in a pressure cooker until boiling and then add the slices;
[0119] 2) After the pressure cooker is vented, start antigen retrieval;
[0120] 3) EDTA timing 3 minutes, citric acid timing 2.5 minutes;
[0121] 4) Cool the slices at room temperature for 20 minutes.
[0122] (4) Immune response:
[0123] 1) Soak in PBS for 5 min × 3 times, then soak in 3% H2O2 for 10 min;
[0124] 2) Soak again in PBS for 5 min × 3 times;
[0125] 3) Add primary antibody and incubate at room temperature at 37°C for 2 h;
[0126] 4) Soak in PBS for 5 min × 3 times;
[0127] 5) Add secondary antibody and incubate at room temperature at 37°C for 30 minutes;
[0128] 6) Soak in PBS for 5 min x 3 times.
[0129] (5)Chemical dyeing:
[0130] 1) DAB staining, microscopic examination of the staining degree, positive expression appears yellow-brown;
[0131] 2) Restain with hematoxylin for 5 minutes, separate with color-separating solution, and reverse blue with anti-blue solution, then wash with water three times.
[0132] (6) Dehydration and sealing:
[0133] 1) Soak in 80% ethanol for 1 min, 95% ethanol for 2 min, and 90% ethanol II for 2 min.
[0134] 2) Soak in anhydrous ethanol I for 5 min and anhydrous ethanol II for 5 min, respectively;
[0135] 3) Soak in clear solution I for 1 min, clear solution II for 5 min, and clear solution III for 5 min in sequence;
[0136] 4) Fix the slides with neutral gum.
[0137] 4. Statistical analysis
[0138] The statistical analysis method is the same as in Example 1.
[0139] 5. Experimental results
[0140] The results of IHC staining of UGT2B15 were analyzed and it was found that the average optical density of UGT2B15 in colorectal adenoma tissue was significantly higher than that in adjacent normal tissue, and the difference was statistically significant (P < 0.001). In the evaluation of its diagnostic efficacy, the ROC curve showed that the AUC was 0.918, at which point its sensitivity could reach 86.67% and its specificity could reach 91.11%. Figure 5 、 Figure 6 、 Figure 7 .
[0141] The above embodiments are only provided for understanding the method and core concept of the present invention. It should be noted that, without departing from the principles of the present invention, a number of improvements and modifications may be made to the present invention by a person skilled in the art, and such improvements and modifications shall fall within the scope of protection of the claims of the present invention.
Claims
1. Use of a detection reagent for detecting the content of the UGT2B15 biomarker in the preparation of a product for diagnosing colorectal adenoma, wherein the sample type detected by the detection reagent is tissue.
2. An electronic device for diagnosing colorectal adenoma, the electronic device comprising: an acquisition and detection module, configured to acquire a sample, perform detection on the sample, and obtain the content of the UGT2B15 biomarker in the sample, wherein the sample is a tissue; The diagnostic module uses a machine learning method and the content of the biomarker to establish a regression equation, construct a model, and then output a diagnostic result based on the model.
3. A method for constructing a diagnostic model for colorectal adenoma, wherein the method uses the content of the UGT2B15 biomarker in a tissue sample, applies a machine learning training method and the content of the biomarker to establish a regression equation, and constructs a diagnostic model.
4. A processor, configured to run a program, wherein the program, when run, executes the construction method according to claim 3 or the diagnostic model according to claim 3.
5. A computer medium, comprising a storage medium, wherein the storage medium is used to store a program, and when the program is run, the device connected to the storage medium is controlled to execute the construction method according to claim 3 or the diagnostic model according to claim 3.
Citation Information
Patent Citations
Application of protein marker in preparation of product for diagnosing Parkinson's disease or predicting Parkinson's disease
CN115856309A
In vitro method for predicting the response of patients suffering from her2+ breast cancer to a treatment with Anti-her2 neoadjuvant therapy
WO2023099350A1