Application and detection methods of protein biomarkers in the preparation of detection products that differentiate between early-stage NASH and Non-NASH
By establishing a detection model using 21 protein biomarkers, the challenge of non-invasive diagnosis of early NASH has been solved, achieving high sensitivity and high specificity for early NASH detection, which is suitable for rapid screening of serum samples.
Patent Information
- Application Number
- CN202111398328.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-23
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2041-11-23
AI Technical Summary
Existing technologies are insufficient for the effective detection of early-stage NASH, and non-invasive diagnostic methods lack high sensitivity and specificity, making it impossible to accurately distinguish between early-stage NASH and Non-NASH.
Using 21 protein biomarkers, including ALCAM, PON3, ROBO1, PTS, CDCP1, TRAIL-R2, and FBP1, a detection model was established through logistic regression. Detection kits or devices were developed to detect protein expression levels in serum samples to distinguish between early NASH and Non-NASH.
It achieves high sensitivity and high specificity for early NASH detection, enabling rapid and objective screening of early NASH, reducing testing costs and time, and improving diagnostic accuracy.
Smart Images

Figure CN115963272B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of non-alcoholic fatty liver disease (NASH) detection, and more specifically, to the application and detection method of a protein biomarker in the preparation of a detection product that distinguishes between early NASH and Non-NASH. Background Technology
[0002] Nonalcoholic steatohepatitis (NASH) is a type of nonalcoholic fatty liver disease (NAFLD), a chronic liver disease caused by fat accumulation and inflammation. Its symptoms are similar to those of alcoholic fatty liver disease, but patients do not have a history of excessive alcohol consumption. The NAFLD spectrum typically includes NAFL (nonalchoholic simple fatty liver), NASH, cirrhosis, and liver cancer. In recent years, with the global spread of obesity, type 2 diabetes, and cardiovascular disease, the number of NAFLD patients in clinical practice has continued to increase. NAFLD has become the most important cause of liver disease worldwide, affecting more than a quarter of the global population.
[0003] The development of NAFLD is a dynamic process. Its main characteristic is the excessive accumulation of fat in the liver. This accumulated fat, combined with other factors such as insulin resistance, places significant metabolic and oxidative stress on the liver, leading to persistent inflammation and hepatocyte apoptosis. If NASH is not controlled at this stage, the liver will initiate its own repair mechanisms to replenish the dead hepatocytes after damage occurs. Fibrosis is a physiological response accompanying this repair process. Chronic inflammation and excessive liver metabolic toxicity continue to cause hepatocyte death, and the cumulative liver fibrosis eventually leads to cirrhosis. Compared to NAFL, NASH patients have a worse prognosis. Studies have shown that 11% of NASH patients are diagnosed with cirrhosis, while only 1% of NAFL patients are diagnosed with cirrhosis. NAFL and early-stage NASH are usually reversible, but NASH and fibrosis develop insidiously, often without obvious symptoms in the early stages. By the time obvious symptoms appear, the disease may have progressed to intermediate or advanced cirrhosis and hepatocellular carcinoma, increasing the risk of death from liver disease.
[0004] Hepatocellular carcinoma accounts for approximately 75% of all liver cancers, making it the most common primary liver cancer. In 2020, liver cancer ranked fifth in mortality worldwide, with its incidence and prevalence increasing annually. In China, liver cancer is the second leading cause of cancer death. Although hepatitis B and hepatitis C are the main causes of liver cancer, with widespread hepatitis B vaccination, NASH (Neuro-Neuro-Hypertension) is predicted to gradually replace hepatitis viruses as the leading cause of liver cancer. Therefore, early screening of high-risk groups for NASH and accurate subtyping of NAFLD are crucial for early intervention and treatment, effectively improving cure rates and reducing the likelihood of disease progression.
[0005] Currently, liver biopsy is the gold standard for differentiating patients from non-alcoholic liver disease (NAFLD), NASH, liver fibrosis, and cirrhosis. Diagnosis is primarily based on pathologists' scores for hepatocellular steatosis, intralobular inflammation, hepatocellular ballooning, and fibrosis. However, liver biopsy is an invasive procedure, expensive, and carries potential issues such as sampling errors and subjective interpretation by pathologists, making it unsuitable as the first-line method for NASH screening and treatment evaluation. Some studies have explored using artificial intelligence and machine learning to assist physicians in evaluating medical imaging results for more accurate diagnosis and reduced time and financial costs; however, no such technology has yet been applied to NAFLD diagnosis, and this technology still cannot resolve the issue of sample sampling errors.
[0006] In addition, several non-invasive diagnostic methods for NAFLD have been proposed, including serological marker testing, imaging examinations, and predictive models. For simple fatty liver, reports indicate that ultrasound, MRI, and TE are effective non-invasive methods for quantifying liver fat content, but their clinical application is not yet widespread due to factors such as cost.
[0007] Keratin (CK-18) is currently an important serological marker for NASH, but it is still under scientific research and has not yet been applied clinically. The development of non-invasive diagnostics for fibrosis is more promising. For example, APRI (aspartate aminotransferase (AST) to platelet ratio index), NAFLD fibrosis score, and FIB-4 (fibrosis-4 index) are mainly based on formulas using routine clinical variables to predict the progression of fibrosis. Some diagnostic methods are designed by detecting new biomarkers. For example, the ELF (enhanced liver fibrosis) test measures the degree of liver fibrosis by measuring the concentration of three matrix switching proteins (hyaluronic acid, tissue inhibitor of metalloproteinase 1, and N-terminal procollagen III-peptide). ELF testing is still in clinical trials and its diagnostic effectiveness is poor in early NASH and patients with low fibrosis. VCTE (vibration-controlled transient elastography; FibroScan) can non-invasively measure liver stiffness and has been approved by the FDA for use in children and adults. NIS4 (non-invasive score 4) diagnoses NASH and liver fibrosis by measuring the levels of four molecules (miR-34a-5p, α-2-macroglobuline, chitinase-like protein 1, and HbA1c) in serum. This test has been shown in clinical trials to have high accuracy in distinguishing between non-NASH and advanced NASH (which is distinguished from early NASH by a higher degree of fibrosis, i.e., a higher clinical fibrosis score). However, this method has not been validated in Asian populations.
[0008] In summary, there is currently no highly sensitive and specific protocol for diagnosing early-stage NASH patients without significant fibrosis. Summary of the Invention
[0009] The main objective of this invention is to provide an application and detection method for protein biomarkers in the preparation of detection products that can distinguish between early NASH and Non-NASH, so as to solve the problem that it is difficult to effectively detect early NASH in the prior art.
[0010] To achieve the above objectives, according to a first aspect of the present invention, there is provided an application of a protein biomarker in the preparation of a detection product for distinguishing between early NASH and NON-NASH, the protein biomarker including any one or more of the following: ALCAM, PON3, ROBO1, PTS, CDCP1, TRAIL-R2, FBP1, Insulin, IL-1ra, SELE, CTSD, KYNU, IGFBP-7, SULT2A1, MVK, GUSB, RBKS, ALDH1A1, FGF-21, HAOX1, and DAG1.
[0011] Further, the protein biomarkers include any one or more of the following: ALCAM, PON3, ROBO1, PTS, CDCP1, TRAIL-R2, and FBP1; preferably, the protein biomarkers further include any one or more of the following: Insulin, IL-1ra, SELE, CTSD, KYNU, IGFBP-7, SULT2A1, MVK, GUSB, RBKS, ALDH1A1, FGF-21, HAOX1, and DAG1.
[0012] Furthermore, the protein biomarkers include any two or more of the following combinations: ALCAM, PON3, ROBO1, PTS, CDCP1, TRAIL-R2, FBP1, Insulin, IL-1ra, and SELE.
[0013] Furthermore, the protein biomarkers are selected from any one of the following groups:
[0014]
[0015]
[0016]
[0017] Furthermore, the testing product is a test kit or a test device.
[0018] To achieve the above objectives, according to a second aspect of the present invention, a kit for distinguishing between early NASH and NON-NASH is provided. The kit includes detection reagents for protein biomarkers, which include any one or more of the following: ALCAM, PON3, ROBO1, PTS, CDCP1, TRAIL-R2, FBP1, Insulin, IL-1ra, SELE, CTSD, KYNU, IGFBP-7, SULT2A1, MVK, GUSB, RBKS, ALDH1A1, FGF-21, HAOX1, and DAG1.
[0019] Further, the protein biomarkers include any one or more of the following: ALCAM, PON3, ROBO1, PTS, CDCP1, TRAIL-R2, and FBP1; preferably, the protein biomarkers further include any one or more of the following: Insulin, IL-1ra, SELE, CTSD, KYNU, IGFBP-7, SULT2A1, MVK, GUSB, RBKS, ALDH1A1, FGF-21, HAOX1, and DAG1.
[0020] Furthermore, the protein biomarkers include any two or more of the following combinations: ALCAM, PON3, ROBO1, PTS, CDCP1, TRAIL-R2, FBP1, Insulin, IL-1ra, and SELE.
[0021] Furthermore, the protein biomarkers are selected from any group of protein biomarkers used in the above applications.
[0022] To achieve the above objectives, according to a third aspect of the present invention, a detection device for distinguishing between early NASH and non-NASH is provided. The detection device has a built-in detection model for distinguishing between early NASH and non-NASH, wherein the detection model is a model for detecting protein biomarkers, and the protein biomarkers include multiple protein biomarkers used in the above applications.
[0023] Furthermore, the detection model is a logistic regression model.
[0024] Furthermore, the detection device includes a storage medium on which the detection model is stored.
[0025] Furthermore, the detection device includes a processor for running the detection model.
[0026] Furthermore, the detection device includes a protein biomarker expression level receiving module for the sample to be tested, and the receiving module includes at least one of the following modes: user manual input mode, alternative list import mode, and file import mode.
[0027] Furthermore, the sample to be tested is a body fluid sample, preferably a serum sample.
[0028] Furthermore, the test samples are derived from any one or more of the following subjects: healthy individuals or NAFLD patients.
[0029] To achieve the above objectives, according to a fourth aspect of the present invention, a detection method for distinguishing between early NASH and non-NASH is provided. The detection method includes: detecting the expression level of a protein biomarker in the body fluid of a subject to obtain the expression level of the protein to be tested; inputting the expression level of the protein to be tested into a detection model for distinguishing between early NASH and non-NASH, and outputting the detection result; wherein the detection model is a model for detecting protein biomarkers, and the protein biomarkers include the protein biomarkers used in the above applications.
[0030] Furthermore, the detection model is a logistic regression model.
[0031] Furthermore, the body fluid is serum.
[0032] Furthermore, the subjects were selected from any one or more of the following: healthy individuals or NAFLD patients.
[0033] By applying the technical solution of this invention and screening different population cohorts, 21 serum protein biomarkers significantly correlated with early NASH were identified. Therefore, using these proteins as biomarkers for effective detection of early NASH has significant clinical application value. For example, by utilizing the expression level (e.g., concentration) of any one or more of these protein biomarkers in serum, combined with clinical data, related detection products can be developed, such as preparing detection kits or establishing relevant diagnostic models, to detect or diagnose whether a subject has early NASH. Attached Figure Description
[0034] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:
[0035] Figure 1a and Figure 1b The differential expression of the protein biomarker according to the present invention 21 in early NASH and Non-NASH populations is shown. Detailed Implementation
[0036] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The present invention will now be described in detail with reference to the embodiments.
[0037] Terminology Explanation:
[0038] NASH: Non-alcoholic steatohepatitis. Clinically, NASH is divided into three stages. This application targets stage 1 NASH, also known as early NASH. The difference between stage 1 and stages 2 and 3 lies in whether fibrosis is present in the liver biopsy results. That is, if only hepatocellular steatosis, intralobular inflammation, and hepatocellular ballooning are found, and the fibrosis score is 0, it is considered early NASH. If the fibrosis score is greater than 0, it is considered advanced NASH.
[0039] Non-NASH: This refers to individuals who do not have non-early NASH. In this application, it includes healthy individuals without liver disease and individuals with simple fatty liver.
[0040] Advanced NASH: In this application, NAFLD specifically refers to a general term for liver diseases including NAFL and NASH. As mentioned above, NASH can be divided into three stages. This application mainly studies the difference between stage 1 NASH and non-NASH. Advanced NASH generally refers to stages 2 and 3 NASH. The main difference between stage 1 and stages 2 and 3 NASH lies in the presence or absence of fibrosis. The samples used in this application have a very low overall fibrosis score (0), so this study focuses on early-stage NASH. A fibrosis score greater than 0 indicates advanced NASH (i.e., mid-to-late stage NASH).
[0041] NAFL: Simple fatty liver, a type of non-alcoholic fatty liver disease.
[0042] NAFLD: A general term for non-alcoholic liver diseases, but in this application it specifically refers to the general term for liver diseases including NASH and NAFL.
[0043] PCR: Polymerase chain reaction
[0044] Uniport (Universal Protein) is a protein database that contains protein sequences, functional information, and research paper indexes. It integrates resources from three major databases: EBI (European Bioinformatics Institute), SIB (the Swiss Institute of Bioinformatics), and PIR (Protein Information Resource).
[0045] Based on the need for more effective non-invasive detection of early NASH mentioned in the background section, this application used proximity extension analysis to detect serum proteins in a cohort of 179 individuals. Combined with comprehensive clinical information and assembled section pathological data, 21 proteins were found to be significantly associated with early NASH. Therefore, based on the expression levels of any one or more of these 21 proteins in serum, detection models were established to distinguish between individuals with early NASH and those with non-NASH. Clinical data validation showed that these different detection models have high accuracy and are therefore suitable for rapid, efficient, and objective screening of relevant populations.
[0046] Based on the above research results, the applicant has proposed the technical solution of this application. In a first typical embodiment of this application, an application of a protein biomarker in the preparation of a detection product for distinguishing between early NASH and NON-NASH is provided, wherein the protein biomarker includes any one or more of the following: ALCAM, PON3, ROBO1, PTS, CDCP1, TRAIL-R2, FBP1, Insulin, IL-1ra, SELE, CTSD, KYNU, IGFBP-7, SULT2A1, MVK, GUSB, RBKS, ALDH1A1, FGF-21, HAOX1, and DAG1.
[0047] In some preferred embodiments, the protein markers include any one or more of the following: ALCAM, PON3, ROBO1, PTS, CDCP1, TRAIL-R2, and FBP1 (these protein markers are newly discovered markers that can be used to distinguish between early NASH and NON-NASH). In other preferred embodiments, in addition to these protein markers, any one or more of the following are further included: Insulin, IL-1ra, SELE, CTSD, KYNU, IGFBP-7, SULT2A1, MVK, GUSB, RBKS, ALDH1A1, FGF-21, HAOX1, and DAG1.
[0048] As mentioned above, these protein biomarkers are significantly correlated with NASH, so any one of them can be used to detect early NASH. Similarly, any two or more random combinations can also be used to detect NASH, and the sensitivity and specificity of using these protein biomarkers to distinguish between NASH and Non-NASH are both high.
[0049] In some preferred embodiments, when using two or more protein biomarkers in combination, the biomarkers are selected from any two or more of the following: ALCAM, PON3, ROBO1, PTS, CDCP1, TRAIL-R2, FBP1, Insulin, IL-1ra, and SELE. This combination of protein biomarkers provides higher sensitivity and specificity when used to differentiate between early NASH and non-NASH.
[0050] In some preferred embodiments, the protein biomarkers are selected from any group of protein biomarkers in Table 1 below.
[0051] Table 1:
[0052]
[0053]
[0054]
[0055]
[0056] It should be noted that, among the various combinations of protein markers in the table above, the following combination is further preferred:
[0057] 1) ALCAM and PON3;
[0058] 2) Insulin, ROBO1, and CDCP1;
[0059] 3) TRAIL-R2, ALCAM, PON3 and CDCP1;
[0060] 4) Insulin, IL-1ra, TRAIL-R2, ALCAM and PON3;
[0061] 5) Insulin, TRAIL-R2, ALCAM, SELE, PON3, CDCP1 and PTS;
[0062] 6) Insulin, TRAIL-R2, ALCAM, SELE, PON3, ROBO1, FBP1 and PTS.
[0063] The protein biomarkers described above show relatively accurate diagnostic results predicted by different clinical classification models. It should be noted that the aforementioned testing products can be any clinically applicable product, such as test kits, testing devices, peptide chips, etc. Specific application methods include, but are not limited to, clinical mass spectrometry, chemiluminescence, or single-molecule immunoassay using Simoa.
[0064] In a second typical embodiment, a kit is provided to distinguish between early NASH and non-NASH. The kit includes detection reagents for protein biomarkers, which include any one or more of the following: ALCAM, PON3, ROBO1, PTS, CDCP1, TRAIL-R2, FBP1, Insulin, IL-1ra, SELE, CTSD, KYNU, IGFBP-7, SULT2A1, MVK, GUSB, RBKS, ALDH1A1, FGF-21, HAOX1, and DAG1.
[0065] Since the above-mentioned protein biomarkers are significantly correlated with NASH, the detection kits designed for these protein biomarkers not only make detection convenient, simple, and rapid, but also have high detection sensitivity and specificity.
[0066] In some preferred embodiments, the protein biomarkers include any one or more of the following: ALCAM, PON3, ROBO1, PTS, CDCP1, TRAIL-R2, and FBP1 (these protein biomarkers are newly discovered biomarkers capable of distinguishing early NASH from non-NASH). In other preferred embodiments, in addition to these protein biomarkers, any one or more of the following are further included: Insulin, IL-1ra, SELE, CTSD, KYNU, IGFBP-7, SULT2A1, MVK, GUSB, RBKS, ALDH1A1, FGF-21, HAOX1, and DAG1. Kits designed specifically to include these protein biomarkers, based on ALCAM, PON3, ROBO1, PTS, CDCP1, TRAIL-R2, and FBP1, exhibit relatively higher detection sensitivity and specificity.
[0067] In some preferred embodiments, when using two or more protein biomarkers in combination, the biomarkers are selected from any two or more of the following: ALCAM, PON3, ROBO1, PTS, CDCP1, TRAIL-R2, FBP1, Insulin, IL-1ra, and SELE. This combination of protein biomarkers provides higher sensitivity and specificity when used to differentiate between early NASH and non-NASH.
[0068] In some preferred embodiments, the protein biomarkers detected by the assay reagents in the kit are selected from any group in Table 1. It should be noted that, among the various combinations of protein biomarkers in the table above, the following combination is further preferred:
[0069] 1) ALCAM and PON3;
[0070] 2) Insulin, ROBO1, and CDCP1;
[0071] 3) TRAIL-R2, ALCAM, PON3 and CDCP1;
[0072] 4) Insulin, IL-1ra, TRAIL-R2, ALCAM and PON3;
[0073] 5) Insulin, TRAIL-R2, ALCAM, SELE, PON3, CDCP1 and PTS;
[0074] 6) Insulin, TRAIL-R2, ALCAM, SELE, PON3, ROBO1, FBP1 and PTS.
[0075] The protein biomarkers in the above combination are all relatively accurate in predicting diagnostic results according to different clinical classification models.
[0076] For the purpose of setting up detection kits, various different types of detection kits can be prepared according to specific needs. The specific form of the kit is not limited; for example, it can be an ELISA kit, an immunofluorescence kit, or an immunogold immunoassay kit. There are also no restrictions on the detection method of the kit. All clinical methods for detecting proteins are applicable to this application. For example, it can be detected using peptide chips, mass spectrometry, or single-molecule immunoassay using SimoA, etc.
[0077] In a preferred embodiment, the reagent for detecting each protein biomarker in the above kit is an antibody for the corresponding protein biomarker. In some preferred embodiments, these antibodies are disposed on a solid-phase support; preferably, the solid-phase support is selected from ELISA plates, membrane carriers, or microspheres, and more preferably, the membrane carrier is selected from nitrocellulose membranes, glass cellulose membranes, or nylon membranes; preferably, the antibodies for the above protein biomarkers are monoclonal antibodies or polyclonal antibodies.
[0078] From the perspective of convenient detection and easy interpretation of test results, the antibodies for each protein biomarker in the kit are preferably pre-coated. Preferably, the pre-coated antibodies are coated on a solid-phase carrier; the specific solid-phase carrier is designed reasonably according to needs. More preferably, the solid-phase carrier includes an ELISA plate (mostly made of polystyrene), a membrane carrier, or microspheres; even more preferably, the membrane carrier includes a nitrocellulose membrane (the most widely used), a glass cellulose membrane, or a nylon membrane; even more preferably, the membrane carrier is also coated with a positive control, and the corresponding protein biomarkers and positive controls are sequentially arranged on the nitrocellulose membrane according to the detection order.
[0079] Depending on the specific detection method of the kit, the specific reagents in the kit will also vary accordingly, but they can all be combined according to the known preparation method of the kit. Preferably, the above kit also includes at least one of the following: (1) enzyme-labeled secondary antibody, more preferably HRP-labeled secondary antibody (corresponding to ELISA detection kit); (2) colloidal gold conjugate pad, the colloidal gold conjugate pad is coated with a specific conjugate of colloidal gold labeled antibody and positive control (corresponding to immunogold detection kit); (3) label pad, the label pad is coated with fluorescently labeled microspheres, the microspheres are loaded with a specific conjugate of positive control (corresponding to immunofluorescence detection kit).
[0080] The aforementioned immunogold assay kit and immunofluorescence assay kit offer relatively convenient detection capabilities, requiring only the establishment of a C-line for the positive control and a T-line for the test sample. The positive control pre-coated at the C-line can be any specific conjugate carrying a detection marker that is carried along with the test sample during the chromatographic process; there are no specific limitations on the antibody used for the positive control. Preferably, the positive control is selected from mouse immunoglobulin, human immunoglobulin, goat immunoglobulin, or rabbit immunoglobulin, and correspondingly, the specific conjugate of the positive control is selected from anti-mouse immunoglobulin, anti-human immunoglobulin, anti-goat immunoglobulin, or anti-rabbit immunoglobulin.
[0081] The aforementioned anti-mouse immunoglobulins, depending on the target animal, can be sheep anti-mouse immunoglobulins, rabbit anti-mouse immunoglobulins, or anti-mouse immunoglobulins from other immunizable animals. Similarly, anti-human, anti-sheep, or anti-rabbit immunoglobulins can also be derived from different species depending on the immunized animal. These immunoglobulins can be any one of IgM, IgG, IgA, IgD, or IgE. These anti-immunoglobulin antibodies can be monoclonal or polyclonal antibodies.
[0082] The specifications of the microplates used in the above kits vary depending on the number of samples to be tested, and can be reasonably selected from 12 to 384-well microplates.
[0083] All of the above-mentioned kits can quantify proteins, specifically various protein markers in serum. Taking the ELISA kit as an example, serum protein markers, such as TRAIL-R2, react with TRAIL-R2 antibodies on the surface of a solid-phase carrier. Then, enzyme-labeled antibodies are added, which also bind to the solid-phase carrier through a reaction. At this point, the amount of enzyme on the solid phase is proportional to the amount of CDCP1 protein in the serum. After adding the substrate for the enzyme reaction, the substrate is catalyzed by the enzyme into a colored product. The amount of the colored product is directly related to the amount of TRAIL-R2 protein in the serum; therefore, qualitative or quantitative analysis can be performed based on the intensity of the color. Because the enzyme's catalytic efficiency is very high, it indirectly amplifies the results of the immunoreaction, resulting in a very high sensitivity of the assay method.
[0084] In addition, the test kit can also be in the form of a test chip, such as having antibodies for multiple or all protein markers simultaneously set on the chip to achieve more efficient detection.
[0085] In a third typical embodiment of this application, a detection device is provided to distinguish between early-stage NASH and non-NASH. This device incorporates a detection model for differentiating between early-stage NASH and non-NASH. The detection model is a model for detecting protein biomarkers, which include multiple protein biomarkers detected in any of the aforementioned kits. Based on the research results of this application, using the 21 discovered protein biomarkers, different detection models can be established based on any number of them according to different clinical classification criteria. All of these detection models can effectively distinguish between NASH and non-NASH.
[0086] The aforementioned detection model can employ various known classifier models, such as the random forest model and the SVC model, both of which can effectively classify the two types of patients. In some preferred embodiments, the detection model is a logistic regression model, which yields more accurate results.
[0087] In some preferred embodiments, the detection device includes a storage medium on which the detection model is stored.
[0088] In some preferred embodiments, the detection device includes a processor for running a detection model.
[0089] It should be noted that, given the 21 protein biomarkers identified in this application, establishing different detection models based on different algorithms is easily achievable by those skilled in the art using existing model building methods. Therefore, any detection model constructed using the 21 protein biomarkers of this application is applicable to this application. Once established, such a detection model can be embedded in any electronic device. Specifically, it can be located on a storage medium or on a processor; regardless of the installation method, the detection device can distinguish between early NASH and non-NASH.
[0090] Depending on the specific device configuration, in some preferred embodiments, the detection device includes a protein biomarker expression level receiving module for the sample to be tested. This receiving module includes at least one of the following modes: user manual input mode, alternative list import mode, and file import mode. Different receiving modes provide subjects with diverse options, improving the convenience of user testing.
[0091] The aforementioned test sample is a bodily fluid sample, specifically, it can be various bodily fluids, such as serum, urine, saliva, pleural effusion, ascites, etc. In this application, the preferred test sample is a serum sample.
[0092] Furthermore, the test samples are derived from any one or more of the following subjects: healthy individuals or NAFLD patients. In practical applications, test samples are often obese patients, some of whom may have liver disease, while others may be healthy. Therefore, the protocol described in this application is also applicable to early NASH screening in obese patients. It should be noted that NAFLD patients here do not include patients with liver fibrosis and cirrhosis, but only those with NASH and NAFLD.
[0093] In a fourth typical embodiment of this application, a detection method for distinguishing between early NASH and non-NASH is provided. This method includes: detecting the expression levels of protein markers in the body fluids of a subject to obtain the expression levels of the target protein; inputting the expression levels of the target protein into a detection model for distinguishing between early NASH and non-NASH, and outputting the detection results; wherein the detection model is a model for detecting protein markers, and the protein markers include those detected in any of the aforementioned kits. Using the 21 protein markers discovered in this application, different detection models can be established based on any number of them according to different clinical classification standards. All of these detection models can effectively distinguish between early NASH and non-NASH.
[0094] The aforementioned detection model can employ various known classifier models, such as the random forest model and the SVC model, both of which can effectively classify the two types of patients. In some preferred embodiments, the detection model is a logistic regression model, which yields more accurate results.
[0095] The bodily fluids of the aforementioned subjects include, but are not limited to, serum, urine, saliva, pleural effusion, and ascites. Serum is preferred in this application.
[0096] The subjects mentioned above can be anyone, including but not limited to any one or more of the following: healthy individuals or NAFLD patients. In practical applications, the test samples are usually obese patients, some of whom may have liver disease, while others may be healthy. Therefore, the protocol in this application is also applicable to early NASH screening in obese patients. It should be noted that NAFLD patients here do not include patients with liver fibrosis and cirrhosis, but only those with early NASH and NAFLD.
[0097] The beneficial effects of this application will be described in detail below with reference to specific embodiments.
[0098] Example 1: Screening of protein biomarkers
[0099] (I) Queue Samples
[0100] This invention utilizes serum samples from 179 obese patients with lipid metabolism disorders who underwent metabolic surgery, along with their corresponding complete clinical information and histopathological data. All samples (n=179) were grouped based on the scores of three indicators in liver biopsy pathological sections: steatosis, inflammation, and ballooning.
[0101] Healthy liver group: The score for each indicator is 0 (n=44);
[0102] NAFL group, also known as simple fatty liver group: steatosis score greater than 0, and scores of the other two items are 0 (n=55);
[0103] Early NASH group: also known as non-alcoholic steatohepatitis group: scores of all three pathological indicators were greater than 0 (n=80).
[0104] Because this cohort was designed for individuals with early-stage NASH, no significant fibrosis was observed in any of the liver biopsy samples. Clinical information included sex, age, body mass index (BMI), waist-to-hip ratio (WHR), insulin resistance index (HOMAIR), C-peptide, alanine aminotransferase (ALT), aspartate aminotransferase (Ast), and total bile acids (Tba).
[0105] (II) Protein Detection
[0106] This embodiment primarily utilizes proximity extension assay (PEA; Olink platform) technology to measure the concentration of corresponding protein biomarkers in blood samples. PEA is a high-throughput immunoassay for detecting protein biomarkers in liquid samples. It uses a pair of matched antibodies conjugated with single-stranded DNA oligonucleotides to detect proteins (i.e., when two antibodies conjugated with single-stranded DNA oligonucleotides bind to the same target protein, the two DNA oligonucleotide strands combine to form a double-stranded DNA molecule, which is then detected by real-time PCR extension amplification). The detection result is expressed as npx (normalized protein expression) values. This invention used a total of 12 detection panels, containing 1066 proteins. Each well contained internal control samples (2 incubation controls, 1 extension control, and 1 detection control). Each detection panel also included negative controls, positive controls, and inter-panel controls for experimental quality control. The detection method followed the kit's instructions. All samples were tested in three batches (see Tables 2 and 3 below for details).
[0107] Table 2:
[0108]
[0109] Table 3:
[0110]
[0111] (III) Data Preprocessing
[0112] For Olink protein data, quality control was first performed on individual samples based on the test results of the internal control samples. If the difference between the test value of a single well of the internal control sample and the median test value of the entire plate of internal control samples exceeded the threshold (0.35 NPX), the quality control for that well was considered to have failed, and the test result of that sample was removed and recorded as a missing value (NA). Furthermore, the limit of detection (LOD) for each protein indicator was determined based on the test results of the negative control samples. When the test value was below the LOD, the protein indicator was considered undetectable and recorded as a missing value (ND). Proteins with a missing percentage >25% were removed. Samples with a missing percentage >80% were removed. Missing values NA in the numerical matrix were filled using the KNN method (a known algorithm for imputing missing values); missing values ND were filled using the LOD. Batch correction was performed on multiple batches of data using the median method to obtain the final data analysis matrix.
[0113] For clinical data, the KNN method is used to impute missing values in the clinical data.
[0114] (iv) Protein biomarker screening
[0115] The data was randomly split into a training set (70%) and a test set (30%). Protein biomarkers were selected from the training set. First, protein variables were filtered by collinearity calculation, i.e., the Pearson correlation coefficient matrix of all variables was calculated, and hclust clustering was performed. Only the most important variable in each cluster was retained, i.e., the variable with the smallest p-value in the t-test (t-test = Student's t-test, p-value = p-value). Then, univariate logistic regression analysis (using the glm method) was performed to calculate the coefficient of each variable in the logistic regression. Variables with significant p-values (p < 0.05) were retained. This resulted in 21 significantly different proteins (as shown in Table 4) as protein biomarkers to distinguish between non-NASH and early NASH. From these 21 protein biomarkers, one or more of the following were further selected as biomarkers: IL-1ra, TRAIL-R2, ALCAM, SELE, PON3, ROBO1, CDCP1, PTS, Insulin, and FBP1.
[0116] Table 4:
[0117]
[0118]
[0119] Note: In the table above, "increase" indicates that the expression level of the corresponding protein is significantly increased in the early NASH population compared with the NON-NASH population; "decrease" indicates a significant decrease.
[0120] Example 2: Evaluation of the performance of protein biomarkers in distinguishing between non-NASH and early NASH.
[0121] (1) Calculate the AUC value of each protein biomarker to distinguish between NON-NASH and early NASH to determine the accuracy of the protein biomarkers.
[0122] All individual proteins identified in the initial screening showed an AUC greater than 0.65 (0.656-0.796) in distinguishing between non-NASH and early NASH (when expression levels decrease in the NASH population, the value is (1-AUC)) (see Table 4). This result indicates that each of the 21 proteins has the ability to distinguish between early NASH and non-NASH. Among these 21 protein biomarkers, CDCP1 performed best in the test set (AUC = 0.796). The concentrations of most proteins gradually increased with disease progression, except for PON3, whose concentration gradually decreased with disease progression. Figure 1a and Figure 1b ).
[0123] (2) Calculate the distinguishing ability of various protein combinations.
[0124] Based on the discovered protein biomarkers, classification models were built using random combinations of different numbers (n = 2-21) of proteins. Various methods were explored to validate the detection capabilities of these random protein combinations, such as logistic regression, random forest, and SVC models, and the prediction results were found to be quite good. The following explanation uses the logistic regression model as an example. In the following examples, different combinations of protein biomarkers were used to build logistic regression models. The models were trained on the training set and validated on the test set, with AUC and other values calculated.
[0125] The models formed by these randomly combined proteins, after being trained on the training set, all showed high AUCs on the test set (Table 5). This result indicates that any combination of two or more of these 21 proteins has a high ability to distinguish between Non-NASH and early NASH. Among them, the 10 preferred proteins are IL-1ra, TRAIL-R2, ALCAM, SELE, PON3, ROBO1, CDCP1, PTS, Insulin, and FBP1. Any combination of two or more of these proteins has a good ability to distinguish between Non-NASH and early NASH, as shown in combinations numbered 5, 9, 13, 16, 30, and 35 in Table 5.
[0126] Table 5:
[0127]
[0128]
[0129]
[0130]
[0131]
[0132] Note: In the table above, AUC stands for Area under the ROC curve.
[0133] TPR: True positive rate;
[0134] TNR: True negative rate;
[0135] NPV: Negative predictive value;
[0136] PPV: Positive predictive value.
[0137] The 21 protein biomarkers discovered in this application, along with the detection models, kits, and chips established using these biomarkers, can differentiate between early-stage NASH and non-NASH in various application scenarios, including NASH screening for healthy individuals, especially obese individuals. Specifically, by inputting the patient's relevant serum protein levels (input methods are not limited to manual input, alternative list input, or file import), and performing corresponding algorithm model calculations, the study can assess whether the subject has NASH.
[0138] As can be seen from the above description, the embodiments of this application achieve the following technical effects: The protein biomarkers screened in this application, through screening in two population groups—early NASH and non-NASH—resulted in 21 protein biomarkers significantly correlated with early NASH. This helps to accurately distinguish between early NASH and non-NASH patients, thereby reducing the risk of disease progression to hepatocellular carcinoma and improving patient survival rates. In the current lack of commercially available non-invasive testing methods that can effectively differentiate between non-NASH and early NASH, the protein biomarkers discovered in this application and their derived testing products can, to some extent, fill this diagnostic gap and effectively improve the differentiation of the non-NASH disease spectrum. Patient clinical information collection and serum protein biomarker concentration measurement are relatively simple, non-invasive, accurate, and highly acceptable to patients, thus possessing significant clinical application value.
[0139] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. The application of protein biomarkers in the preparation of detection products for differentiating early NASH from non-NASH, characterized in that, The protein biomarkers are selected from any one of the following groups:
2. The application according to claim 1, characterized in that, The testing product is a test kit or a test device.
3. A detection device for distinguishing between early NASH and non-NASH, characterized in that, The detection device has a built-in detection model that distinguishes between NASH and NON-NASH, wherein the detection model is a model for detecting protein biomarkers, and the protein biomarkers are any one of the protein biomarkers described in the application of claim 1 or 2.
4. The detection device according to claim 3, characterized in that, The detection model is a logistic regression model.
5. The detection device according to claim 4, characterized in that, The detection device includes a storage medium, and the detection model is stored on the storage medium.
6. The detection device according to claim 4, characterized in that, The detection device includes a processor for running the detection model.
7. The detection device according to claim 3, characterized in that, The detection device includes a protein biomarker expression level receiving module for the sample to be tested, and the receiving module includes at least one of the following modes: user manual input mode, alternative list import mode, or file import mode.
8. The detection device according to claim 7, characterized in that, The sample to be tested is a bodily fluid sample.
9. The detection device according to claim 8, characterized in that, The sample to be tested is a serum sample.
10. The detection device according to claim 7, characterized in that, The test samples are derived from any one or more of the following subjects: healthy individuals or NAFLD patients.
11. A detection method for distinguishing between early NASH and NON-NASH for non-diagnostic purposes, characterized in that, The detection method includes: The expression levels of protein markers in the body fluids of the subjects were detected to obtain the expression levels of the target proteins; The expression level of the protein to be tested is input into a detection model used to distinguish between early NASH and NON-NASH, and the detection results are output. Wherein, the detection model is a model for detecting protein biomarkers, and the protein biomarkers are any one set of the protein biomarkers in the application described in claim 1 or 2.
12. The detection method according to claim 11, characterized in that, The detection model is a logistic regression model.
13. The detection method according to claim 11, characterized in that, The body fluid in question is serum.
14. The detection method according to claim 11, characterized in that, The subjects were selected from any one or more of the following: healthy individuals or NAFLD patients.
Citation Information
Patent Citations
Application of CD166 as liver cancer diagnosis serum marker and kit thereof
CN105628919A
Methods for evaluation of gestational progress and preterm abortion for clinical intervention and applications thereof
CN113396332A
Methods for Evaluation of Gestational Progress and Preterm Abortion for Clinical Intervention and Applications Thereof
US20210398682A1
Treating chronic liver disease
US20220220206A1