A liver injury model special for hujian buzhu granules and a construction method thereof
By combining a specific intervention regimen and dynamic data acquisition with hepatoprotective buzhure granules in an SD rat model, and utilizing a Stacking ensemble prediction model, the limitations of specificity and predictive ability in the evaluation of liver injury models in existing technologies were addressed. This enabled accurate prediction of the early efficacy of hepatoprotective buzhure granules, improving the efficiency and robustness of drug development.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YECHENG COUNTY PEOPLES HOSPITAL
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies, general liver injury models lack specificity and systematicity when evaluating hepatoprotective granules. The efficacy evaluation dimensions are singular and severely lagging, making it impossible to effectively predict the integrated efficacy in the early stages of medication, resulting in long drug development cycles and high costs.
Using an acetaminophen-induced liver injury model in SD rats, and combining acute treatment and long-term preventive intervention with hepatoprotective buzure granules, early feature vectors were obtained through a dynamic data acquisition and processing unit. The Stacking ensemble prediction model was used to predict the comprehensive efficacy of hepatoprotective buzure granules, which included a combination of random forest regression, extreme gradient boosting regression base learners, and ridge regression learners.
A highly structured, dedicated evaluation system was developed, which can accurately predict drug efficacy at an early stage, improve drug development efficiency, solve the problems of high heterogeneity and poor predictability of existing model evaluation results, and has broad adaptability and robustness.
Smart Images

Figure SMS_1 
Figure SMS_2
Abstract
Description
Technical Field
[0001] This invention relates to the field of animal model technology, and in particular to a liver injury model specifically designed for hepatoprotective Buzure Granules and its construction method. Background Technology
[0002] Drug-induced liver injury (DILI) is a common clinical complication and a leading cause of acute liver failure. Furthermore, DILI is a major contributor to new drug development failure, drug withdrawal from the market, and acute liver failure. Conversely, establishing reliable liver injury models can help assess drug hepatotoxicity and the efficacy of hepatoprotective drugs.
[0003] Currently, preclinical research mainly relies on two types of models: traditional animal models (such as rats and mice) and emerging in vitro complex models (such as 3D organoids and organ-on-a-chip).
[0004] While traditional animal models can provide systemic physiological responses, significant species differences limit their accuracy in predicting human DILI. Many drugs shown to be safe in animals can induce liver damage in humans. Furthermore, animal experiments face significant ethical challenges, high costs, and long processing times.
[0005] Meanwhile, in vitro models, such as liver organoids, have shown advantages in simulating the metabolism and microstructure of human hepatocytes. However, these advanced models currently focus more on toxicity mechanism research and early safety warning, and still have functional limitations in systematically simulating gut-hepatic axis interactions, host systemic inflammatory responses, and quantitatively evaluating the overall efficacy of specific protective drugs.
[0006] In particular, when evaluating specific drugs such as hepatoprotective granules (e.g., those exhibiting hepatoprotective activities such as anti-hepatic fibrosis and anti-inflammation), existing animal models not only lack specificity and systematicity but also suffer from lagging and limited treatment evaluation. This is because existing general animal models are disconnected from the specific intervention scenarios of these drugs, failing to internalize the dosing regimen as standard parameters for the model, resulting in insufficient specificity and reproducibility of the evaluation. Furthermore, existing evaluation systems rely entirely on static indicators at the end of treatment, failing to capture dynamic responses in key dimensions such as gut microbiota and host metabolism in the early stages of drug administration. Consequently, they cannot predict and quantify the final integrated efficacy early, leading to long development cycles and high costs.
[0007] Therefore, there is an urgent need for a dedicated liver injury model and its construction method that is deeply bound to specific hepatoprotective drugs such as hepatoprotective granules, in order to solve the fundamental problems of lagging evaluation, single dimension and lack of predictive ability in the existing technology. Summary of the Invention
[0008] The purpose of this invention is to provide a liver injury model specifically for hepatoprotective Buzure granules and its construction method, in order to solve the key technical problems of existing general liver injury models when evaluating this specific drug, such as lack of specificity and systematicity, single and seriously lagging efficacy evaluation dimensions, and inability to effectively predict the endpoint integrated efficacy in the early stage of drug use.
[0009] To achieve the above objectives, the present invention provides the following solution:
[0010] A liver injury model specifically designed for hepatoprotective Buzure Granules, the key feature of which is that the model specifically includes:
[0011] Animal model main unit: The above animal model main unit uses SD rats with liver injury induced by acetaminophen as the basic model, and pre-sets two liver-protecting Buzure Granule intervention programs: acute treatment and long-term prevention.
[0012] Dynamic data acquisition and processing unit: The above-mentioned dynamic data acquisition and processing unit is configured to collect and process fecal and serum samples at an early time point after intervention with liver-protecting Buzure granules, and output an early feature vector consisting of feature values including the relative abundance change rate of Lactobacillus and Akkermansia, serum CA / DCA ratio and butyrate concentration.
[0013] Comprehensive efficacy prediction unit: The comprehensive efficacy prediction unit integrates a Stacking ensemble prediction model with fixed parameters. The Stacking ensemble prediction model includes a first-layer random forest regression and extreme gradient boosting regression base learner and a second-layer ridge regression learner. Taking the early feature vector output by the dynamic data acquisition and processing unit as input, it can directly predict and output a comprehensive efficacy score that characterizes the liver-protecting Buzure granules at the preset efficacy endpoint.
[0014] Furthermore, in the main animal model unit, the above-mentioned acute treatment intervention protocol involved a single administration of hepatoprotective buzhure granules within 30 minutes after acetaminophen-induced liver injury; the above-mentioned long-term prevention intervention protocol involved pretreatment with hepatoprotective buzhure granules for several consecutive days before acetaminophen-induced liver injury.
[0015] Furthermore, the aforementioned early time points in the dynamic data acquisition and processing unit include 3 hours and 6 hours after the intervention.
[0016] Furthermore, the early feature vector output by the aforementioned dynamic data acquisition and processing unit consists of eight feature values: the relative abundance change rate of Lactobacillus spp. at 3h and 6h after intervention, the relative abundance change rate of Akkermansia spp., the serum CA / DCA ratio, and the butyrate concentration.
[0017] Specifically, in the aforementioned integrated efficacy prediction unit, the mathematical expression of the Stacking integrated prediction model is shown in Equation 6.
[0018] CES pred = w4 + w5 × RF(X) T1 ) + w6×XGBR (X T1 Equation 6
[0019] In Equation 6, CES pred The predicted comprehensive efficacy score; RF() and XGBR() represent the trained random forest and XGBoost base model functions, respectively; w4, w5, and w6 are the optimal integration parameters automatically learned by the meta-learner based on meta-features, and their values are determined after the model is solidified.
[0020] Specifically, the aforementioned comprehensive efficacy prediction unit was validated on an independent test set, and the Pearson correlation coefficient between its predicted comprehensive efficacy score and the actual comprehensive efficacy score was greater than 0.85.
[0021] The key to constructing the liver injury model specifically for the liver-protecting Buzure Granules described above lies in the following steps:
[0022] S1. Animal model preparation and intervention: An acetaminophen-induced liver injury SD rat model was constructed, and hepatoprotective Buzure granules were administered according to the pre-set acute treatment or long-term prevention plan.
[0023] S2. Construction of dynamic feature dataset: Samples are collected at early time points after intervention, and early feature vectors for each animal are detected and calculated;
[0024] S3. Construction of endpoint efficacy dataset: Samples are collected at the preset efficacy endpoint, multidimensional efficacy indicators are detected and the endpoint comprehensive efficacy score of each animal is calculated.
[0025] S4. Predictive Model Construction, Consolidation and Integration: Using the aforementioned early feature vectors as input and the aforementioned endpoint comprehensive efficacy score as the target, train and consolidate a Stacking integrated predictive model; integrate the consolidated predictive model into the aforementioned comprehensive pharmacodynamic prediction unit, thereby constructing the comprehensive pharmacodynamic prediction unit of the aforementioned dedicated liver injury model.
[0026] Specifically, in step S3, the overall efficacy score of the above endpoint is calculated according to Equation 2:
[0027] CES T2 =w1×f(ALT T2 )+w2×f(Ab 乳杆菌T2 )+w3×f(GCA T2 Equation 2,
[0028] In Equation 2, f represents the recovery percentage relative to the model group, specifically including:
[0029] f(ALT T2 )=(1 - ALT T2 / M ALT Formula 3, )×100%
[0030] f(Ab 乳杆菌T2 )=(Ab 乳杆菌T2 / C Ab Formula 4, )×100%
[0031] f(GCA T2 )=(1 - GCA T2 / M GCA Formula 5, )×100%
[0032] In Equations 2 to 5, CES T2 The overall efficacy score representing the endpoint, f(ALT) T2 f(Ab) represents the recovery rate of serum alanine aminotransferase activity. 乳杆菌T2 ) represents the recovery rate of absolute abundance of Lactobacillus spp. in the gut, f(GCA) T2 The value represents the improvement rate of liver glycocholic acid content. w1, w2, and w3 are weights calculated using principal component analysis of the effective sample data from the liver-protecting Buzure granule-induced group. ALT T2 M represents the endpoint serum alanine aminotransferase activity in this animal. ALT The average serum ALT activity represents the value of the control group in the model; Ab 乳杆菌T2 C represents the absolute abundance of *Lactobacillus* in the animal's gut at the endpoint. Ab GCA represents the average absolute abundance of *Lactobacillus* spp. in the control group; T2 M represents the final liver glycocholic acid content of the animal. GCA The average value of glycocholic acid content in the liver of the representative model control group.
[0033] More specifically, step S4 includes:
[0034] S41. Modeling Dataset Construction and Preprocessing: Based on the results of steps S2 and S3, the early feature vector of each animal is used as the input feature, and the endpoint comprehensive efficacy score is used as the target variable to construct the modeling dataset; and the above input features are standardized and preprocessed.
[0035] S42. Training and Consolidation of Stacking Ensemble Prediction Models: Using the Stacking ensemble learning framework, the following operations are performed:
[0036] S421. Base Learner Training and Optimization: Divide the above dataset into training and test sets; on the training set, use cross-validation and grid search to optimize the hyperparameters of the random forest regression model and the extreme gradient boosting regression model; the optimization parameters of the random forest regression model include the number of decision trees and the maximum depth of the trees, and the optimization parameters of the extreme gradient boosting regression model include the number of boosting iterations, the maximum depth of the trees, and the learning rate.
[0037] S422, Meta-learner training: Based on the optimized base learner, cross-validation is used to generate a meta-feature matrix, and the ridge regression model is trained using the above meta-feature matrix as the meta-learner.
[0038] S423. Model solidification: Using all training set data, retrain the above base learners and meta learners with optimal hyperparameters to obtain a parameter-solidified Stacking ensemble prediction model; the mathematical expression of the above solidified model is shown in Equation 6.
[0039] Preferably, for the random forest regression model, the search range for the number of decision trees is set to 100-500, and the search range for the maximum depth of the trees is set to 5-15; for the extreme gradient boosting regression model, the search range for the number of boosting iterations is set to 100-300, the search range for the maximum depth of the trees is set to 3-8, and the search range for the learning rate is set to 0.01-0.1.
[0040] The present invention discloses the following technical effects:
[0041] The liver injury model and its construction method specifically for hepatoprotective Buzure granules provided by this invention systematically integrate standardized drug intervention protocols, dynamic monitoring of multi-timepoint multidimensional biomarkers, and artificial intelligence prediction models. This constructs a dedicated evaluation system capable of accurately predicting early efficacy, effectively solving the technical problems of weak specificity, poor predictive ability, and limited system dimensions in existing general models when evaluating specific drugs. Specific technical effects are as follows:
[0042] First, this invention constructs a dedicated liver injury model deeply integrated with Hugan Buzure Granules. Existing general models are severely disconnected from specific intervention scenarios such as acute treatment or long-term prevention of drugs. This invention, however, innovatively internalizes two clearly defined intervention protocols for Hugan Buzure Granules into the model's standard parameters and integrates the entire process from early dynamic data collection to endpoint efficacy calculation. This makes the model constructed by this invention no longer an isolated disease carrier, but a highly structured and reproducible dedicated evaluation system, fundamentally ensuring the specificity of the evaluation for Hugan Buzure Granules and overcoming the inherent defect of high heterogeneity in evaluation results of general models.
[0043] Secondly, this invention endows drug efficacy evaluation with powerful early predictive capabilities, significantly improving the efficiency of drug development. Traditional evaluation relies entirely on static end-stage indicators, which are time-consuming and lack predictive power. This invention captures the dynamics of gut microbiota and serum metabolic responses at two key time points, 3 hours and 6 hours after drug administration, constructing an early feature vector containing eight features. By training a fixed predictive model using the Stacking ensemble learning algorithm, it can accurately predict the overall efficacy 24 hours later based on the aforementioned early data, providing a revolutionary and efficient tool for early drug screening and mechanism research.
[0044] Third, the method established in this invention demonstrates high robustness and scalability. As shown in the embodiments, the dedicated prediction model maintains excellent predictive performance in both scheme A and scheme B, two drastically different intervention scenarios, proving the broad adaptability of the method to different medication logics. Furthermore, the core value of the pre-designed low, medium, and high dose groups in the model lies in providing the machine learning model with diverse, high-quality training data covering the entire therapeutic spectrum, ensuring the model's generalization ability and predictive robustness.
[0045] In summary, this invention has made significant progress in terms of the specificity of evaluation, the forward-looking nature of prediction, and the robustness of the method by creating a system of dedicated intervention programs, multi-time-point dynamic monitoring, and artificial intelligence prediction, providing an innovative solution for the precise and efficient development of hepatoprotective drugs. Detailed Implementation
[0046] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as a limitation of the present invention, but rather as a more detailed description of certain aspects, features, and embodiments of the present invention.
[0047] It should be understood that the terminology used in this invention is merely for describing particular embodiments and is not intended to limit the invention. Furthermore, with respect to numerical ranges in this invention, it should be understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. Any stated value or intermediate value within a stated range, as well as each smaller range between any other stated value or intermediate value within said range, is also included in this invention. The upper and lower limits of these smaller ranges may be independently included or excluded from the range.
[0048] Unless otherwise stated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. While only preferred methods and materials have been described herein, any methods and materials similar or equivalent to those described herein may be used in the implementation or testing of this invention. All references to this specification are incorporated by way of citation to disclose and describe methods and / or materials associated with those references. In the event of any conflict with any incorporated reference, the content of this specification shall prevail.
[0049] Various modifications and variations can be made to the specific embodiments described in this specification without departing from the scope or spirit of the invention, as will be apparent to those skilled in the art. Other embodiments derived from this specification will also be apparent to those skilled in the art. This specification and embodiments are merely exemplary.
[0050] The terms “include,” “including,” “have,” “contain,” etc., used in this article are all open-ended terms, meaning that they include but are not limited to.
[0051] Example 1
[0052] This embodiment provides a liver injury model specifically designed for hepatoprotective buzhure granules (HBG) and its construction method, the specific steps of which include:
[0053] S1. Animal Model Preparation and Intervention:
[0054] S11. Laboratory Animals and Grouping:
[0055] One hundred SPF-grade male SD rats were randomly divided into three groups of 20 each: a blank control group, a liver injury model group, and three HBG treatment groups.
[0056] The HBG treatment group can be divided into three dose groups according to the research objectives, specifically including:
[0057] HBG low-dose group: 25 mg / kg; HBG medium-dose group: 50 mg / kg; HBG high-dose group: 100 mg / kg.
[0058] S12. Preparation of liver injury model:
[0059] This embodiment provides an intervention plan (Plan A) for an acute treatment model, used to simulate immediate drug treatment after acute liver injury. Specifically, it includes:
[0060] Except for the control group, rats in all other groups were given a single intraperitoneal injection of acetaminophen solution at a dose of 500 mg / kg to establish an acute liver injury model; the control group was given an equal volume of physiological saline at the same time point.
[0061] Within 30 minutes after modeling, the animals in the HBG treatment group were given a single oral gavage of the corresponding dose of hepatoprotective Buzhure granules suspension.
[0062] The control group and the liver injury model group were given the same volume of physiological saline at the same time point, and no further medication was given until the endpoint sampling.
[0063] This embodiment provides a long-term prevention model (Plan B) to evaluate the preventive protective effect of the drug against liver injury. Specifically, it includes:
[0064] Before formal modeling, the animals in the HBG treatment group were given the corresponding dose of liver-protecting Buzure granules suspension by gavage once a day for 15 consecutive days as a pretreatment.
[0065] The control group and the liver injury model group were given the same volume of physiological saline daily;
[0066] Within 30 minutes after the last administration, liver injury model groups and HBG treatment groups were established, while the blank group was given normal saline.
[0067] S2. Construction of Dynamic Feature Dataset:
[0068] Stool and serum samples were collected at 3 h and 6 h after HBG administration (defined as early time points, denoted as T1) for early characteristic detection, specifically including:
[0069] Early indicators of gut microbiota dynamics: High-throughput sequencing of the V3-V4 region of 16S rRNA gene was performed on T1 fecal samples. The rate of change of the relative abundance of Lactobacillus and Akkermansia before self-induction was calculated using Equation 1.
[0070] Rate of change = (RA') x - RA x ) / RA x ×100% Formula 1;
[0071] In Equation 1, RA x Represents the baseline relative abundance of a specific bacterial genus before the administration of a liver injury inducer; x is the representative symbol for the specific bacterial genus; RA' x This represents the relative abundance of a specific bacterial genus at a predetermined time point after the administration of the liver injury inducer.
[0072] Early metabolic response substances in serum: Targeted metabolomics was performed on T1 serum to detect the molar ratio of primary bile acids (cholic acid) to secondary bile acids (deoxycholic acid) (i.e., CA / DCA ratio) and the serum concentration of short-chain fatty acid butyric acid (i.e., butyric acid concentration).
[0073] S3. Construction of the endpoint efficacy dataset:
[0074] Blood, liver tissue, and intestinal contents samples were collected at the pre-defined efficacy evaluation endpoint (T2; for both Protocol A and Protocol B, this was 24 hours after model establishment) for multidimensional efficacy assessment. Specifically, this included:
[0075] (1) Host inflammation and injury dimension: serum alanine aminotransferase activity;
[0076] (2) Intestinal flora homeostasis: endpoint absolute abundance of Lactobacillus and Akkermansia;
[0077] (3) Liver metabolic dimension: content of glycocholic acid and glutamic acid in liver tissue.
[0078] The comprehensive efficacy score is calculated using the formula shown in Equation 2:
[0079] CES T2 =w1×f(ALT T2 )+w2×f(Ab 乳杆菌T2 )+w3×f(GCA T2 Equation 2,
[0080] In Equation 2, f represents the recovery percentage relative to the model group, specifically including:
[0081] f(ALT T2 )=(1 - ALT T2 / M ALT Formula 3, )×100%
[0082] f(Ab 乳杆菌T2 )=(Ab 乳杆菌T2 / C Ab Formula 4, )×100%
[0083] f(GCA T2 )=(1 - GCA T2 / M GCA Formula 5, )×100%
[0084] In Equations 2 to 5, CES T2 The overall efficacy score representing the endpoint, f(ALT) T2 f(Ab) represents the recovery rate of serum alanine aminotransferase activity. 乳杆菌T2 ) represents the recovery rate of absolute abundance of Lactobacillus spp. in the gut, f(GCA) T2 The value represents the improvement rate of liver glycocholic acid content. w1, w2, and w3 are weights calculated from the effective sample data of the medium-dose HBG group through principal component analysis, and are set to 0.4, 0.3, and 0.3, respectively.
[0085] ALT T2 M represents the endpoint serum alanine aminotransferase activity in this animal. ALTThe average serum ALT activity in the control group of the representative model;
[0086] Ab 乳杆菌T2 C represents the absolute abundance of *Lactobacillus* in the animal's gut at the endpoint. Ab The average value representing the absolute abundance of Lactobacillus spp. in the gut of the control group;
[0087] GCA T2 M represents the final liver glycocholic acid content of the animal. GCA The average value of glycocholic acid content in the liver of the representative model control group.
[0088] It should be noted that in this embodiment, the relative abundance of the bacterial community in step S2 is a proportion calculated based on high-throughput sequencing reads and is used to calculate early change characteristics; while the absolute abundance used for efficacy evaluation in step S3 refers to the copy number of the bacterial community genome measured by PCR quantification method. The two are different measurement indicators.
[0089] S4. Predictive model construction, solidification, and integration:
[0090] S41. Modeling Dataset Construction and Preprocessing:
[0091] The early feature set obtained for each animal at time point T1 (including: the change rate of Lactobacillus spp., the change rate of Akkermansia spp., CA / DCA ratio, butyrate concentration at 3h, and the corresponding 6h features, for a total of 8 features) is taken as a feature vector, denoted as X. T1 ;
[0092] The corresponding endpoint comprehensive efficacy score (CES) calculated at time point T2. T2 As the target variable, it forms a set of several groups (X) T1 CES T2 A complete modeling dataset;
[0093] To ensure the effectiveness of model training, the early feature set is standardized before training so that the mean of each feature is 0 and the standard deviation is 1.
[0094] S42, Training and Consolidation of Stacking Ensemble Prediction Models:
[0095] A dedicated prediction model is built using the Stacking ensemble learning framework;
[0096] The model consists of two layers:
[0097] The first layer consists of multiple heterogeneous base learners, designed to learn the complex relationship between early features and endpoint efficacy from different dimensions;
[0098] The second layer consists of a meta-learner, which is used to optimally integrate the outputs of the base learners to form the final prediction.
[0099] S421, Base Learner Training and Optimization:
[0100] Random forest regression and extreme gradient boosting regression were selected as the base learners for the first layer;
[0101] Before model training, the complete dataset is randomly divided into a training set and an independent test set in a 7:3 ratio;
[0102] Subsequently, on the training set, a five-fold cross-validation strategy was adopted, combined with a grid search method, to optimize the key hyperparameters of the two base learners respectively;
[0103] For the random forest regression model, the core parameters to optimize include:
[0104] For the random forest regression model, the search range for the number of decision trees is set to 100–500, and the search range for the maximum depth of the trees is set to 5–15.
[0105] For the extreme gradient boosting regression model, the search range for the number of boosting iterations is set to 100–300, the search range for the maximum tree depth is set to 3–8, and the search range for the learning rate is set to 0.01–0.1.
[0106] S422, Meta-learner training:
[0107] After optimizing the parameters of the base learner, cross-validation is used to generate meta-features, specifically:
[0108] The training set is divided into five equal parts. Four of these parts are used as training subsets to fit base learners, and the remaining one part is used to predict the validation subset.
[0109] After traversing all five samples, obtain the data-leakage-free prediction values of each training sample on the two base learners, and concatenate these two prediction values into a two-dimensional meta-feature matrix.
[0110] Using this meta-feature matrix as the new input, with the original CES... T2 To achieve this, a ridge regression model is trained as the meta-learner for the second layer, and its regularization strength parameter is also selected through cross-validation optimization.
[0111] S423, Model Solidification:
[0112] To obtain the final deployable prediction model, the two base learners, Random Forest Regression and Limiting Gradient Boosting Regression, were retrained using all training set data and the optimal hyperparameters determined in step S421.
[0113] X the early features of the entire training set T1 Input these two trained base learners, obtain their predicted outputs, and use them to generate the final meta-features;
[0114] Finally, using this meta-feature and the corresponding endpoint efficacy score, the final meta-learner (ridge regression model) is trained.
[0115] Thus, a Stacking integrated prediction pipeline with a fixed structure and fixed parameters is constructed.
[0116] The pipeline receives early feature vectors (X) T1 As input, it sequentially executes the predictions of two base learners, and then inputs their predictions into a meta-learner for weighted integration, ultimately outputting the overall efficacy score (CES) at the endpoint. T2 The predicted value of CES pred The mathematical expression of this composite model can be summarized as follows:
[0117] CES pred = w4 + w5 × RF(X) T1 ) + w6×XGBR (X T1 Equation 6
[0118] In Equation 6, RF() and XGBR() represent the trained random forest and XGBoost base model functions, respectively. w4, w5, and w6 are the optimal integration parameters automatically learned by the meta-learner (ridge regression) based on the meta-features. Their values are determined after the model is solidified and are 0.02, 0.63, and 0.35, respectively.
[0119] S43. Model Validation and Performance Confirmation:
[0120] The early features X of the independent test set reserved in step S41 T1 The data is input into the Stacking integrated prediction model solidified in step S423 to obtain its predicted efficacy score (CES). pred ;
[0121] Calculate CES pred The Pearson correlation coefficient (r) and root mean square error between the comprehensive efficacy score and the true endpoint of the test set samples are used to determine the effectiveness of a specific predictive model. For a model to be effective, the Pearson correlation coefficient (r) should be greater than 0.85.
[0122] Meanwhile, paired-samples t-tests should be used to compare the RMSE of the Stacking ensemble model and the single base learner model on the test set to confirm that the prediction performance of the ensemble model is significantly better than that of any single base learner model (p < 0.05).
[0123] Example 2
[0124] This embodiment aims to demonstrate how to use the model constructed in Embodiment 1 to build and verify a dedicated model based on Scheme A. Specifically, it includes:
[0125] S1. Animal Model Preparation and Intervention:
[0126] S11. Experimental animals and grouping: Perform according to scheme A in step S1 of Example 1.
[0127] S2. Construction of dynamic feature dataset: Follow step S2 of Example 1 to collect and test samples at 3h and 6h after HBG administration; Table 1 shows the feature data of some samples in the HBG medium-dose group at this early time point.
[0128] Table 1: Early characteristics, actual efficacy, and predictive data of some samples (Plan A)
[0129]
[0130] S3. Construction of endpoint efficacy dataset: Follow the steps in S3 of Example 1. Some results are shown in Table 1.
[0131] S4. Predictive model construction, solidification, and integration:
[0132] S41. Modeling dataset construction and preprocessing: Perform step S41 of Example 1.
[0133] S42. Training and solidification of the Stacking integrated prediction model: Perform step S42 of Example 1.
[0134] S43. Model Validation and Performance Confirmation: Perform step S43 as described in Example 1, and use the early features X of the independent test set reserved in step S41. T1 The data is input into the Stacking integrated prediction model solidified in step S423 to obtain its predicted efficacy score (CES). pred Some results are shown in Table 1;
[0135] The Pearson correlation coefficient between the predicted and actual values of the model in this embodiment on the test set was calculated to be r=0.89, and the root mean square error (RMSE) was 8.37.
[0136] Using paired-samples t-tests, the prediction error of the Stacking ensemble model was significantly lower than that of the single random forest model (p=0.0032<0.01) and the XGBoost model (p=0.0124<0.05).
[0137] Example 3
[0138] This embodiment aims to demonstrate how to use the model constructed in Embodiment 1 to build and verify a dedicated model based on Scheme B. Specifically, it includes:
[0139] S1. Animal Model Preparation and Intervention:
[0140] S11. Experimental animals and grouping: Perform according to scheme B in step S1 of Example 1.
[0141] S2. Construction of dynamic feature dataset: Follow step S2 of Example 1 to collect and test samples at 3h and 6h after HBG administration; Table 2 shows the feature data of some samples in the HBG medium-dose group at this early time point.
[0142] Table 2: Early characteristics, actual efficacy, and predictive data for some samples (Program B)
[0143]
[0144] S3. Construction of endpoint efficacy dataset: Follow the steps in S3 of Example 1. Some results are shown in Table 2.
[0145] S4. Predictive model construction, solidification, and integration:
[0146] S41. Modeling dataset construction and preprocessing: Perform step S41 of Example 1.
[0147] S42. Training and solidification of the Stacking integrated prediction model: Perform step S42 of Example 1.
[0148] S43. Model Validation and Performance Confirmation: Perform step S43 as described in Example 1, and use the early features X of the independent test set reserved in step S41. T1 The data is input into the Stacking integrated prediction model solidified in step S423 to obtain its predicted efficacy score (CES). pred Some results are shown in Table 2;
[0149] The Pearson correlation coefficient between the predicted and true values of the model in this embodiment on the test set was calculated to be r=0.91, and the root mean square error (RMSE) was 7.84.
[0150] Using paired-samples t-tests for comparison, the prediction error of the Stacking ensemble model was significantly lower than that of the single random forest model (p=0.0023<0.01) and the XGBoost model (p=0.0087<0.05).
[0151] As can be seen from the above embodiments, the predictive performance of the dedicated model constructed using scheme B is slightly better than that of the model using scheme A. This is because long-term drug pretreatment can more significantly stabilize the gut microbiota and the host's metabolic basis, making the body's early response more active, orderly, and with relatively smaller individual differences after being subjected to the same liver injury induction. This makes the association between early characteristics and endpoint efficacy clearer, which is conducive to machine learning models capturing patterns and making more accurate predictions.
[0152] The embodiments of the present invention fully demonstrate that the dedicated liver injury model and method constructed by the present invention can be flexibly and reliably applied to two different drug intervention scenarios, namely acute treatment and long-term prevention, and have broad applicability and robustness.
[0153] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A liver injury model specifically designed for hepatoprotective Buzhu Relief Granules, characterized in that, The model specifically includes: Animal model main unit: The animal model main unit uses SD rats with liver injury induced by acetaminophen as the basic model, and pre-sets two liver-protecting Buzure Granule intervention programs: acute treatment and long-term prevention. Dynamic data acquisition and processing unit: The dynamic data acquisition and processing unit is configured to collect and process fecal and serum samples at an early time point after intervention with liver-protecting Buzure granules, and output an early feature vector consisting of feature values including the relative abundance change rate of Lactobacillus and Akkermansia, serum CA / DCA ratio and butyrate concentration. Comprehensive efficacy prediction unit: The comprehensive efficacy prediction unit integrates a Stacking ensemble prediction model with fixed parameters; the Stacking ensemble prediction model includes a first-layer random forest regression and extreme gradient boosting regression base learner and a second-layer ridge regression learner. Taking the early feature vector output by the dynamic data acquisition and processing unit as input, it can directly predict and output a comprehensive efficacy score characterizing the liver-protecting Buzure granules at the preset efficacy endpoint.
2. The model according to claim 1, characterized in that, The acute treatment intervention in the main animal model unit involves a single administration of hepatoprotective buzhure granules within 30 minutes after acetaminophen-induced liver injury; the long-term prevention intervention involves pretreatment with hepatoprotective buzhure granules for several consecutive days before acetaminophen-induced liver injury.
3. The model according to claim 1, characterized in that, The early time points mentioned in the dynamic data acquisition and processing unit include 3 hours and 6 hours after the intervention.
4. The model according to claim 3, characterized in that, The early feature vector output by the dynamic data acquisition and processing unit consists of eight feature values: the relative abundance change rate of Lactobacillus spp. at 3h and 6h after intervention, the relative abundance change rate of Akkermansia spp., the serum CA / DCA ratio, and the butyrate concentration.
5. The model according to claim 1, characterized in that, In the comprehensive efficacy prediction unit, the mathematical expression of the Stacking integrated prediction model is shown in Equation 6. CES pred = w4 + w5 × RF(X) T1 ) + w6×XGBR (X T1 Equation 6 In Equation 6, CES pred The predicted comprehensive efficacy score; RF() and XGBR() represent the trained random forest and XGBoost base model functions, respectively; X T1 w4, w5, and w6 are the early feature vectors; w4, w5, and w6 are the optimal integration parameters automatically learned by the meta-learner based on the meta-features, and their values are determined after the model is solidified.
6. The model according to claim 1, characterized in that, The comprehensive efficacy prediction unit was validated on an independent test set, and the Pearson correlation coefficient between its predicted comprehensive efficacy score and the actual comprehensive efficacy score was greater than 0.
85.
7. The method for constructing the model as described in any one of claims 1-6, characterized in that, Includes the following steps: S1. Animal model preparation and intervention: An acetaminophen-induced liver injury SD rat model was constructed, and hepatoprotective Buzure granules were administered according to the pre-set acute treatment or long-term prevention plan. S2. Construction of dynamic feature dataset: Samples are collected at early time points after intervention, and early feature vectors for each animal are detected and calculated; S3. Construction of endpoint efficacy dataset: Samples are collected at the preset efficacy endpoint, multidimensional efficacy indicators are detected and the endpoint comprehensive efficacy score of each animal is calculated. S4. Predictive model construction, solidification and integration: Using the early feature vector as input and the endpoint comprehensive efficacy score as target, train and solidify a Stacking integrated predictive model; integrate the solidified predictive model into the comprehensive efficacy prediction unit, thereby constructing the comprehensive efficacy prediction unit of the dedicated liver injury model.
8. The construction method according to claim 7, characterized in that, In step S3, the endpoint comprehensive efficacy score is calculated according to Equation 2: CES T2 = w1 × f(ALT T2 ) + w2 × f(Ab 乳杆菌T2 ) + w3 × f(GCA T2 ) Equation 2 In Equation 2, f represents the recovery percentage relative to the model group, specifically including: f(ALT T2 )=(1 - ALT T2 / M ALT )×100% Equation 3 f(Ab 乳杆菌T2 )=(Ab 乳杆菌T2 / C Ab )×100% Equation 4 f(GCA T2 ) = (1 - GCA T2 / M GCA ) × 100% Equation 5 In Equations 2 to 5, CES T2 The overall efficacy score representing the endpoint, f(ALT) T2 f(Ab) represents the recovery rate of serum alanine aminotransferase activity. 乳杆菌T2 ) represents the recovery rate of absolute abundance of Lactobacillus spp. in the gut, f(GCA) T2 The values represent the improvement rate of liver glycocholic acid content. w1, w2, and w3 are weights calculated using principal component analysis of the effective sample data from the liver-protecting Buzure granule-induced group. ALT T2 M represents the endpoint serum alanine aminotransferase activity in this animal. ALT The average serum ALT activity represents the value of the control group in the model; Ab 乳杆菌T2 C represents the absolute abundance of *Lactobacillus* in the animal's gut at the endpoint. Ab GCA represents the average absolute abundance of *Lactobacillus* spp. in the control group; T2 M represents the final liver glycocholic acid content of the animal. GCA The average value of glycocholic acid content in the liver of the representative model control group.
9. The construction method according to claim 7, characterized in that, Step S4 specifically includes: S41. Modeling Dataset Construction and Preprocessing: Based on the results of steps S2 and S3, the early feature vector of each animal is used as the input feature, and the endpoint comprehensive efficacy score is used as the target variable to construct the modeling dataset; and the input features are standardized and preprocessed. S42. Training and Consolidation of Stacking Ensemble Prediction Models: Using the Stacking ensemble learning framework, the following operations are performed: S421. Base Learner Training and Optimization: The dataset is divided into a training set and a test set; on the training set, cross-validation and grid search are used to optimize the hyperparameters of the random forest regression model and the extreme gradient boosting regression model; wherein, the optimization parameters of the random forest regression model include the number of decision trees and the maximum depth of the trees, and the optimization parameters of the extreme gradient boosting regression model include the number of boosting iterations, the maximum depth of the trees, and the learning rate. S422, Meta-learner training: Based on the optimized base learner, a meta-feature matrix is generated using cross-validation, and the ridge regression model is trained using the meta-feature matrix as the meta-learner. S423. Model Stabilization: Using all training set data, the base learner and meta-learner are retrained with optimal hyperparameters to obtain a parameter-stabilized Stacking ensemble prediction model; the mathematical expression of the stabilized model is shown in Equation 6. CES pred = w4 + w5 × RF(X) T1 ) + w6×XGBR (X T1 Equation 6 In Equation 6, CES pred The predicted comprehensive efficacy score; RF() and XGBR() represent the trained random forest and XGBoost base model functions, respectively; X T1 w4, w5, and w6 are the early feature vectors; w4, w5, and w6 are the optimal integration parameters automatically learned by the meta-learner based on the meta-features, and their values are determined after the model is solidified.
10. The construction method according to claim 9, characterized in that, For the random forest regression model, the search range for the number of decision trees is set to 100–500, and the search range for the maximum tree depth is set to 5–15; for the extreme gradient boosting regression model, the search range for the number of boosting iterations is set to 100–300, the search range for the maximum tree depth is set to 3–8, and the search range for the learning rate is set to 0.01–0.1.