Anti-hepatitis B virus drug metabolism kinetics prediction method based on Transform model
Through the pharmacokinetic prediction method based on the Transformer model, the accurate prediction problem of the metabolic process of tenofovir in patients with hepatitis B is solved, dynamically simulated the patient's metabolism, optimized the drug combination and dosage design, and improved the treatment effect.
Patent Information
- Application Number
- CN202510363568.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art is difficult to accurately predict the metabolic process of tenofovir prodrugs in patients with hepatitis B, especially in the case of liver damage, which affects the therapeutic effect and is difficult to optimize dosage design and combined treatment in drug development.
The metabolic kinetic prediction method of anti-hepatitis B virus drug based on the Transformer model was used to construct hepatocytes and in vivo pharmacokinetic models, combining feature encoding and multi-headed attention layer to predict the concentration of TFV-DP in human liver cells.
Dynamic simulation of the metabolic conditions of different patient groups has been achieved, helping drug development and optimize drug combination and dosage design, and improving the treatment effect of hepatitis patients.
Smart Images

Figure CN120280179A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - field of artificial intelligence and pharmacokinetics, and particularly relates to a method for predicting the pharmacokinetics of anti - hepatitis B virus drugs based on the Transformer model. Background Art
[0002] Chronic hepatitis B virus (HBV) infection is a widespread health problem globally. The prodrug of tenofovir (TFV) is the main treatment option. The TFV prodrug is hydrolyzed in vivo to TFV and further phosphorylated in hepatocytes to form its active metabolite tenofovir diphosphate (TFV - DP). TFV - DP inhibits virus replication by incorporating into the viral DNA strand and terminating DNA strand elongation, thus playing an antiviral therapeutic role. Therefore, the concentration of TFV - DP in target organs and target cells is an important indicator for evaluating the therapeutic effect of the TFV prodrug.
[0003] However, the generation process of TFV - DP involves multiple metabolic steps in hepatocytes. In patients with hepatitis B, liver damage (such as complications like cirrhosis or liver fibrosis) may affect the conversion efficiency of TFV and the generation of TFV - DP, thereby affecting the therapeutic effect. In clinical practice, direct measurement of drug concentration in the human liver is usually very difficult, and there are significant differences in drug metabolism among different individuals and different disease states; this makes it difficult to predict metabolism during the drug R & D stage, and thus difficult to assist in multiple aspects such as drug screening, optimization, and dose design in combination drug use. For this reason, a method for predicting the pharmacokinetics of anti - hepatitis B virus drugs based on the Transformer model is proposed. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a method for predicting the pharmacokinetics of anti - hepatitis B virus drugs based on the Transformer model, which solves the problems in the prior art.
[0005] The purpose of the present invention can be achieved by the following technical solutions:
[0006] A method for predicting the pharmacokinetics of anti - hepatitis B virus drugs based on the Transformer model for non - diagnostic purposes, comprising the following steps:
[0007] Obtain the in - vitro and in - vivo pharmacokinetic data of three TFV prodrugs in mice, beagle dogs, and humans through experiments and literature research to form an experimental data set;
[0008] Build a hepatocyte PK and in - vivo mPBPK model, and verify it. On this basis, develop a liver - damaged mouse / human cell PK and mPBPK model;
[0009] Based on the obtained experimental data set, cell PK and mPBPK models, construct a data set for Transformer model training;
[0010] Perform feature encoding on the data in the data set for training, and splice them to obtain feature vectors, and then splice them with position encoding to obtain input features;
[0011] Construct a Transformer model based on the encoder-only architecture, and input the input features into the Transformer model for training and validation;
[0012] Use the trained Transformer model to predict the concentration of TFV-DP in human liver cells.
[0013] Further, the three TFV prodrugs are: tenofovir disoproxil fumarate, tenofovir alafenamide, and tenofovir ambufenamide.
[0014] Further, four types of features are integrated in the data set for Transformer model training: drug concentration data, key metabolic parameters in the cell PK / mPBPK model, liver injury score, and time series, and a multi-variable input sequence is formed.
[0015] Further, the steps of feature encoding include:
[0016] Classify the data according to categorical features, numerical features, and time features;
[0017] According to the classification results, perform One-Hot encoding on species, drug indicators, and liver injury scores; perform Z-score standardization encoding on drug concentration data and key metabolic parameters in the cell PK / mPBPK model; directly retain the time values for the time series.
[0018] Further, before feature encoding, each parameter in the key metabolic parameters in the cell PK / mPBPK model is individually mapped to a fixed dimension using a fully connected layer.
[0019] Further, the Transformer model is stacked by 6 encoder layers, and a single encoder layer includes a multi-head attention layer, a liver injury-specific adjustment layer, a feed-forward neural network, a residual, and layer normalization;
[0020] In the input feature input encoder layer, first, the multi-head attention layer is passed through to capture the global dependencies in the input features. Then, residual and layer normalization are performed, such that the output of the multi-head self-attention layer is added to the original input and passed through layer normalization. Then, it is input into the liver injury-specific adjustment layer for correction. Next, the feed-forward neural network further processes the output of the liver injury-specific adjustment layer. Finally, residual connection and layer normalization are performed, such that the output of the feed-forward neural network is added to the original input and passed through layer normalization.
[0021] A hepatitis B virus antiviral drug pharmacokinetic prediction system based on the Transformer model, comprising:
[0022] Obtain the in vitro and in vivo pharmacokinetic data of three TFV prodrugs in mice, beagle dogs, and humans through experiments and literature research to form an experimental data set;
[0023] Build hepatocyte PK and in vivo mPBPK models and validate them. On this basis, develop liver injury mouse / human cell PK and mPBPK models;
[0024] Based on the obtained experimental data set, cell PK, and mPBPK models, construct a data set for Transformer model training;
[0025] Encode the features of the data in the data set for training, splice them to obtain feature vectors, and then splice them with position encoding to obtain input features;
[0026] Build a Transformer model based on the encoder-only architecture and input the input features into the Transformer model for training and validation;
[0027] Use the trained Transformer model to predict the concentration of TFV-DP in human liver cells.
[0028] A computer storage medium stores a readable program that, when run, can execute the above-mentioned hepatitis B virus antiviral drug pharmacokinetic prediction method based on the Transformer model.
[0029] An electronic device, comprising: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus;
[0030] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the above-mentioned hepatitis B virus antiviral drug pharmacokinetic prediction method based on the Transformer model.
[0031] A computer program product includes computer instructions that direct a computing device to perform the operations corresponding to the above-mentioned hepatitis B virus drug pharmacokinetic prediction method based on the Transformer model.
[0032] Advantages of the present invention:
[0033] 1. The prediction method of the present invention can dynamically simulate the metabolism of different patients (including groups with different liver functions, etc.) when using TFV prodrugs, accurately predict the pharmacokinetic process of drugs, and thus help drug R & D personnel screen and optimize drug combinations, and rationally design the dosage and usage instructions for combination therapy.
[0034] 2. The method of the present invention predicts the metabolism process of tenofovir (TFV) and its active metabolite TFV-DP in patients under different disease states. For hepatitis B patients with liver damage (such as complications like cirrhosis and liver fibrosis), the drug metabolism ability of the liver may be affected, which can assist clinicians in formulating more accurate personalized treatment plans, ensuring the best therapeutic effect of the drug in patients, and thus improving the treatment effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0036] Figure 1 is the flowchart of the hepatitis B virus drug pharmacokinetic prediction method based on the Transformer model of the present invention;
[0037] Figure 2 is the schematic diagram of the hepatocyte PK model of the present invention;
[0038] Figure 3 is the schematic diagram of the in vivo PBPK model of the present invention;
[0039] Figure 4 is the schematic diagram of the Transformer model including a liver damage-specific adjustment layer of the present invention;
[0040] Figure 5 is the schematic diagram of the multi-head attention layer of the present invention;
[0041] Figure 6 is the schematic diagram of the feed-forward neural network layer of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the protection scope of the present invention.
[0043] Example 1
[0044] As Figure 1 shown, a method for predicting the pharmacokinetics of anti-hepatitis B virus drugs based on the Transformer model includes the following steps:
[0045] S1. Obtain the in vitro and in vivo pharmacokinetic data of three tenofovir (TFV) prodrugs in mice, beagle dogs and humans through experiments and literature research to form an experimental data set; The three tenofovir (TFV) prodrugs selected are: tenofovir disoproxil fumarate (TDF), tenofovir alafenamide (TAF) and tenofovir ambufenamide (TMF).
[0046] The steps for obtaining the data set specifically include:
[0047] S11. Obtain the in vivo pharmacokinetic data of three tenofovir prodrugs in mice, beagle dogs and humans through experiments and literature research;
[0048] Obtain the plasma pharmacokinetic data of the three TFV prodrugs: Use C57BL / 6 mice (iv. 14.5 μmol / kg; ig. 60.7 μmol / kg; n = 6, male) and beagle dogs (iv. 16.3 μmol / kg; ig. 16.3 μmol / kg; n = 4, male). Collect blood samples from mice (two time points for each mouse) or dogs (through the forelimb vein) into tubes containing EDTA at 0.083, 0.25, 1, 4, 10 and 24 hours after dosing for LC-MS / MS analysis.
[0049] Study the concentrations of TFV and TFV-DP in the liver of mice by single-dose and oral administration. Sacrifice the mice 6 hours after oral administration of the TFV prodrug. Immediately harvest the liver samples into 50% acetonitrile and homogenize them separately. Tissue homogenate (W / V: 0.1 g·mL-1) is analyzed by LC-MS / MS.
[0050] Obtain the pharmacokinetic data of the TFV prodrug in mice with liver injury: To simulate hepatitis B with liver injury, HBV transgenic mice are intraperitoneally injected with 7% CCl4 (dissolved in soybean oil, V / V), 3 times a week for 4 consecutive weeks. Then, these model mice are treated with the TFV prodrug, and the same samples (plasma and liver) as those of C57BL / 6 mice are collected.
[0051] Collect the multi-dose pharmacokinetic characteristics of TFV prodrugs in beagle dogs through literature research, as well as the pharmacokinetic data of TDF and TAF (administered as single or multiple oral doses to healthy subjects) in humans.
[0052] S12. Obtain the in vitro pharmacokinetic-related data of TFV prodrugs in mice, beagle dogs, and human primary hepatocytes.
[0053] Obtain mouse, beagle dog, and human primary hepatocytes. All primary hepatocytes are seeded in 24-well plates at a density of 4×105 cells / well and incubated with 5 μM TMF / TAF / TDF for 2, 4, 8, 12, and 24 hours, respectively. After incubation for different times, the cells are collected with 70% methanol. The concentrations of TFV prodrugs, TFV, and TFV-DP in the cells are determined by LC-MS / MS and corrected for protein concentration.
[0054] S13. Determine the hydrolysis rate of TFV prodrugs;
[0055] Transfer 45 μL of plasma from different species (mice, dogs, and humans) and add 5 μL of each TFV prodrug solution (operate on ice). After vortex mixing, the drug-containing plasma is immediately incubated at 37°C. After incubation for 0, 2, 4, 8, 15, 30, 60, and 120 minutes, 150 μL of ice-cold acetonitrile containing internal standard is immediately added to terminate the reaction, and then LC-MS / MS analysis is performed.
[0056] S2. Build hepatocyte PK and in vivo mPBPK models and validate them; on this basis, develop liver-injured mouse / human cell PK and mPBPK models.
[0057] Use Monolix Suite software to establish hepatocyte PK models and mPBPK models (divided into compartments such as intestine, plasma, and liver), and describe the drug transmembrane transport, metabolism, and clearance processes through differential equations. Select the Stochastic Approximation-Expectation Maximization (SAEM) algorithm to perform parameter estimation. Optimize the parameters of the liver injury model based on Sobol sensitivity analysis to adapt to the metabolic characteristics under the state of liver fibrosis. The hepatocyte PK model is used to simulate the pharmacokinetic behaviors of TFV prodrugs (TDF, TAF, TMF), TFV, and TFV-DP in primary mouse, beagle dog, and human hepatocytes.
[0058] (1) Construct the hepatocyte PK model
[0059] This model was used to simulate the pharmacokinetic behaviors of TFV prodrugs (TDF, TAF, TMF), TFV, and TFV-DP in primary mouse, beagle dog, and human hepatocytes. TFV prodrugs have good liposolubility and mainly enter hepatocytes through transporters. In the cellular PK model, the transmembrane process of the prodrug was characterized as first-order kinetics, and the efflux of the prodrug was ignored due to its rapid metabolism. Due to the poor membrane permeability of TFV, it enters and exits hepatocytes in a membrane-limited manner. Two parameters, PS and KP, were used to describe this process. After entering the cell, the prodrug was rapidly metabolized to TFV by carboxylesterase 1 (CES1). Then TFV was phosphorylated by a series of phosphatases to generate TFV-DP. The processes of converting TFV to TFV-MP and converting TFV-MP to TFV-DP were represented by k1. A part of TFV-DP was converted to TFV by dephosphorylation, and another part bound to DNA and then was eliminated. These two steps were represented by k2 and k3 respectively. The structure of the in vitro PK model is as Figure 2 shown and was obtained from the following formula:
[0060]
[0061] -k1*A in_TFV +k2*A in_TFVdp (4)
[0062]
[0063] where the variables and parameters of the hepatocyte PK model are shown in Table 1 below:
[0064] Table 1 Variables and Parameters of the Hepatocyte PK Model
[0065]
[0066]
[0067] (2) Construct an in vivo physiologically based pharmacokinetic (mPBPK) model.
[0068] The organism was divided into four functional compartments: intestinal vasculature, plasma, liver, and other tissues. The model simulated the oral absorption process of tenofovir (TFV) prodrugs: after the prodrug entered the blood circulation through intestinal epithelial cells, its metabolic behaviors showed four aspects of characteristic differences due to structural differences: (1) intestinal absorption efficiency and bioavailability; (2) degradation rate of the drug in plasma; (3) rate of transmembrane transport to the liver; (4) metabolic rate of the prodrug converted to TFV in hepatocytes. Based on the above mechanisms, an mPBPK model that accurately characterized the in vivo metabolic dynamics of different TFV prodrugs (TDF, TAF, TMF) was finally established, as Figure 3 shown.
[0069] The prodrug absorption is calculated as follows:
[0070]
[0071]
[0072]
[0073] Among them, the mPBPK model variables and parameters are shown in Table 2 below:
[0074] Table 2 mPBPK model variables and parameters
[0075]
[0076]
[0077]
[0078] (3) Develop a liver injury mouse / human cell PK model and a physiological pharmacokinetic (mPBPK) model for liver injury mice and human cells.
[0079] Based on the developed cell PK and mPBPK models, the Sobol sensitivity analysis method is used to quantify the impact of liver-related parameters (such as metabolic rate, clearance rate, etc.) on the accumulation of tenofovir diphosphate (TFV-DP) in the liver. To ensure the reliability of the evaluation, a wide range of changes (0.1 to 10 times) is set for the key test parameters, and the model is optimized by calibrating the sensitive parameters to fit the experimental PK data of liver injury mice. To predict the accumulation pattern of TFV-DP in the human liver, the parameter adjustment range is further introduced into the human-derived mPBPK model to simulate the dynamic changes of drug metabolism under pathological conditions such as liver fibrosis.
[0080] S3. Based on the experimental data set obtained in S1 and the cell PK and mPBPK models established in S2, construct a data set for Transformer model training;
[0081] The data set integrates four types of features: drug concentration data (prodrug of TFV, TFV, TFV-DP), key metabolic parameters in the cell PK / mPBPK model, liver injury score (Child-Pugh score), and time series, and forms a multi-variable input sequence.
[0082] (1) Drug concentration data.
[0083] Concentration data of each species (mouse, beagle dog, human) and TFV prodrugs (TDF, TAF, TMF) and their metabolites (TFV, TFV-DP) at different time points (0, 2, 4, 8, 12, 24 hours), covering mouse, beagle dog and human experimental data.
[0084] (2) Extract key parameters from the cell PK / mPBPK model as the parameter dataset, including:
[0085] 2.1) Metabolism and transformation:
[0086] Hepatic membrane permeability (PSTFV): Describes the drug transmembrane transport efficiency and directly affects the intracellular drug concentration.
[0087] Partition coefficient (kpTFV): Reflects the equilibrium state of drug distribution inside and outside the cell.
[0088] Intracellular metabolic rate (k, k1, k2, k3): Completely depicts the metabolic chain of prodrug → TFV → TFV-MP → TFV-DP, as well as the clearance pathway.
[0089] Hepatic tissue metabolic rate (Kli, k1li, k2li): Adapts to the metabolic characteristics under liver injury conditions, such as decreased phosphorylation rate or increased clearance rate.
[0090] 2.2) Whole body pharmacokinetics:
[0091] Vascular / plasma metabolic rate constant (k): Describes the stability of the drug in the circulatory system.
[0092] TFV-DP clearance rate (k3li): Quantifies the elimination efficiency of the active metabolite and correlates with the degree of liver injury.
[0093] (3) Liver injury status score
[0094] The Child-Pugh score is used to quantify the severity of liver fibrosis or cirrhosis and serves as a dynamic correction factor. Scoring range: Grade A (5 - 6 points), Grade B (7 - 9 points), Grade C (10 - 15 points), reflecting the degree of liver function injury.
[0095] (4) Time series
[0096] Concatenate the data of each sample to form a multivariate time series. The input dataset is organized as:
[0097] X t =[C t ,P,S,t]
[0098] C tDrug concentrations of different species (mice, beagles, humans) in (TFV prodrug, TFV, TFV-DP); P is the key metabolic parameter in the cellular PK / mPBPK model; S is the liver injury score (Child-Pugh score); t is the time point value.
[0099] S4 encodes the features of the data in the dataset constructed by S3, splices them to obtain feature vectors, and then splices them with positional encoding to obtain input features;
[0100] (1) Encode the features of the data according to categorical features, numerical features, and time features;
[0101] Among them, the classification of the data is shown in Table 3 below:
[0102] Table 3 Data Classification
[0103]
[0104]
[0105] The information in Table 3 is specifically as follows:
[0106] 1) Represent the species and drug index types using One-Hot encoding:
[0107] Species encoding example: Mice are represented as [1, 0, 0]; Beagles are represented as [0, 1, 0]; Humans are represented as [0, 0, 1].
[0108] Drug index type categories: TDF (Tenofovir Disoproxil Fumarate), TAF (Tenofovir Alafenamide), TMF (Tenofovir Ambufenamide), TFV (Tenofovir), TFV-DP (Tenofovir Diphosphate). Encoding example:
[0109] TDF is represented as: [1, 0, 0, 0, 0]; TAF is represented as [0, 1, 0, 0, 0]; TMF is represented as [0, 0, 1, 0, 0];
[0110] TFV is represented as: [0, 0, 0, 1, 0]; TAF is represented as [0, 0, 0, 0, 1].
[0111] 2) Convert each grade of the liver injury score Child-Pugh score into a binary vector using One-Hot encoding;
[0112] Example: Grade A is represented as [1, 0, 0]; Grade B is represented as [0, 1, 0]; Grade C is represented as [0, 0, 1].
[0113] 3) Use Z-score normalization to process the drug concentration data and the key metabolic parameters in the cellular PK / mPBPK model, so that the mean of each feature is 0 and the standard deviation is 1, eliminating the scale differences of different features.
[0114]
[0115] Among them, X is the original data, μ is the mean, and σ is the standard deviation.
[0116] 4) Temporal feature: Retain the time value.
[0117] (2) First, map the key metabolic parameters in the cellular PK / mPBPK model to a fixed dimension, and then splice the species and drug index types, drug concentration data, the key metabolic parameters in the mapped cellular PK / mPBPK model, liver injury scores, and time series after encoding to obtain the feature vector F; the specific steps are as follows:
[0118] 1) For the key metabolic parameters in the cellular PK / mPBPK model, different metabolic parameters represent different biological processes. To retain the specificity of each parameter, a fully connected layer is used for each parameter separately to map it to a fixed dimension (32 dimensions). Use the activation function ReLu to capture the non-linear relationship between the input parameter and the target:
[0119] h i = ReLU(W i ·p i + b i )
[0120] Among them, p i is the normalized value of the i-th parameter, and W i and b i are learnable weights.
[0121] 2) Splice the above species and drug index types, drug concentration data, the key metabolic parameters in the mapped cellular PK / mPBPK model, liver injury scores, and time series after encoding to obtain the feature vector F.
[0122] 3) Divide the dataset after feature encoding into a training set, a test set, and a validation set. Among them, the ratio of the training set, the test set, and the validation set is 7:1.5:1.5.
[0123] (3) Positional encoding
[0124] Directly splice the feature vector F and add the positional encoding PE to obtain the input feature E input , which is used for training in the Transformer model.
[0125] Define positional encoding. For each position pos, a vector is generated, where the value of each dimension is determined by sine and cosine functions. Define the formula to calculate the positional encoding matrix:
[0126] For even dimensions (index 2i), calculate through the sine function:
[0127]
[0128] For odd dimensions (index 2i + 1), calculate through the cosine function:
[0129]
[0130] where pos is the position index (starting from 0), i is the dimension index in the positional encoding vector, and d is the feature dimension.
[0131] Concatenate the positional encoding PE and the feature vector F (formed by F i ) on the dimension. The dimension after concatenation will be the sum of the embedding dimension and the positional encoding dimension:
[0132]
[0133] S5. Build a Transformer model based on the encoder-only architecture, and input the input features into the Transformer model for training and validation;
[0134] As Figure 4 shown, the Transformer model is stacked by 6 encoder layers (Ne = 6). A single encoder layer includes a multi-head attention layer (8 heads), a liver injury-specific adjustment layer, a feed-forward neural network, a residual, and layer normalization;
[0135] The input features first pass through the multi-head attention layer to capture the global dependencies in the input features, and then go through a residual and layer normalization, which adds the output of the multi-head self-attention layer to the original input and normalizes it through layer normalization; then input into the liver injury-specific adjustment layer for correction; then the feed-forward neural network further processes the output of the liver injury-specific adjustment layer; finally, a residual connection and layer normalization are performed to add the output of the feed-forward neural network to the original input and normalize it through layer normalization.
[0136] (1) Multi-head self-attention layer
[0137] As Figure 5As shown, the multi-head self-attention layer can process multiple attention heads in parallel, capturing various different relationships in the input features. When there are h attention heads and the data matrix dimension is d, the data matrix processed by each attention head is d / h. Multiple "heads" calculate different attention scores respectively, and then these results are concatenated and followed by a linear transformation. In this way, the model can capture different types of relationships and dependency information in multiple different representation subspaces. The output of the multi-head attention layer is directly added to the input to form a residual connection. Subsequently, layer normalization is performed to ensure that the data has stable mean and variance when passing through each layer.
[0138] Query, Key, Value calculation:
[0139] For the input feature E entering the multi-head attention layer input perform a linear transformation to obtain the query matrix Q, key matrix K, and value matrix V:
[0140] Q = E input W Q
[0141] K = E input W K
[0142] V = E input W V
[0143] where W Q , W K , W V are the weight matrices of the linear transformation.
[0144] Split the multi-heads. If the number of heads in the multi-head self-attention layer is H and the data matrix dimension is d, the dimension d h of each head = d / H, and split Q, K, V into different subspaces according to the number of heads:
[0145]
[0146] where is the projection matrix of each head, and Q h , K h , V h are the query, key, and value matrices of the h-th head.
[0147] Calculate scaled dot-product attention. For each head h, calculate the attention score and the weight matrix.
[0148]
[0149] where Represents the similarity of all position pairs Used to scale the gradient and prevent the dot product from being too large, resulting in softmax saturation
[0150] Concatenation and projection of the multi-head output. After concatenating the outputs of all heads, they are combined into the final output through a linear transformation:
[0151] MultiHead(E input ) = Concat(Attention1, ……, Attention H )W O
[0152] where Concat(x) concatenates the outputs of H heads into a matrix of the original dimension d. W O Projects the concatenated result back to the original hidden space
[0153] Subsequently, residual connection and layer normalization (Add&Norm) are performed; the output of the multi-head self-attention layer is added to the original input and layer normalization is applied:
[0154] E MA_norm = LayerNorm(E input + MultiHead(E input )
[0155] where E MA_norm represents the output after passing through the multi-head self-attention layer and layer normalization; it is achieved through LayerNorm(x)
[0156] (2) Liver injury-specific adjustment layer
[0157] The model can adapt to different liver injury states, and the "directional correction" of liver injury to drug metabolism. The adjustment layer dynamically injects the liver injury score into the hidden representation through a linear transformation:
[0158] E adjusted = E MA_norml + W s ·S + b s
[0159] where E adjusted represents the output of the liver injury-specific adjustment layer, W s ·S projects the scalar S into a d-dimensional vector. b s is the bias term
[0160] (3) Feed-Forward Network
[0161] After passing through the liver injury-specific adjustment layer, a feed-forward neural network is set up to further process the information; as Figure 6 shown, the feed-forward neural network usually consists of two fully connected layers (one is the activation layer and the other is the output layer) and an activation function (ReLU). For the two fully connected layers, the first layer (W1) is regarded as extracting high-level features from the input features, and the second layer (W2) is regarded as combining the high-level features into the output. The feed-forward neural network is used to further process the output after the attention mechanism to increase the expressive power of the model.
[0162] FFN(E adjusted ) = ReLU(E adjusted ·W1 + b1)W2 + b2
[0163] where FFN(E adjusted ) represents the output of the feed-forward neural network layer, which is implemented through the activation function ReLU(x).
[0164] Subsequently, residual connection and layer normalization (Add&Norm) are performed. The output of the feed-forward neural network is added to the original input and passed through layer normalization:
[0165] F ffn_norm = LayerNorm(F adjusted + FFN(E adjusted ))
[0166] where F ffn_norm represents the output after passing through the feed-forward neural network and layer normalization, which is implemented through LayerNorm(x).
[0167] After the model is trained, model validation is performed;
[0168] The difference between the predicted concentration and the true concentration is measured, and its performance is evaluated on the validation set and its generalization ability is evaluated on the test set. Evaluation metrics:
[0169] 1) The mean squared error (MSE) is used as the loss function to measure the difference between the predicted concentration and the true concentration:
[0170]
[0171] where y pred is the predicted liver drug concentration of the model; y true is the experimentally measured true data; N is the number of samples.
[0172] 2) Its performance is evaluated on the validation set, and finally its generalization ability is evaluated on the test set. Evaluation metric: root mean squared error (RMSE):
[0173]
[0174] 3) Mean Absolute Error (MAE):
[0175]
[0176] S6. Use the trained Transformer model to predict the concentration of TFV-DP in human liver cells.
[0177] Based on a similar inventive concept, an embodiment of the present invention further provides a computer storage medium storing a readable program which, when running, can execute the above-mentioned anti-hepatitis B virus pharmacokinetic prediction method based on the Transformer model.
[0178] Based on a similar inventive concept, an embodiment of the present invention provides an electronic device, including: a processor, a memory, a communication interface and a communication bus, and the processor, the memory and the communication interface complete communication with each other through the communication bus;
[0179] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations corresponding to the above-mentioned anti-hepatitis B virus pharmacokinetic prediction method based on the Transformer model.
[0180] Based on a similar inventive concept, an embodiment of the present invention further provides a computer program product including computer instructions, and the computer instructions instruct a computing device to execute the operations corresponding to the above-mentioned anti-hepatitis B virus pharmacokinetic prediction method based on the Transformer model.
[0181] Embodiment 2
[0182] Based on the anti-hepatitis B virus pharmacokinetic prediction method based on the Transformer model proposed in Embodiment 1, in this embodiment, an anti-hepatitis B virus pharmacokinetic prediction system based on the Transformer model is proposed, including:
[0183] Obtain the in vitro and in vivo pharmacokinetic related data of three TFV prodrugs in mice, beagle dogs and humans through experiments and literature research to form an experimental data set;
[0184] Build hepatocyte PK and in vivo mPBPK models and verify them, and on this basis, develop liver injury mouse / human cell PK and mPBPK models;
[0185] Based on the obtained experimental data set and cell PK and mPBPK models, construct a data set for Transformer model training;
[0186] Encode the features of the data in the training dataset and splice them to obtain feature vectors, and then splice them with position encoding to obtain input features;
[0187] Build a Transformer model based on the encoder-only architecture, and input the input features into the Transformer model for training and validation;
[0188] Use the trained Transformer model to predict the concentration of TFV-DP in human liver cells.
[0189] Example 3
[0190] In this example, it is used to introduce the specific application of the anti-hepatitis B virus drug pharmacokinetic prediction method based on the Transformer model.
[0191] (1) Screening of candidate drugs in the drug R & D stage
[0192] Application scenario: In the research and development of anti-hepatitis B virus drugs (such as TFV prodrugs), it is necessary to quickly evaluate the metabolic stability and toxicity risks of different candidate drugs (TDF, TAF, TMF) in patients with liver injury, and narrow the experimental scope.
[0193] Specific implementation:
[0194] Step 1: Input the in vitro hepatocyte metabolism data of candidate drugs (such as TFV prodrug hydrolysis rate, TFV-DP generation rate) and cell PK / mPBPK model parameters (such as k1, k3li).
[0195] Step 2: Predict the TFV-DP accumulation curve of candidate drugs in the liver under different liver injury scores (Child-Pugh A / B / C grades) through the Transformer model.
[0196] Step 3: Compare the predicted concentrations of each candidate drug with the toxicity threshold (such as >100 μM may cause liver injury), and screen out drugs with high metabolic efficiency and low toxicity risk (such as TAF accumulates more stably in patients with liver fibrosis).
[0197] Technical advantage: Avoid the time-consuming and consumable traditional animal experiments, and shorten the drug R & D cycle by 30% - 50%.
[0198] (2) Virtual bioequivalence study for drug regulation and declaration
[0199] Application scenario: Support the accelerated approval of drugs for liver injury subgroups, and reduce the ethics and resource consumption of real-world trials.
[0200] Specific implementation:
[0201] Step 1: Train a model based on the data of healthy subjects and verify its extrapolation ability in the virtual population with liver injury (e.g., the prediction error RMSE < 15%).
[0202] Step 2: Input the physiological parameters of liver injury patients (e.g., a 20% decrease in plasma protein binding rate), simulate the bioequivalence curve, and compare it with the reference drug for consistency.
[0203] Step 3: Generate a virtual bioequivalence report and submit it to regulatory agencies (such as FDA, NMPA) as supplementary data.
[0204] Technical advantage: Accelerate the market launch process of drugs for the liver injury subgroup.
[0205] The method of the present invention can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CDROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and to be downloaded through a network and stored in a local recording medium, so that the method described herein can be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.
[0206] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.
Claims
1. A method for predicting the pharmacokinetics of anti - hepatitis B virus drugs based on a Transformer model for non - diagnostic purposes, characterized in that, It includes the following steps: Obtain the in vitro and in vivo pharmacokinetic-related data of three TFV prodrugs in mice, beagle dogs, and humans through experiments and literature research to form an experimental dataset; Build hepatocyte PK and in vivo mPBPK models and validate them. On this basis, develop liver injury mouse / human cell PK and mPBPK models; Based on the obtained experimental dataset, cell PK, and mPBPK models, construct a dataset for Transformer model training; Perform feature encoding on the data in the dataset for training, and splice them to obtain feature vectors, and then splice them with position encoding to obtain input features; Construct a Transformer model based on the encoder-only architecture, and input the input features into the Transformer model for training and validation; Use the trained Transformer model to predict the concentration of TFV-DP in human liver cells.
2. The method for predicting the pharmacokinetics of anti-hepatitis B virus drugs based on the Transformer model for non-diagnostic purposes according to claim 1, wherein The three TFV prodrugs are: tenofovir disoproxil fumarate, tenofovir alafenamide, and tenofovir ambutolamide.
3. The method for predicting the pharmacokinetics of anti-hepatitis B virus drugs based on the Transformer model for non-diagnostic purposes according to claim 1, wherein, Four types of features are integrated in the dataset for Transformer model training: drug concentration data, key metabolic parameters in the cell PK / mPBPK model, liver injury score, and time series, and form a multi-variable input sequence.
4. The method for predicting the pharmacokinetics of anti-hepatitis B virus drugs based on the Transformer model for non-diagnostic purposes according to claim 3, wherein The steps of feature encoding include: Classify the data according to categorical features, numerical features, and time features; According to the classification results, use One-Hot encoding for species, drug indicators, and liver injury scores; use Z-score normalization for drug concentration data and key metabolic parameters in the cell PK / mPBPK model; directly retain the time values for time series.
5. The method for predicting the pharmacokinetics of anti-hepatitis B virus drugs based on the Transformer model for non-diagnostic purposes according to claim 4, wherein, Before feature encoding, each parameter in the key metabolic parameters in the cell PK / mPBPK model is individually mapped to a fixed dimension using a fully connected layer.
6. The method for predicting the pharmacokinetics of anti-hepatitis B virus drugs based on the Transformer model for non-diagnostic purposes according to claim 1, wherein The Transformer model is stacked by 6 encoder layers. A single encoder layer includes a multi-head attention layer, a liver injury-specific adjustment layer, a feed-forward neural network, a residual, and layer normalization; The input features are input into the encoder layer. First, the multi-head attention layer is used to capture the global dependencies in the input features, and then residual and layer normalization are performed, so that the output of the multi-head self-attention layer is added to the original input and passed through layer normalization; then it is input into the liver injury-specific adjustment layer for correction; then the feed-forward neural network further processes the output of the liver injury-specific adjustment layer; finally, residual connection and layer normalization are performed, so that the output of the feed-forward neural network is added to the original input and passed through layer normalization.
7. A pharmacokinetic prediction system for anti-hepatitis B virus drugs based on the Transformer model, characterized in that, It includes: Obtain the in vitro and in vivo pharmacokinetic-related data of three TFV prodrugs in mice, beagle dogs, and humans through experiments and literature research to form an experimental dataset; Build hepatocyte PK and in vivo mPBPK models and validate them. On this basis, develop liver injury mouse / human cell PK and mPBPK models; Based on the obtained experimental dataset, cell PK, and mPBPK models, construct a dataset for Transformer model training; Feature-encode the data in the dataset for training, and splice them to obtain a feature vector, and then splice it with the position encoding to obtain the input feature; Construct a Transformer model based on the encoder-only architecture, and input the input feature into the Transformer model for training and validation; Use the trained Transformer model to predict the concentration of TFV-DP in human liver cells.
8. A computer storage medium stores a readable program, characterized in that, When the program runs, it can execute the hepatitis B virus drug pharmacokinetic prediction method based on the Transformer model according to any one of claims 1-6.
9. An electronic device, characterized in that, Including: A processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations corresponding to the hepatitis B virus drug pharmacokinetic prediction method based on the Transformer model according to any one of claims 1-6.
10. A computer program product, comprising computer instructions, characterized in that, The computer instruction instructs the computing device to execute the operations corresponding to the hepatitis B virus drug pharmacokinetic prediction method based on the Transformer model according to any one of claims 1-6.