CYP2C9 enzyme drug metabolism function prediction method based on transfer learning and application
Through transfer learning and data augmentation technology, combined with residual network and channel attention mechanism, a convolutional neural network model was constructed, which solved the problem of drug metabolism function prediction of multi-locus variants of CYP2C9 gene, and achieved high-precision metabolic function classification to assist in clinical use.
Patent Information
- Application Number
- CN202510327616.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art is difficult to effectively use transfer learning methods to predict drug metabolism function for multiple loci and unintegrated variants of CYP2C9 gene, resulting in the problem of small training sample size and low prediction accuracy.
A transfer learning-based method is adopted, combining data augmentation, residual network and channel attention mechanism to generate a simulated double type data set, and the source domain model weight is transferred to the target domain through the transfer learning algorithm, a convolutional neural network model is constructed, and the characteristics of multi-locus variation information of CYP2C9 gene are captured, and drug metabolism function prediction is performed.
The accuracy of prediction of multi-site variation of CYP2C9 gene is improved, and metabolic typing prediction of unintegrated variants is achieved, which assists in clinical precise medication and reduces the incidence of adverse drug reactions.
Smart Images

Figure CN120279985A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of functional genomics, and particularly relates to a method and application for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning. Background Art
[0002] Drug injuries caused by the irrational use of drugs have become an important public health issue of global concern. According to a meta-analysis study on preventable drug injuries in 2020: in the field of medical health, approximately 1 in every 30 patients is affected by preventable drug injuries, and more than 25% of the patients suffer serious injuries or even life-threatening situations. How to maximize the therapeutic effect of drugs in the body while reducing the incidence of their adverse reactions is a major challenge faced by clinical individualized drug therapy. The effect of individualized drug therapy is related to multiple factors, such as environmental factors (co-medication, diet, smoking, etc.), physical conditions (gender, age, concomitant diseases, etc.) and genetic factors (gene mutations, etc.). Among them, genetic factors can explain 20-30% of the different drug treatment effects. Therefore, predicting the related drug metabolism function based on gene mutation information helps to provide a scientific basis for clinical drug use, prevent the occurrence of drug adverse reactions, and achieve personalized and precise drug use.
[0003] Cytochrome P450 (CYP450) is the main drug-metabolizing enzyme in the human body, responsible for catalyzing the metabolism of a variety of endogenous substrates, exogenous compounds and other drugs. Among the genes with pharmacogenetic activity released by the Clinical Pharmacogenetics Implementation Consortium (CPIC), multiple genes belong to the CYP450 subfamily, such as CYP2D6, CYP2C9, CYP2C19, etc. Among them, the CYP2C9 enzyme encoded by the CYP2C9 gene is an important isoenzyme in the CYP2C subfamily, accounting for about 20% of the P450 protein content in human liver microsomes, and is involved in the metabolism of about 17% of common clinical drugs, such as the anticoagulant warfarin, the antidiabetic drug tolbutamide, and the non-steroidal anti-inflammatory drug diclofenac, etc. There is sufficient research showing that the genetic polymorphism of its encoding gene CYP2C9 can significantly change the drug metabolism function of the CYP2C9 enzyme, thus affecting the therapeutic effect of drugs and possibly leading to the occurrence of drug adverse reactions, which is more obvious for drugs with a narrow therapeutic window such as the anticoagulant warfarin.
[0004] Internationally, according to the different drug metabolism rates of CYP2C9 enzyme in different individuals, it is generally divided into three metabolic phenotypes: normal metabolizer, intermediate metabolizer, and poor metabolizer. For example, when poor metabolizers of CYP2C9 enzyme, especially homozygous carriers of CYP2C9*3, use drugs related to its metabolic substrates (such as warfarin, glipizide), serious drug adverse reactions may occur. On the contrary, when using drugs related to its metabolic precursors (such as losartan, cyclophosphamide), drug efficacy may be insufficient, leading to treatment failure. To reduce the occurrence of clinical drug adverse reactions, as of now, the PharmVar website has included a total of 75 star alleles of CYP2C9, and most of its variations are located in the exon region. According to the different functional activities of CYP2C9 enzyme encoded by star alleles in in vitro and in vivo experiments, they are divided into four types: normal function, decreased function, non-functional, and unintegrated. Among them, unintegrated means that multiple-site variations of this type of CYP2C9 gene have been found, but its drug metabolism functional activity is not yet clear.
[0005] Therefore, before clinical medication, determining the drug metabolism phenotype of CYP2C9 enzyme in patients based on CYP2C9 gene variation information can effectively guide doctors to adjust the doses of CYP2C9 enzyme substrate drugs or precursor drugs according to the metabolizer type of patients, improve drug efficacy and reduce the incidence of drug adverse reactions, and achieve individualized precision medication.
[0006] Currently, the recognized method for converting CYP2C9 gene variation information into CYP2C9 enzyme drug metabolism types is to use the AS (Activity Score) scoring method. Its basic principle is to detect the allelic variations carried by CYP2C9 through experiments, then assign corresponding scores according to the metabolic activities of each allelic variation, and finally calculate the overall AS score. This method maps the known single or double-site variation information of the CYP2C9 gene to the corresponding metabolic phenotypes in a way of quantifying metabolic functions. However, allelic variations that have not yet determined the drug metabolism function of CYP2C9 enzyme (such as unintegrated variations like *7, *10, *17, *18, *19, *20, *21, *22, *27, *32, *34, etc., accounting for 46.7% of the total number of CYP2C9 alleles included in the PharmVar database) cannot use this method. In addition, there is currently a lack of an effective comprehensive rating method for drug metabolism based on multiple-site variations or even the whole gene sequence of the CYP2C9 gene.
[0007] In the field of functional genomics, many scholars have used machine learning methods to study how gene variation information maps to phenotypes, such as MutComput, ECNet, VarCoPP, etc. However, these methods only predict the pathogenicity of single or two variant sites and are not applicable to the complex multiple variant sites in CYP2C9 gene variation information and the unintegrated variant sites in the PharmVar database. The training data of the CYP2C9 enzyme metabolism prediction model has the characteristic of a small number of sample sets. If general data sets are used for conventional training, it is difficult to achieve good training effects and may affect the final prediction accuracy. Summary of the Invention
[0008] Aiming at the deficiencies of the prior art, the present invention provides a method and application for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning, which solves the problems of small training sample size and difficult multi-site mutation prediction in the prediction of related drug metabolism functions according to CYP2C9 gene variation information, combines data augmentation, residual network and channel attention mechanism to capture the characteristics of multi-site variation information of CYP2C9 gene, and improves the prediction accuracy. It breaks through the limitation of traditional methods that only perform manual interpretation of drug metabolism types based on known single or double-site mutations of CYP2C9 gene, and performs artificial intelligence interpretation and prediction of drug metabolism types for mutations at multiple sites or even the entire gene sequence of CYP2C9 gene, especially realizes the metabolic typing of unintegrated variants, and assists in guiding precise clinical medication.
[0009] To achieve the above object, the technical solution adopted by the present invention is:
[0010] A method and application for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning, comprising the following steps:
[0011] Step 1: Obtain CYP2C9 gene variation information and generate a source domain data set and a target domain data set through preprocessing;
[0012] The source domain data set is generated in the following way: Randomly select CYP2C9 star alleles with known metabolic functions, construct a simulated diplotype for pre-training, and mutate the variant sites related to non-star alleles based on the known population variation frequency to generate a source domain data set of the simulated diplotype after data augmentation;
[0013] Step 2: Construct two convolutional neural network models, including:
[0014] Five-class convolutional neural network model: Used for training the source domain data set, including a residual convolution module, an attention mechanism module and a CYP2C9 enzyme drug metabolism function prediction module;
[0015] Binary classification convolutional neural network model: It is used for training on the target domain dataset. Its structure is the same as that of the five-classification convolutional neural network model, and the output layer is adjusted to binary classification;
[0016] Step 3: Use the source domain dataset to train the five-classification convolutional neural network model until the model converges to obtain the initial weights;
[0017] Step 4: Through the transfer learning algorithm, transfer the weights of the trained five-classification convolutional neural network model to the binary classification convolutional neural network model. Freeze the weights of the residual convolution module and the attention mechanism module, and only update the weights of the CYP2C9 enzyme drug metabolism function prediction module with a small learning rate to obtain the transferred convolutional neural network model;
[0018] Step 5: Use the target domain dataset to train the transferred convolutional neural network model, adjust the model parameters to adapt to the target domain dataset until the model converges to obtain the trained target domain convolutional neural network model;
[0019] Step 6: Input the single-site, double-site, multi-site variant information or the full gene sequence of the CYP2C9 gene with unknown metabolic function into the target domain convolutional neural network model, and output the metabolic function classification result.
[0020] Furthermore, in the step 1, the target domain dataset is generated in the following way: Obtain the CYP2C9 gene variant information from databases, literature or experiments. Use the bioinformatics annotation software and database comparison annotation method, and adopt the independent one-hot encoding method to convert the haplotype variant information into a data matrix with the shape of M×P×F, where M is the number of target domain data samples, P is the gene sequence length, and F is the number of data features.
[0021] Furthermore, in the step 1, the source domain dataset is generated in the following way: Randomly select a pair of CYP2C9 star alleles with normal metabolic function, decreased metabolic function or no metabolic function, and construct haplotypes with variants related to the star alleles to create simulated diploids for pre-training; then according to the population-level alternate allele frequencies published in the GnomAD database, perform SNV and INDEL alternate allele sampling on the variant sites not related to any star alleles, select sites in a uniform distribution manner and assign corresponding mutation types, and finally form a source domain dataset containing N simulated diploids.
[0022] Furthermore, the convolutional neural network model includes two to four cascaded residual convolution modules. Each residual convolution module provides two mapping methods: residual mapping and identity mapping. The residual mapping includes a convolutional layer, a batch normalization layer and a ReLU activation layer. The identity mapping adds the input to the result of the residual mapping to avoid gradient disappearance.
[0023] Furthermore, the implementation process of the attention mechanism module includes: performing global average pooling operation on the input CYP2C9 feature map to compress the spatial information of each channel into a scalar; using one-dimensional convolution operation to calculate the attention weights between channels; multiplying the calculated attention weights by the original CYP2C9 feature map to achieve weighted adjustment of channel features.
[0024] Furthermore, the implementation process of the CYP2C9 enzyme drug metabolism function prediction module includes: calculating and outputting the non-functional score and normal functional score of the haplotype metabolism function through two fully connected layers; setting the critical value of the haplotype metabolism function based on maximizing sensitivity and specificity, and then converting the score into three metabolic function classification results: normal metabolizer, intermediate metabolizer, and poor metabolizer.
[0025] Furthermore, the range of the critical value is from 0 to 1. Preferably, the critical values are set as follows: the critical value of the non-functional score is 0.7720637, and the critical value of the normal functional score is 0.3042464.
[0026] Furthermore, the training strategy of the transfer learning algorithm includes: updating the weights of the model with an initial learning rate of 0.00001 during the training process; automatically adjusting the learning rate using the ReduceLROnPlateau function. When the validation set loss has not improved for 5 to 10 consecutive epochs, the learning rate is automatically reduced by one order of magnitude, and overfitting of the model is avoided through an early stopping mechanism.
[0027] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning are implemented.
[0028] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, the steps of the method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning are implemented.
[0029] The present invention also provides an application of the method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning. The method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning is applied to the metabolic function rating of drugs metabolized by CYP2C9 enzyme to assist in guiding precise clinical medication.
[0030] The beneficial effects of the above solution are as follows: In the present invention, a source domain dataset is obtained by performing site-directed mutagenesis on CYP2C9 allele variations with the mutation frequency of the site in the population as the probability, so as to solve the problem of insufficient sample quantity during training. Subsequently, the transfer learning method, ResNet network, and ECA attention mechanism are combined to capture the characteristics of multiple-site variations of the CYP2C9 gene, and the model performance is further optimized through parameter fine-tuning, realizing the prediction of the metabolic function type of the information on multiple-site variations of the CYP2C9 gene. The present invention not only has high classification accuracy, but also can mine deeper features, realize fast and accurate classification, and achieve precise prediction of the drug metabolism function of unknown mutation combinations. Brief Description of the Drawings
[0031] Figure 1 It is a flowchart of the method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning in the present invention;
[0032] Figure 2 It is a schematic diagram for generating the source domain dataset in the present invention;
[0033] Figure 3 It is a framework diagram of the convolutional neural network in the present invention;
[0034] Figure 4 It is a flowchart of the residual block in the present invention;
[0035] Figure 5 It is a flowchart of the attention mechanism module in the present invention;
[0036] Figure 6 It is a graph showing the change of the accuracy of the model with the number of epochs in the source domain dataset;
[0037] Figure 7 It is a graph showing the change of the loss of the model with the number of epochs in the source domain dataset;
[0038] Figure 8 It is a graph showing the change of the accuracy of the model with the number of epochs in the target domain dataset;
[0039] Figure 9 It is a graph showing the change of the loss of the model with the number of epochs in the target domain dataset. Detailed Embodiment
[0040] The present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0041] It should be noted that unless otherwise specified, all technical and scientific terms used in this embodiment have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.
[0042] In the following examples, unless otherwise specified, the reagents or instruments used are commercially available and can be obtained through ordinary commercial channels; the experimental methods used are conventional methods in the art, and those skilled in the art can understand how to conduct the experiments and obtain the corresponding results according to the descriptions in the examples.
[0043] Example 1
[0044] As Figure 1 shown, a method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning includes the following steps:
[0045] Step 1: Obtain CYP2C9 gene variant information and generate a source domain dataset and a target domain dataset through preprocessing; the source domain dataset is generated in the following way: randomly select CYP2C9 star alleles with known metabolic functions, construct a simulated diplotype for pre-training, and mutate the variant sites related to non-star alleles based on the known population variant frequencies to generate a source domain dataset of the simulated diplotype with data augmentation;
[0046] Step 2: Construct two convolutional neural network models, including:
[0047] Five-class convolutional neural network model: used for training the source domain dataset, including a residual convolutional module, an attention mechanism module, and a CYP2C9 enzyme drug metabolism function prediction module; two-class convolutional neural network model: used for training the target domain dataset, with the same structure as the five-class convolutional neural network model, and the output layer is adjusted to two-class;
[0048] Step 3: Use the source domain dataset to train the five-class convolutional neural network model until the model converges to obtain the initial weights;
[0049] Step 4: Through the transfer learning algorithm, transfer the weights of the trained five-class convolutional neural network model to the two-class convolutional neural network model, freeze the weights of the residual convolutional module and the attention mechanism module, and only update the weights of the CYP2C9 enzyme drug metabolism function prediction module with a small learning rate to obtain the transferred convolutional neural network model;
[0050] Step 5: Use the target domain dataset to train the transferred convolutional neural network model, adjust the model parameters to adapt to the target domain dataset until the model converges to obtain the trained target domain convolutional neural network model;
[0051] Step 6: Input the single-site, double-site, multi-site variant information or the full gene sequence of the CYP2C9 gene with unknown metabolic function into the target domain convolutional neural network model, and output the metabolic function classification result.
[0052] The following is a detailed description of each step:
[0053] Step 1: Obtain CYP2C9 gene variant information and perform preprocessing
[0054] The CYP2C9 gene variant information can be obtained from various sources, such as using the first-generation sequencing method, the second-generation sequencing method, the third-generation sequencing method on the peripheral blood of volunteers, or using the CYP2C9 allele variant information included in the PharmVar database but without verifying its metabolic function data. Specifically as follows:
[0055] Method 1: Obtain CYP2C9 gene variant information using the first-generation sequencing method
[0056] Collect whole blood samples from seven volunteers (denoted as I-1, I-2, I-3, I-4, I-5, I-6, I-7). All research subjects included in the experiment have signed informed consent forms. Take 200 mL of EDTA-anticoagulated whole blood samples and extract the genome. Set up 11 PCR reactions to amplify the promoter sequence and 9 exon sequences of the CYP2C9 gene of the volunteers respectively. After the PCR products are qualified by agarose gel electrophoresis detection, they are handed over to Sangon Biotech Co., Ltd. for Sanger sequencing. After obtaining the sequencing data of the volunteers, compare and analyze them with the reference sequence (NG_008385.1) of CYP2C9 downloaded from the PharmVar database to obtain CYP2C9 gene variant information.
[0057] Method 2: Obtain CYP2C9 gene variant information using the second-generation sequencing method
[0058] Use the second-generation sequencing method on the genomes of seven volunteers (denoted as II-1, II-2, II-3, II-4, II-5, II-6, II-7) to obtain the original FastQ files. Subsequently, use the sequence alignment software BWA to align the base sequences in the FastQ files to the hg19 (GRCh37) human reference genome to obtain BAM files. Then use the GATK software to sort the BAM files according to the genome coordinates and select the gene variant information, and finally obtain the CYP2C9 gene variant information of the volunteers.
[0059] Method 3: Obtain CYP2C9 gene variant information using the third-generation sequencing method
[0060] First, select one volunteer (denoted as III-1), collect his genome sample, and send the sample to a professional sequencing service company, and perform high-throughput sequencing on it using the third-generation sequencing method to obtain the volunteer genome variant data, and finally screen out the variant information related to the CYP2C9 gene.
[0061] Method 4: Collect CYP2C9 allele variant information in the PharmVar database
[0062] The CYP2C9 alleles obtained from the literature are included in the PharmVar database and classified according to their reported metabolic functions as normal metabolic function, decreased metabolic function, no metabolic function, uncertain metabolic function, and unknown metabolic function. Data on CYP2C9 allele variations without verified metabolic functions are collected.
[0063] Then, the CYP2C9 gene variation information obtained by the above method and the CYP2C9 allele variation information with normal, decreased, and no metabolic functions downloaded from the PharmVar database are preprocessed to enhance data features.
[0064] The basic process of preprocessing CYP2C9 gene variation information is as follows:
[0065] First, since the CYP2C9 gene variation information we obtained is limited to the sites with base changes relative to the reference sequence in this gene, in order to reconstruct the CYP2C9 gene sequence of the sample, we need to fill in the normal base sequences between these known variation sites, which includes the entire sequence of the promoter, the sequences of nine exons, and the sequences of 400 base pairs on each side of their respective left and right sides, so as to obtain the CYP2C9 gene sequence in the sample, with a total length of 16384.
[0066] Second, use ANNOVAR and VEP annotation software to compare the variation information with multiple databases to obtain nine data features of the CYP2C9 gene variation sites, which are:
[0067] 1. Whether the variation site is located in the coding region; 2. Whether the variation site is a rare variation in the population, 3. Whether the variation site is harmful; 4. Whether the variation site belongs to an Indel variation; 5. Whether the variation site is located at a methylation site; 6. Whether the variation site is located at a DNase sensitive site; 7. Whether the variation site is located at a transcription factor binding site; 8. Whether the variation site is affected; 9. Whether the variation site is highly affected.
[0068] Finally, use one-hot encoding to convert the base information and annotation information of CYP2C9 into a data matrix. The matrix information for model prediction is obtained after preprocessing the obtained CYP2C9 gene multi-site variation information, and the dataset downloaded from the PharmVar database will be used as the target domain dataset for model training and divided into a target domain training set and a target domain validation set according to a ratio of 6:4.
[0069] The target domain dataset is generated as follows: Obtain CYP2C9 gene variation information from databases, literature, or experiments. Using the method of comparing and annotating with a bioinformatics annotation software and a database, and adopting the independent hot encoding method, convert the haplotype variation information into a data matrix with the shape of M×P×F, where M is the number of target domain data samples, P is the gene sequence length, and F is the number of data features.
[0070] Preferably, in this embodiment, 75 CYP2C9 gene haplotypes with multi-locus variation information are obtained from the PharmVar database, among which 40 haplotypes have been verified by biological experiments, and 35 haplotypes have not been verified by biological experiments. A CYP2C9 haplotype includes variation information at 2 to 30 different sites, which are distributed in series on the CYP2C9 gene sequence. Using the method of comparing and annotating with a bioinformatics annotation software and a database, and adopting the independent hot encoding method, convert the single sample of haplotype variation information into a data matrix with the shape of 1×P×F, specifically 1×8176×13; Use this method to obtain small-sample real data for the 40 verified haplotypes, and use it as the target domain dataset. As a preference, the target domain dataset uses the homozygous diplotype of the biologically verified CYP2C9 star alleles.
[0071] In this embodiment, the databases used for annotation include refGene, wgEncodeAwgDnaseMasterSites, wgEncodeRegTfbsClusteredV3, tfbsConsSites, wgEncodeHaiMethyl450Gm12878SitesRep1, wgEncodeHaibMethylRrbsGm12878HaibSitesRep1, gnomad211_genome, dbnsfp33a.
[0072] In addition, since the number of CYP2C9 star alleles is small, there are only 55 functionally verified CYP2C9 star alleles, which have the characteristic of being difficult to extract important data features. Therefore, we adopt the method of combining the star alleles with known drug metabolism functions in pairs to form a diplotype to obtain a large-sample simulated source domain dataset. The process of generating the source domain dataset is as Figure 2 shown, and the specific method is:
[0073] Create simulated diploids for pre-training by randomly selecting a pair of CYP2C9 star alleles with known metabolic functions (normal metabolic function, decreased metabolic function, or no metabolic function), and constructing haplotypes with variants related to the star alleles. To introduce additional diversity into the training data, we sample alternate alleles (SNVs and INDELs) at variant sites not related to any star allele, select sites in a uniform distribution and assign corresponding mutation types, and the probability of alternate alleles occurring is equal to the population-level alternate allele frequency published in the GnomAD database. Finally, a source domain dataset containing N simulated diploids is formed. In this process, rare harmful variants that do not currently exist in any star allele may be added. Preferably, 40 validated haplotypes are used as the target domain dataset, and other variant expansions are added in the above manner to obtain a source domain dataset with the shape of N×P×F. A total of 30,000 genotypes are selected for each AS (0, 0.5, 1, 1.5, 2), and a total of 21,000 simulated samples are used as the training set, and another 9,000 genotypes are used as the test set.
[0074] Step 2: Construct a five-class convolutional neural network model for training on the source domain dataset, and a two-class convolutional neural network for training on the target domain dataset.
[0075] Taking the two-class convolutional neural network as an example, see Figure 3 , the convolutional neural network used for training is divided into three parts: a residual convolutional module, an attention mechanism module, and a CYP2C9 enzyme drug metabolism function prediction module, as shown in Table 1. Table 1 CYP2C9 Enzyme Drug Metabolism Function Prediction Model Framework
[0076] The CYP2C9 enzyme drug metabolism function prediction model first processes the input batch sequence data (with a size of batch×8192×13) through an initial convolutional layer, and then sequentially inputs it into two to four cascaded residual convolutional modules to gradually extract the features of the CYP2C9 batch sequence data. In this embodiment, three cascaded residual convolutional modules are set, and each residual convolutional module contains two residual blocks. The residual blocks achieve feature extraction through two mapping methods: residual mapping and identity mapping. Among them, the residual mapping includes a one-dimensional convolutional layer (the number of convolutional kernels is determined according to the position, the window length is 30, and uniform padding is used), a batch normalization layer, and a ReLU activation layer; the identity mapping is achieved through a bypass connection, providing a "shortcut" for the model, that is, directly passing the input data of the residual convolutional module to the output end and adding it to the result after being processed by the residual mapping. This design enables the model to learn small adjustments of the data rather than complete transformations, thus effectively alleviating the gradient vanishing problem of deep networks. The specific process of the residual block is as Figure 4 shown.
[0077] After the feature extraction of three residual convolutional modules, the CYP2C9 batch sequence data will be converted into a feature map with a size of batch×1024×256. Subsequently, this feature map is input into the attention mechanism module to further optimize the feature values. The specific process of the attention mechanism is as Figure 5 shown. The attention mechanism module includes a global average pooling layer, a one-dimensional convolutional layer, and a Sigmoid activation layer. Its specific process is as follows: First, perform a global average pooling operation on the input CYP2C9 feature map to compress the spatial information of each channel into a scalar; then calculate the attention weights between channels through a one-dimensional convolutional operation, and the size of the convolutional kernel is determined adaptively, so as to maintain good prediction performance while reducing the number of parameters; finally, multiply the calculated attention weights by the original CYP2C9 feature map to achieve weighted adjustment of the channel features. The output size of the attention mechanism module is the same as the input size.
[0078] Finally, the weighted CYP2C9 feature map is input into the CYP2C9 enzyme drug metabolism function prediction module to complete the classification task of the sequence data. This module includes two fully connected layers: the first layer contains 128 neurons, and the second layer contains 2 neurons for outputting two function scores. These two scores respectively represent the probabilities of the haplotype metabolism function being "non-functional" and "normal function". By maximizing the sensitivity and specificity, the critical values of the corresponding non-functional score and normal function score are determined to be 0.7720637 and 0.3042464 respectively. Based on these two scores, the model can classify the metabolism function into three categories: normal metabolism type, intermediate metabolism type, and poor metabolism type.
[0079] Step 3: Input the source domain dataset into the five-class convolutional neural network model constructed in Step 2 for training until the model converges to obtain the trained model weights.
[0080] When training the five-class convolutional neural network model on the source domain dataset, set the batch size to 256 (i.e., train 256 double-type data at a time), and its data tensor shape is [256, 16384, 13]; the learning rate is 0.0001 as the initial learning rate and use the ReduceLROnPlateau method to dynamically adjust the learning rate. When the validation set loss has not improved for 5 to 10 consecutive epochs, the learning rate is reduced by one order of magnitude; preferably, set the validation set loss not to improve for 5 consecutive epochs, and the learning rate is automatically reduced by one order of magnitude; select the multi-class cross-entropy loss function as the loss function; use the Adam method as the optimizer. In addition, use an early stopping mechanism during training to save computer resources and avoid overfitting of the model.
[0081] Step 4: Use the transfer learning algorithm to transfer the trained model weights in Step 3 to the two-class convolutional neural network model to obtain the transferred convolutional neural network model.
[0082] The basic principle of the transfer learning algorithm is the process of using the converged model weights obtained from the source domain dataset as the initial weights of the new model and retraining the target domain dataset. Specifically, load the five-class convolutional neural network model obtained in Step 3, use its weights as the initial weights of the two-class convolutional neural network model, and then freeze the weights of the residual convolutional module and the attention mechanism module in the two-class convolutional neural network model, that is, keep their weights unchanged during model training, while the weights in the CYP2C9 enzyme drug metabolism function prediction module are allowed to change with a smaller learning rate, and finally complete the training task of transfer learning.
[0083] Step 5: Input the target domain dataset into the transferred convolutional neural network model in Step 4 for training, adjust the model parameters to adapt to the target domain dataset, and obtain the trained target domain convolutional neural network model.
[0084] When training a binary classification convolutional neural network model on the target domain dataset, the batch size is set to 32 (i.e., training 32 haplotype data at a time), and the shape of its data tensor is [32, 8176, 13]; the learning rate is 0.00001 as the initial learning rate, and the ReduceLROnPlateau method is used to dynamically adjust the learning rate. When the validation set loss has not improved for 5 consecutive epochs, the learning rate is automatically reduced by one order of magnitude; the loss function selects the multi-class cross-entropy loss function; the optimizer uses the Adam method. In addition, an early stopping mechanism is used during training to avoid model overfitting. The accuracy and loss of the model in the source domain and target domain datasets vary with the epoch as shown in Figures 6 to 9 shown.
[0085] Finally, the data matrix obtained by preprocessing the CYP2C9 gene multi-locus mutation information obtained by methods 1 to 4 in step 1 is input into the trained binary classification convolutional neural network model to predict the CYP2C9 gene multi-locus mutation information. The prediction results are as follows: Table 2 Prediction results table for methods 1, 2, 3, and 4
[0086] Step 6: Input the CYP2C9 gene multi-locus mutation information with unknown metabolic function into the target domain convolutional neural network to predict the CYP2C9 enzyme drug metabolic function and output the metabolic function classification result.
[0087] In summary, the present invention provides a method for predicting the CYP2C9 enzyme drug metabolic function based on transfer learning, which can realize the prediction of the CYP2C9 gene multi-locus mutation information and its metabolic function. To solve the problems of few training samples and easy overfitting of the model, the present invention provides a method for generating simulated sample data to expand the training samples, and through a binary classification convolutional neural network model with high accuracy, the drug metabolic function type of the CYP2C9 gene multi-locus mutation combination can be predicted, which helps to provide a scientific basis for clinical medication at the genetic level, prevent the occurrence of drug adverse reactions, and achieve precision medication.
[0088] Example 2
[0089] This example provides a computer-readable storage medium with a computer program stored thereon. When the computer program is executed by a processor, it implements the steps of the method for predicting the CYP2C9 enzyme drug metabolic function based on transfer learning described in Example 1.
[0090] Example 3
[0091] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, the steps of a method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning described in Embodiment 1 are implemented.
[0092] Embodiment Four
[0093] This embodiment provides an application of a method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning. A method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning in Embodiment 1 is applied to the metabolic function rating of drugs metabolized by CYP2C9 enzyme to assist in guiding precise clinical medication.
[0094] The drug metabolism function prediction method of the present invention can be applied to the prediction of the metabolic functions of all drugs that can be metabolized by CYP2C9, such as anticoagulant drugs (such as warfarin), hypoglycemic drugs (such as glipizide), and antihypertensive drugs (such as losartan). Since the CYP2C9 gene polymorphism significantly affects the metabolic rates of all drugs that can be metabolized by CYP2C9, such as warfarin and glipizide, this prediction method recommends drug doses according to the evaluated metabolic function types (such as normal metabolizer, intermediate metabolizer, and poor metabolizer).
[0095] Based on the above technical innovation points, the present patent has achieved the following significant advantages in solving the problem of predicting the drug metabolism function of CYP2C9 enzyme:
[0096] (1) Break through the small sample limit and expand the prediction coverage
[0097] Simulate the generation of diploid data, expand the source domain dataset to 30,000 samples, and transfer the general features learned by the source domain five-class convolutional neural network model to the target domain dataset with only 40 verified samples through the transfer learning algorithm, solving the overfitting problem caused by insufficient data in traditional deep learning methods.
[0098] The classification accuracy of the model on the target domain dataset reaches more than 90%, and the performance is significantly improved compared with the traditional AS scoring method and conventional machine learning models. It can classify the metabolic functions of unvalidated alleles (such as *7, *10, *17, etc.) in the PharmVar database, filling the clinical gap.
[0099] (2) Accurately capture the complex effects of multi-site mutations
[0100] Combined with data augmentation, residual network, and channel attention mechanism, the residual network alleviates the vanishing gradient through bypass identity mapping, supports the training of deep networks, and can effectively extract multi-site mutation features; the attention mechanism dynamically weights key mutation sites, improving the prediction sensitivity of the model to mutations with 2 - 30 sites.
[0101] (3) Dynamically optimize the classification threshold to improve the prediction accuracy
[0102] Setting the critical value based on the maximization of sensitivity and specificity (non-functional score 0.7720637, normal functional score 0.3042464) can reduce the prediction error for poor metabolizers.
[0103] The present invention solves the problem of small data sample size through transfer learning combined with simulated data generation, uses residual network and attention mechanism to achieve complex variant analysis, and combines dynamic threshold optimization to improve the prediction accuracy, finally forming a set of efficient and accurate CYP2C9 enzyme drug metabolism function prediction system.
[0104] Finally, it should be noted that the parts not detailed in the present invention are all prior arts. Those of ordinary skill in the art can understand that the above are only preferred examples of the invention and are not used to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing examples, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, etc. made within the spirit and principle of the invention shall be included within the protection scope of the invention.
Claims
1. A method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning, characterized in that, It includes the following steps: Step 1: Obtain CYP2C9 gene variant information and generate a source domain dataset and a target domain dataset through preprocessing; The source domain dataset is generated in the following way: Randomly select CYP2C9 star alleles with known metabolic functions, construct simulated haplotypes for pre-training, and mutate the variant sites related to non-star alleles based on the known population variant frequencies to generate a source domain dataset of simulated haplotypes with data augmentation; Step 2: Construct two convolutional neural network models, including: Five-class convolutional neural network model: Used for training the source domain dataset, including a residual convolutional module, an attention mechanism module, and a CYP2C9 enzyme drug metabolism function prediction module; Two-class convolutional neural network model: Used for training the target domain dataset, with the same structure as the five-class convolutional neural network model, and the output layer adjusted to two-class; Step 3: Use the source domain dataset to train the five-class convolutional neural network model until the model converges to obtain the initial weights; Step 4: Through the transfer learning algorithm, transfer the weights of the trained five-class convolutional neural network model to the two-class convolutional neural network model, freeze the weights of the residual convolutional module and the attention mechanism module, and only update the weights of the CYP2C9 enzyme drug metabolism function prediction module with a small learning rate to obtain the transferred convolutional neural network model; Step 5: Use the target domain dataset to train the transferred convolutional neural network model, adjust the model parameters to adapt to the target domain dataset until the model converges to obtain the trained target domain convolutional neural network model; Step 6: Input the single-site, double-site, multi-site variant information or the full gene sequence of the CYP2C9 gene with unknown metabolic function into the target domain convolutional neural network model and output the metabolic function classification result.
2. A method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning according to claim 1, wherein In the above Step 1, the target domain dataset is generated in the following way: Obtain CYP2C9 gene variant information from databases, literature, or experiments, use the bioinformatics annotation software and database comparison annotation method, and adopt the independent hot encoding method to convert the haplotype variant information into a data matrix with the shape of M×P×F, where M is the number of target domain data samples, P is the gene sequence length, and F is the number of data features.
3. A method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning according to claim 1 or 2, characterized in that, In the above Step 1, the source domain dataset is generated in the following way: Randomly select a pair of CYP2C9 star alleles with normal metabolic function, decreased metabolic function, or no metabolic function, and construct a haplotype with variants related to the star alleles to create a simulated haplotype for pre-training; Then, according to the population-level alternate allele frequencies published in the GnomAD database, perform SNV and INDEL alternate allele sampling on the variant sites not related to any star alleles, select sites in a uniform distribution manner and assign corresponding mutation types, and finally form a source domain dataset containing N simulated haplotypes.
4. A method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning according to claim 1, characterized in that The convolutional neural network model includes two to four serially connected residual convolutional modules. Each residual convolutional module provides two mapping methods: residual mapping and identity mapping. The residual mapping includes a convolutional layer, a batch normalization layer, and a ReLU activation layer. The identity mapping adds the input to the result of the residual mapping to avoid gradient vanishing.
5. A method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning according to claim 1, characterized in that The implementation process of the attention mechanism module includes: performing global average pooling on the input CYP2C9 feature map to compress the spatial information of each channel into a scalar; using one-dimensional convolution operation to calculate the attention weights between channels; multiplying the calculated attention weights by the original CYP2C9 feature map to achieve weighted adjustment of channel features.
6. A method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning according to claim 1, characterized in that The implementation process of the CYP2C9 enzyme drug metabolism function prediction module includes: calculating and outputting the non-functional score and normal functional score of the haplotype metabolism function through two fully connected layers; setting the critical value of the haplotype metabolism function based on maximizing sensitivity and specificity, and then converting the score into three metabolic function classification results: normal metabolizer, intermediate metabolizer, and poor metabolizer.
7. A method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning according to claim 6, characterized in that, The critical values are: the critical value of the non-functional score is 0.7720637, and the critical value of the normal functional score is 0.3042464.
8. A method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning according to claim 1, characterized in that The training strategy of the transfer learning algorithm includes: During the training process, the weights of the model are updated with an initial learning rate of 0.00001. The ReduceLROnPlateau function is used to automatically adjust the learning rate. When the validation set loss has not improved for 5 to 10 consecutive epochs, the learning rate is automatically reduced by one order of magnitude, and the early stopping mechanism is used to avoid overfitting of the model.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning according to any one of claims 1-8.
10. A computer device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, When the processor executes the computer program, it implements the steps of the method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning according to any one of claims 1-8.
11. Application of a method for predicting CYP2C9 enzyme drug metabolism function based on transfer learning, characterized in that: Apply the method for predicting the drug metabolism function of CYP2C9 enzyme based on transfer learning according to any one of claims 1-8 to the metabolic function rating of drugs metabolized by CYP2C9 enzyme.