A method for constructing an enhancer-promoter regulatory network prediction model
By combining the sequence features, distance features, and chromatin openness features of enhancers and promoters, and employing a gradient boosting decision tree model, the problems of high data requirements and information leakage in existing technologies are solved, and efficient and accurate prediction of enhancer-promoter interactions is achieved.
Patent Information
- Application Number
- CN202411370937.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-29
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-09-29
AI Technical Summary
Existing technologies for predicting enhancer-promoter interactions suffer from problems such as high data requirements, neglect of chromatin openness information, and information leakage during cross-validation, resulting in insufficient prediction accuracy.
By acquiring the sequence features, distance features, and chromatin openness features of enhancers and promoters, a symmetric decision tree model with categorical feature gradient boosting is used to predict enhancer-promoter interactions, thus avoiding information leakage and resolving the imbalance problem.
It improves the accuracy and reliability of enhancer-promoter interaction prediction, reduces the requirements for experimental data, avoids information leakage, solves the imbalance problem, and achieves efficient and accurate prediction.
Smart Images

Figure CN119446286B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to bioinformatics technology, in particular, to the field of gene regulation in bioinformatics, more particularly, to an enhancer-promoter regulatory network prediction model construction method. BACKGROUND
[0002] Eukaryotic gene expression is determined by both the coding region of the coding gene sequence and the non-coding region that regulates gene expression. Among them, the cis-acting element is a sequence that exists in the adjacent sequence of the gene and can affect the expression of the gene in the non-coding region. The cis-acting element includes the promoter, the enhancer, etc., and their role is to participate in the regulation of gene expression. Accurate identification of the interaction between genomic elements is a core task to decipher transcriptional regulation and human disease risk.
[0003] Under the prior art, researchers have proposed various schemes to calculate and predict the interaction between candidate cis-regulatory elements. According to the technical means adopted by different schemes, the existing technologies can be divided into three categories:
[0004] The first category is to train a prediction model based on the sequence construction features of enhancers and promoters, which has no requirements for experimental data. For example, the schemes adopted in references [1], [2] and [3]. The prediction models in these schemes use the sequence features of extracted enhancers and promoters to infer the interaction between enhancers and promoters. Among them, the schemes in references [1] and [2] extract the sequence features of enhancers and promoters through convolutional neural networks, respectively, and then input the concatenation of the two into a recurrent neural network to predict whether there is an interaction. The scheme in reference [3] converts the sequence features into k-mer sentences with a step size of 1, then embeds each sentence into a vector through a paragraph vector, and concatenates all the vectors to classify through a normalized exponential function. It should be noted that the intrinsic mechanism of sequence-based prediction is that the interaction between enhancers and promoters depends on the binding of transcription factors with clear sequence preference. Although the first category of schemes is based on sequences and is user-friendly because only sequence data needs to be input, this type of method ignores information such as chromatin openness that is closely related to gene regulation, making the prediction of enhancer and promoter interaction have limitations.
[0005] The second type uses sequence characteristics of enhancers and promoters and various types of experimental data as features, such as enhancer and promoter sequences, protein sequencing, chromatin open region sequencing, methylation sequencing, and transcriptome sequencing, etc. For example, refer to the schemes in [0] and [0]. Among them, the scheme in [4] uses a gradient boosting tree classifier to classify whether the enhancer and promoter have interaction based on the features, and the scheme in [5] uses a random forest algorithm to predict whether the enhancer and promoter have interaction based on the features. The second type of scheme needs to use various types of experimental data as features, needs to input various omics data, needs to obtain various omics data through experiments, has high requirements for data, and is prone to overfitting problems.
[0006] The third type uses a combination of sequence and chromatin open region sequencing of enhancers and promoters for prediction, which can reduce the requirement for experimental data and consider convenience, for example, the scheme in [6]. Among them, the scheme in [6] uses omics data signals of chromatin open region data, sequence features of enhancers and promoters for learning, and jointly infers the interaction of enhancer-promoter pairs. In the scheme of [6], the sequence features of enhancers and promoters and the chromatin open region features are extracted by convolutional neural network respectively, then these features are spliced and input into a recurrent neural network to predict whether there is interaction. Although the third type of scheme considers sequence and chromatin open region information, increases the requirement for chromatin open region sequencing on the basis of sequence data, and increases the information related to regulation, it has relatively low requirements for experimental sequencing types, reduces the requirement for data, and considers convenience.
[0007] These algorithms have made progress in enhancer-promoter interaction prediction, helping to decipher gene regulation networks and disease mechanisms. Among them, the most typical sequence feature is represented as k-mers, i.e. a substring of length k contained in a biological sequence, and k-mers frequency features have been proven to be very effective in predicting various biological interactions, such as lncRNA classification prediction in [0], RBP binding prediction in [0], etc., which can effectively represent sequence features. However, these schemes still have some problems. For example, these algorithms in the past have been affected by inappropriate cross-validation schemes, i.e. when the training set and the test set contain the same enhancer, promoter genomic sites, the generated model may incorrectly perform well by effectively remembering the average activity related to each site during the training process, resulting in information leakage of enhancer-promoter pair information in validation and test sets, leading to defective data and false high accuracy in testing.
[0008] It should be noted that the background art is only used to introduce the related information of the present application, so as to help understand the technical solutions of the present application, but does not mean that the related information must be prior art. In the absence of evidence that the related information has been disclosed before the filing date of the present application, the related information should not be regarded as prior art.
[0009] References:
[0010] [1]Singh S,Yang Y,Póczos B,et al.Predicting enhancer-promoterinteraction from genomic sequence with deep neural networks[J].QuantitativeBiology,2019,7(2):122-137.
[0011] [2]Hong Z,Zeng X,Wei L,et al.Identifying enhancer–promoterinteractions with neural network based on pre-trained DNA vectors andattention mechanism[J].Bioinformatics,2020,36(4):1037-1043.
[0012] [3]Zeng W,Wu M,iang R.Prediction of enhancer-promoter interactionsvia natural language processing[J].BMC genomics,2018,19:13-22.
[0013] [4]Whalen S,Truty R M,Pollard K S.Enhancer–promoter interactions areencoded by complex genomic signatures on looping chromatin[J].Naturegenetics,2016,48(5):488-496.
[0014] [5]Cao Q, Anyansi C, Hu X, et al. Reconstruction of enhancer-target networks in 935 samples of human primary cells, tissues and cell lines [J]. Nature genetics, 2017, 49(10): 1428-1436.
[0015] [6]Li W, Wong W H, Jiang R. DeepTACT: predicting 3D chromatin contacts via bootstrapping deep learning [J]. Nucleic acids research, 2019, 47(10): e60-e60.
[0016] [7]Kirk J M, Kim S O, Inoue K, et al. Functional classification of long non-coding RNAs by k-mer content [J]. Nature genetics, 2018, 50(10): 1474-1482.
[0017] [8]Bressin A, Schulte-Sasse R, Figini D, et al. TriPepSVM: de novo prediction of RNA-binding proteins based on short amino acid motifs [J]. Nucleic acids research, 2019, 47(9): 4406-4417. SUMMARY
[0018] Therefore, the purpose of the present application is to overcome the defects of the prior art, and to provide an enhancer promoter regulatory network prediction model construction scheme.
[0019] The purpose of the present application is achieved by the following technical solutions:
[0020] According to a first aspect of the present application, a method for constructing a prediction model of enhancer-promoter regulatory network is provided, the prediction model being used to predict whether there is interaction between enhancers and promoters, the method comprising: S1, obtaining an original data set, the original data set comprising a plurality of enhancer-promoter pair data of a plurality of biological samples, and dividing the original data set into a plurality of subsets based on chromosome pairs, wherein all enhancer-promoter pairs on the same chromosome pair are divided into the same subset; S2, preprocessing all subsets obtained in step S1, obtaining a feature vector of each sample based on each enhancer-promoter pair, wherein each subset comprises a plurality of data samples, each data sample is an enhancer-promoter pair, the feature vector of each data sample is a feature vector formed by splicing the sequence feature of the corresponding enhancer-promoter pair, the distance feature between the enhancer-promoter pair, and the chromatin openness feature corresponding to the enhancer-promoter pair, and the label of each data sample is whether there is interaction between the corresponding enhancer-promoter pair; S3, based on all preprocessed subsets, using a categorical feature gradient boosting method, taking the feature vector of the data sample as input and whether the corresponding enhancer-promoter pair of the sample has interaction as output to iteratively construct a plurality of symmetric decision trees to form a prediction model, wherein in each iteration, any one subset is taken as a validation set and the other subsets are taken as training sets to train symmetric decision trees.
[0021] Preferably, in the step S1, the subsets are divided as follows: S11, obtaining the number of enhancers and promoters on each chromosome pair in the data set and counting the number of enhancer-promoter pairs on each chromosome pair; wherein the enhancer-promoter pair is an interacting enhancer-promoter pair; S12, sorting all chromosome pairs according to the number of enhancer-promoter pairs thereon, and complementarily combining all chromosome pairs two by two to balance the number of enhancer-promoter pairs in each combination, wherein when complementarily combining, the two ends of the sorted chromosome pairs are simultaneously moved towards the middle, and each time, the chromosome pairs at the two ends are selected as a combination, and if there is a single chromosome pair in the middle, it forms a combination by itself; S13, taking a chromosome pair combination as a subset to construct a plurality of subsets.
[0022] Preferably, in the step S2, each enhancer-promoter pair in each subset is preprocessed in the following way: S21, obtaining basic information of the enhancer and the promoter: the chromosome where the enhancer is located, the position of the enhancer, the chromosome where the promoter is located, and the position of the promoter; S22, based on the position information of the enhancer and the promoter, extracting the enhancer sequence and the promoter sequence from the genome; S23, based on the enhancer sequence and the promoter sequence respectively, obtaining the k-mer sequence features corresponding to the enhancer sequence and the promoter sequence, thereby constructing the sequence features corresponding to the enhancer and the promoter respectively; S24, based on the position information of the enhancer and the promoter, obtaining the center position of the enhancer and the center position of the promoter respectively, and calculating the distance feature of the enhancer and the promoter based on the center position of the enhancer and the center position of the promoter; S25, based on the position information of the enhancer and the promoter, obtaining the chromatin openness data corresponding to the enhancer region and the promoter region respectively, and based on this, calculating the chromatin openness feature of the enhancer and the promoter respectively; S26, concatenating the sequence features of the enhancer and the promoter, the distance features of the enhancer and the promoter, and the chromatin openness features of the enhancer and the promoter to obtain the feature vector.
[0023] Preferably, in the step S22, the enhancer sequence and the promoter sequence are extracted in the following way: find the human gene chromosome corresponding to the chromosome where the enhancer-promoter pair is located in the human genome, and obtain the center position of the enhancer and the center position of the promoter from the human gene chromosome, then based on the enhancer center position, extend the first preset length forward and backward on the human gene chromosome to extract the enhancer sequence, and based on the promoter center position, extend the second preset length forward and backward on the human gene chromosome to extract the promoter sequence.
[0024] Preferably, the first preset length is 3000bp, and the second preset length is 2000bp.
[0025] Preferably, in the step S23, the k-mer sequence features of the enhancer sequence and the promoter sequence are obtained in the following way: convert the percentage frequency of each k-mer in the enhancer sequence and the promoter sequence into a k-mer feature vector respectively, and all k-mer feature vectors in each sequence form a multi-dimensional feature vector corresponding to the sequence, thereby obtaining the sequence features corresponding to the enhancer sequence and the promoter sequence respectively.
[0026] Preferably, in the step S25, the chromatin openness feature of the enhancer is calculated in the following way:
[0027]
[0028] wherein, end en and start enend and start represent the end and start positions of the enhancer, DHS en signal of the enhancer region, m represents the number of signals in the enhancer region on the chromosome; and the chromatin openness feature of the promoter is calculated by the following method:
[0029]
[0030] wherein, end pro and start pro represent the end and start positions of the promoter, DHS pro signal of the promoter region, m' represents the number of signals in the promoter region on the chromosome.
[0031] Preferably, in the step S3, the CatBoost method is used to train the symmetric decision tree.
[0032] Preferably, in the step S3, the parameters of the decision tree are updated by the following loss:
[0033]
[0034] wherein, y i is the true label of the i-th sample in the training set, p i is the probability of the i-th sample being predicted as a positive sample.
[0035] Compared with the prior art, the present application has the advantages of high accuracy, low requirement for input data, user-friendliness, avoidance of information leakage problems existing in most existing method evaluations through effective data division, effective solution to the extremely uneven positive and negative sample problem in prediction through weighted gradient boosting of symmetric decision trees, and realization of efficient and accurate enhancer-promoter interaction prediction. BRIEF DESCRIPTION OF DRAWINGS
[0036] The embodiments of the present application will be further described below with reference to the accompanying drawings, in which:
[0037] Figure 1 It is an enhancer-promoter regulatory network model according to the embodiment of the present application. DETAILED DESCRIPTION
[0038] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings through specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.
[0039] As described in the background, the existing solutions still have the following problems, and a typical common problem is the information leakage problem brought by cross-validation, that is, when the training set and the test set contain the same enhancer-promoter genomic sites, the good performance of the generated model may be false by effectively remembering the average signal related to each site during the training process, because the enhancer-promoter pair information has information leakage in the validation and test sets, resulting in defective data and false high accuracy when testing.
[0040] In order to solve some problems existing in the prior art, the inventors propose a new enhancer-promoter regulatory network prediction model construction scheme, in which the distance feature between enhancers and promoters and the chromatin openness feature of enhancers and promoters are introduced on the basis of enhancer and promoter sequence features to improve the representation ability of the features, and a gradient boosting method is used to train a symmetric decision tree to obtain a prediction model. Among them, the inventors found that in the existing method for representing enhancer and promoter sequence features, k-mers is a simple and effective method, k-mers is a substring of length k contained in a biological sequence, and k-mers frequency feature has been proven to be very effective in predicting various biological interactions, which can effectively represent sequence features, so k-mers is still used to represent enhancer-promoter sequence features in the scheme of the present application.
[0041] In summary, according to one embodiment of the present application, the present application provides an enhancer-promoter regulatory network prediction model construction method, the prediction model is used to predict whether there is interaction between enhancers and promoters, the method comprises: S1, obtaining an original data set, the original data set contains interaction pair data of multiple chromosomes of the same biological sample, and the original data set is divided into multiple subsets based on enhancer-promoter pairs, wherein all enhancer-promoter pairs on the same chromosome pair are divided into the same subset; S2, preprocessing all subsets obtained in step S1 to obtain a feature vector of each data sample based on each enhancer-promoter pair, wherein each subset contains multiple data samples, each data sample is an enhancer-promoter pair, the feature vector of each data sample is a feature vector formed by splicing the sequence feature of the corresponding enhancer-promoter pair, the distance feature between the enhancer-promoter pair, and the chromatin openness feature corresponding to the enhancer-promoter pair, and the label of each data sample is whether there is interaction between the corresponding enhancer-promoter pair; S3, based on all preprocessed subsets, using a categorical feature gradient boosting method, taking the feature vector of the sample as the input and whether the corresponding enhancer-promoter pair of the sample has interaction as the output to iteratively construct multiple symmetric decision trees to form a prediction model, wherein in each iteration, any one subset is taken as a validation set and the other subsets are taken as training sets to train a symmetric decision tree.
[0042] In order to better understand the scheme of the present application, the present application is described in detail from the aspects of subset division, subset preprocessing and model training in combination with the drawings and examples.
[0043] I. Subset division
[0044] As introduced in the background, in the scheme of the prior art, the scheme of taking only sequence data as model input ignores information closely related to gene regulation such as chromatin openness, which makes the enhancer-promoter interaction prediction limited. The scheme of using multiple types of experimental data as features is relatively high in experimental data, and most of the schemes are affected by inappropriate cross-validation schemes, there is information leakage in the test and validation sets, which leads to defective data and false high accuracy in testing, thereby affecting the final model effect.
[0045] In order to better understand the concept of information leakage, a simple example is used for illustration. According to whether there is interaction, the enhancer-promoter pairs are divided into positive pairs and negative pairs, wherein the positive pairs: the enhancer-promoter pairs have interaction; the negative pairs: the enhancer-promoter pairs have no interaction. Taking enhancer A as an example, the positive pair of enhancer A is enhancer A-promoter B, and the negative pairs are enhancer A-promoter C, enhancer A-promoter D and enhancer A-promoter E. When dividing the data in the training process, if the positive pair enhancer A-promoter B is divided into the training set, and the negative pairs enhancer A-promoter C, enhancer A-promoter D and enhancer A-promoter E are divided into the test set, information leakage will occur, which makes the test set more likely to be predicted as negative pairs for all enhancer-promoter pairs related to enhancer A.
[0046] Therefore, in order to avoid information leakage, in the scheme of the present application, the positive pairs and negative pairs related to the same enhancer or promoter are divided into the same cross-validation subset when performing cross-validation division. According to an embodiment of the present application, in the scheme of the present application, in order to avoid inappropriate cross-validation problems, the existing data collected for training the model is re-divided, and the positive pairs and negative pairs of enhancer-promoter pairs in the same chromosome are always allocated to the same subset according to the chromosome position, so as to ensure that the positive pairs and negative pairs related to the same enhancer or promoter cannot appear in different subsets at the same time to cause data leakage.
[0047] In addition, since only a small part of all possible enhancer and promoter pairs formed by all enhancers and all promoters in pairs exists interaction, the number of enhancer and promoter pairs with interaction (positive pairs) is far less than that of pairs without interaction (negative pairs), the machine learning task is prone to typical imbalance problem (there are many enhancer-promoter pairs in the data set, the interaction of these pairs is known, that is, whether there is interaction between them is known, and the number of enhancer-promoter pairs with interaction is much less than that of enhancer-promoter pairs without interaction). In order to solve the imbalance problem, bootstrapping and oversampling are the preferred solutions of existing methods, however, these methods artificially change the real distribution of interaction samples and non-interaction samples, which not only deviates from the original situation, but also increases the memory and time cost. Therefore, in the scheme of the present application, the similarity size of different subsets is maintained by pairing large chromosomes with small chromosomes, so as to avoid the imbalance problem without changing the real distribution of samples.
[0048] According to an embodiment of the present application, in the scheme of the present application, the subsets are divided as follows: the number of enhancers and promoters on each chromosome pair in the data set is obtained and the number of enhancer-promoter positive pairs on each chromosome pair is counted; wherein the enhancer-promoter positive pair is an enhancer-promoter pair with interaction; all chromosome pairs are sorted according to the number of enhancer-promoter positive pairs thereon, and all chromosome pairs are complementarily combined two by two to balance the number of enhancer-promoter positive pairs in each combination, wherein when complementarily combining, the chromosome pairs at both ends of the sorted are simultaneously moved towards the middle, each time the chromosome pairs at both ends are selected as a combination, and if there is a single chromosome pair in the middle, it forms a combination by itself; a chromosome pair combination is taken as a subset, and a plurality of subsets are constructed.
[0049] According to one embodiment of the present application, the present application collects experimental data from TargetFinder and BENGI, including Hi-C, pcHi-C, CHIA-PET, eQTLs and CRISPR / Cas9, etc. chromatin openness feature data sets, each of which contains multiple sub-data sets: 1. TargetFinder data sets include GM12878 (B lymphocytes), HUVEC (umbilical vein endothelial cells), HeLa-S3 (ectodermal line cervical cancer cells), IMR90 (human lung fibroblasts), K562 (mesodermal line leukemia cells), NHEK (epidermal keratinocytes); 2. BENGI data sets include CD34.CHiC (CD34 cell-CHiC), GM12878.CHiC (B lymphocyte-CHiC), GM12878.CTCF-ChIAPET (B lymphocyte-ChIAPET), GM12878.HiC (B lymphocyte-HiC), GM12878.RNAPII-ChIAPET (B lymphocyte-ChIAPET), HeLa.CTCF-ChIAPET (ectodermal line cervical cancer cell-ChIAPET), HeLa.HiC (ectodermal line cervical cancer cell-HiC), HeLa.RNAPII-ChIAPET (ectodermal line cervical cancer cell-ChIAPET), HMEC.HiC (umbilical vein endothelial cell-HiC), IMR90.HiC (human lung fibroblast-HiC), K562.HiC (mesodermal line leukemia cell-HiC), NHEK.HiC (epidermal keratinocyte-HiC). Each of the sub-data sets contains multiple biological samples, each of which is a cell.
[0050] Take the K562 cell line of the K562 sub-data set as an example, there are 23 pairs of chromosomes, one of which is the sex chromosome (chr1-22+chrX). Each chromosome contains genes, which are DNA fragments with genetic effects, are the basic genetic units that control biological traits, record and transmit genetic information, and all genes are composed of sequences of four bases. Each chromosome has many enhancers and promoters to regulate gene transcription. Both enhancers and promoters are regulatory elements that affect gene transcription, but they differ in their location and manner of action. Among them, the promoter is located upstream of the gene coding region, where RNA polymerase II binds and starts transcription, while the enhancer is located upstream or downstream of the promoter, and can exist at a distance. They not only increase the rate of transcription, but also adjust the timing and location of transcription. Through the above sequencing technologies such as HiC, ChiC, RNAP II-ChIAPET and CTCF-ChIAPET, the positions of the promoters and enhancers in the chromosome (such as promoter 128745bp, enhancer 168754bp) and whether the promoter and enhancer interact (the tags can be set to 0 / 1, 0 indicating no interaction and 1 indicating interaction) can be obtained. First, count the number of enhancer-promoter pairs on the same pair of chromosomes measured by the corresponding experimental technology, then sort the chromosome pairs in ascending or descending order according to the number of pairs, then set the chromosome pairs at the front and back as complementary groups, and the complementary size of the chromosome pairs is assigned to the same sub-set (for example, the chromosome pairs ranked 1 and 23 are assigned to the same group, the chromosome pairs ranked 2 and 22 are assigned to the same group, and so on. The chromosome pair ranked 12 forms a group), so that the positive and negative pairs on the same chromosome are assigned to the same sub-set, so that these groups contain approximately the same number of positive pairs.
[0051] II. Preprocessing of sub-sets
[0052] The inventors have found that enhancer and promoter sequence data, enhancer and promoter distance, and chromatin accessibility are closely related to gene regulation and can be used to study enhancer and promoter interactions, so in the scheme of the present application, the distance feature and the chromatin accessibility feature of the enhancer and the promoter are added to the feature level, that is, the sequence feature, the distance feature and the chromatin accessibility feature of the enhancer and the promoter are used to predict whether the enhancer-promoter pair exists. Among them, as the name implies, the distance feature refers to the distance between the enhancer and the promoter. Chromatin accessibility reveals the degree of openness or accessibility of chromatin, and this difference in degree can have an important impact on the transcriptional activity of genes, and the chromatin accessibility of the promoter and the enhancer indicates the degree to which the promoter and the enhancer can be approached.
[0053] In each subset, the content of each data sample before preprocessing includes: 1) the chromosome where the promoter is located; 2) the position of the promoter; 3) the chromosome where the enhancer is located; 4) the position of the enhancer. The sample label is whether the promoter and the enhancer interact, marked as 0 / 1. Based on these basic information, the following preprocessing is performed on each subset:
[0054] Firstly, the enhancer / promoter sequence and its upstream and downstream sequences are extracted, specifically, as shown in the following table, the corresponding chromosome and the center position of the corresponding enhancer and promoter in the chromosome are found in the human reference genome (Hg19 or Hg38), the enhancer is filled with 3000bp before and after, the promoter is filled with 2000bp before and after, and then the sequence of the position where the enhancer and the promoter are located is extracted from the human genome using the bedtools genome data analysis tool. Figure 1
[0055] Then, the enhancer sequence and the promoter sequence are processed to obtain sequence features, specifically, in the scheme of the present application, still referring to the following table, the percentage frequency of each k-mer (k-mer is a sequence fragment with a length of k) in the enhancer and promoter sequence is calculated as a k-mer feature vector. Taking k as 3 as an example, the length of the substring is 3, there are 4 possibilities of ATGC at each position, that is, the total possibility is 4 to the power of 3, 64 possible substrings, such as AAA, AAT, AAG, AAC, ATA, ATT, ATG, ATC, etc. In the sequence region of the enhancer or the promoter, the percentage frequency of each substring in the short sequence can be counted, so as to encode the sequence features into a 64-dimensional vector, each dimension corresponds to a substring, and the value of each dimension is the percentage frequency of the substring. Similarly, if k is 4, the length of the substring is 4, there are 4 possibilities of ATGC at each position, that is, the total possibility is 4 to the power of 4, 256 possible substrings, such as AAAA, AATA, AAGA, AACA, ATAA, ATTA, ATGA, ATCA, etc. In the sequence region of the enhancer or the promoter, the percentage frequency of each substring in the short sequence can be counted, so as to encode the sequence features into a 256-dimensional vector, each dimension corresponds to a substring, and the value of each dimension is the percentage frequency of the substring. The same applies to substrings of other lengths, which will not be described here. The k-mer feature vectors of the enhancer and the promoter are normalized respectively to obtain the sequence features. Figure 1
[0056] Next, the chromatin accessibility of the enhancer and the promoter is calculated respectively, that is, the DNase (chromatin openness feature) signal of the enhancer and the promoter is added to the feature vector, wherein the chromatin openness feature of the enhancer is calculated by the following method:
[0057]
[0058] where end and start represent the end and start positions of the enhancer, DHS represents the signal of the enhancer region, and m represents the number of signals in the enhancer region on the chromosome. en and start en represent the end and start positions of the promoter, DHS represents the signal of the promoter region, and m' represents the number of signals in the promoter region on the chromosome. en pro pro pro where end and start represent the end and start positions of the enhancer, DHS represents the signal of the enhancer region, and m represents the number of signals in the enhancer region on the chromosome.
[0059]
[0060] where end and start represent the end and start positions of the enhancer, DHS represents the signal of the enhancer region, and m represents the number of signals in the enhancer region on the chromosome. pro and start pro represent the end and start positions of the promoter, DHS represents the signal of the promoter region, and m' represents the number of signals in the promoter region on the chromosome. pro
[0061] Next, the distance between the center positions of the enhancer and the promoter is added to the feature vector, where the center position of the enhancer and the center position of the promoter are obtained based on the basic information of the enhancer and the promoter, respectively, and the distance feature of the enhancer and the promoter is calculated based on the center position of the enhancer and the center position of the promoter.
[0062] Finally, the two percentage-normalized k-mer feature vectors, i.e., the sequence features of the enhancer and the promoter, the distance features of the enhancer and the promoter, and the chromatin openness features of the enhancer and the promoter, are spliced to obtain the feature vector.
[0063] III. Model training
[0064] According to one embodiment of the present application, the present application uses a gradient boosting-based symmetric decision tree to predict whether an enhancer-promoter pair interacts, and uses a categorical feature boosting algorithm package (CatBoost) to implement it. Because, in the study of enhancer-promoter interaction, the inventors found that CatBoost has made many innovative optimizations in model training compared to traditional gradient boosting decision tree algorithms. In particular, when dealing with category features, CatBoost introduces an order-based target encoding method, which avoids the common high-dimensional problem of category features and the risk of overfitting. CatBoost also has a unique ability to handle data partial order problems, which can effectively prevent target leakage in training data. In addition, CatBoost uses a symmetric tree (Symmetric Trees) structure, making the model training and prediction more efficient, while reducing memory consumption.
[0065] Still referring to Figure 1 The inter- and non-interacting enhancer and promoter pairs are distinguished by learning weighted gradient boosting symmetric decision trees by inputting the feature vectors into a classifier of a categorical feature boosting algorithm package (CatBoost). CatBoost is a supervised machine learning method, which is a gradient boosting decision tree algorithm optimized for categorical features. CatBoost has two main functions, which are suitable for category data and use gradient boosting (Boost), and is a gradient boosting framework with less parameters, support for categorical variables and high accuracy based on symmetric decision trees (oblivious trees) as base learners. The decision tree model is essentially a tree composed of multiple judgment nodes. At each node of the tree, a parameter judgment is made, and the best judgment of the value of the variable of interest can be made at the last branch (leaf node) of the tree. Generally, a decision tree contains a root node, several internal nodes and several leaf nodes, and the leaf nodes correspond to the decision classification results. Branches make judgments, and leaves make conclusions. In a symmetric decision tree, the judgment conditions of each node at each level are the same, the fitting mode is relatively simple, the prediction speed can be improved, and fast prediction can be achieved. The core idea of the gradient boosting algorithm is to iteratively construct and combine multiple decision trees to gradually reduce the prediction residual and thus improve the overall performance of the model. In the process of iteratively constructing decision trees, each subsequent tree improves the results of the previous tree to obtain better results. For example, the fitted decision tree may make decisions on the distance feature at a certain level node, and decide to be the left node of the next level if the distance is greater than a certain value, and decide to be the right node of the next level if the distance is less than or equal to a certain value.
[0066] According to one embodiment of the present application, in the present application, multiple subsets of the same biological sample after pretreatment are divided into multiple cross-validation groups, one cross-validation group is used for testing, and the other cross-validation groups are used for constructing decision trees; multiple decision trees are iteratively constructed, and each decision tree corrects the prediction value of the previous decision tree. In the training process, the iteration number is 1000 times, the tree depth is 10, and the learning rate is 0.1. Since the technical theory of CatBoost is known to those skilled in the art, CatBoost will not be described in detail here.
[0067] In the scheme of the present application, in order to solve the imbalance problem, the parameter'scale_pos_weight' in the categorical feature boosting algorithm package (CatBoost) is set to the initial imbalance ratio, the weighted decision tree is adjusted, the model is helped to pay attention to a small amount of data samples of the category, the imbalance problem is solved without changing the distribution of the data itself, and fast and accurate prediction is realized. The loss function is a log loss function (Logloss), i.e. a log likelihood loss, and the expression is:
[0068]
[0069] where y i is the true label of the i-th data sample, p i is the probability of the i-th sample being predicted as a positive sample.
[0070] As can be seen from the above embodiments, the present application aims to solve the limitations of the existing enhancer-promoter interaction prediction technology, such as high requirements for experimental data of input features, neglecting chromatin openness and other epigenetic information, and information missing defects in cross-validation scheme, and proposes a new scheme, which comprehensively considers the sequence features, distance features and chromatin accessibility features of enhancers and promoters, and uses a weighted gradient boosting symmetric decision tree algorithm to solve the imbalance problem and reassign subsets, thereby improving the accuracy and reliability of enhancer-promoter interaction prediction.
[0071] Compared with the prior art, the scheme of the present application has the following beneficial effects: 1. The enhancer-promoter interaction is predicted by comprehensively considering the sequence features, distance features and chromatin accessibility features of enhancers and promoters; the distance features and chromatin accessibility features improve the prediction accuracy of the promoter-enhancer interaction, and the acquisition methods of the three types of features are user-friendly, and the types of experimental data are not high; 2. The positive and negative pairs of enhancer-promoter interactions on the same chromosome are allocated to the same subset by chromosome position, and the large chromosomes are paired with the small chromosomes to control the similar size of different subsets, effectively avoiding the information leakage of the validation and test sets; 3. The weighted gradient boosting symmetric decision tree is realized by using the category feature boosting algorithm package (CatBoost, Category Boosting), and by setting the penalty term, the model pays more attention to the data samples classified into the wrong category; effectively solving the problem of extremely unbalanced positive and negative samples in enhancer-promoter interaction prediction, and improving the prediction accuracy of the promoter-enhancer interaction.
[0072] In order to verify the effect of the present application, the performance of the model constructed by the scheme of the present application and other models is compared in different biological samples on different data sets, and the area under the precision-recall curve (AUPR) is used as the performance indicator for evaluating the performance of the present application, wherein AUPR is a classic indicator for evaluating the performance of the model on unbalanced data.
[0073] Table 1 is the AUPR comparison of the model constructed by the scheme of the present application (denoted as EPnet) and TargetFinder, EP2vec, SPEID, EPIVAN model in different biological samples in the TargetFinder dataset. It can be seen that the average AUPR of EPnet in the TargetFinder dataset is 0.81, which is higher than the comparison methods TargetFinder, EP2vec, SPEID and EPIVAN, etc.
[0074] Table 1
[0075]
[0076]
[0077] Table 2 is the AUPR comparison of the model constructed by the scheme of the present application (denoted as EPnet) and TargetFinder, EP2vec, SPEID, EPIVAN model in different biological samples in the BENGI dataset. It can be seen that the average AUPR of EPnet in the BENGI is 0.5841, which is also higher than the comparison methods.
[0078] Table 2
[0079]
[0080] It should be noted that although the above describes each step in a specific order, it does not mean that each step must be performed in the above specific order, in fact, some of these steps can be performed concurrently, or even change the order, as long as the required function can be achieved.
[0081] The present application can be a system, a method and / or a computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions embodied therewith, which instructions are used to program a processor to implement the various aspects of the present application.
[0082] A computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves.
[0083] Having described above several embodiments of the present application, any modifications and variations that fall within the scope of the described embodiments can be apparent to those skilled in the art. The foregoing description of various embodiments of the present application is exemplary and explanatory only and is not intended to be limiting. Various modifications and variations of the described embodiments of the application will be apparent to those skilled in the art from the foregoing description and teachings. It is intended that the scope of the application should be limited only by the broadest interpretation of the appended claims to be supported by this description, including the full range of equivalents to which such claims are entitled. The terminology used or introduced herein is for the purpose of describing particular embodiments only and is not intended to be limiting.
Claims
1. A method for constructing a prediction model of enhancer-promoter regulatory network, the prediction model being used to predict whether there is interaction between enhancers and promoters, characterized in that, The method comprises: S1, obtaining an original data set, the original data set comprising multiple enhancer-promoter pair data of multiple biological samples, and dividing the original data set into multiple subsets based on chromosome pairs, wherein all enhancer-promoter pairs on the same chromosome pair are divided into the same subset; wherein the subsets are divided as follows: S11, obtaining the number of enhancers and promoters on each chromosome pair in the original data set and counting the number of enhancer-promoter positive pairs on each chromosome pair; wherein the enhancer-promoter positive pair is an interacting enhancer-promoter pair; S12, sorting all chromosome pairs according to the number of enhancer-promoter positive pairs thereon, and complementarily combining all chromosome pairs two by two to balance the number of enhancer-promoter positive pairs in each combination, wherein when complementarily combining, the two ends of the sorted are simultaneously moved towards the middle, and each time, the chromosome pairs at the two ends are selected as a combination, and if there is a single chromosome pair in the middle, it forms a combination by itself; S13, taking one chromosome pair combination as one subset to construct multiple subsets; S2, preprocessing all subsets obtained in step S1 to obtain a feature vector of each data sample based on each enhancer-promoter pair, wherein each subset comprises multiple data samples, each data sample is an enhancer-promoter pair, the feature vector of each data sample is a feature vector formed by splicing the sequence feature of the corresponding enhancer-promoter pair, the distance feature between the enhancer-promoter pair, and the chromatin openness feature corresponding to the enhancer-promoter pair, and the label of each data sample is whether the corresponding enhancer-promoter pair has interaction; S3, based on all preprocessed subsets, using a categorical feature gradient boosting method, taking the feature vector of the data sample as the input and whether the enhancer-promoter pair corresponding to the sample has interaction as the output to iteratively construct multiple symmetric decision trees to form a prediction model, wherein in each iteration, any one subset is taken as a validation set and other subsets are taken as training sets to train symmetric decision trees.
2. The method of claim 1, wherein, In the step S2, each enhancer-promoter pair in each subset is preprocessed as follows: S21, obtaining basic information of the enhancer and the promoter: the chromosome where the enhancer is located, the position of the enhancer, the chromosome where the promoter is located, and the position of the promoter; S22, based on the position information of the enhancer and the promoter, extracting the enhancer sequence and the promoter sequence from the genome; S23, based on the enhancer sequence and the promoter sequence respectively, obtaining the k-mer sequence feature corresponding to the enhancer sequence and the promoter sequence, thereby constructing the sequence feature corresponding to the enhancer and the promoter respectively; S24, based on the position information of the enhancer and the promoter, obtaining the center position of the enhancer and the center position of the promoter respectively, and calculating the distance feature of the enhancer and the promoter based on the center position of the enhancer and the center position of the promoter; S25, based on the position information of the enhancer and the promoter, obtaining the chromatin openness data corresponding to the enhancer and the promoter regions respectively, and based on this, calculating the chromatin openness feature of the enhancer and the promoter respectively; S26, splice the sequence features of the enhancer and the promoter, the distance features of the enhancer and the promoter, and the chromatin openness features of the enhancer and the promoter to obtain a feature vector.
3. The method of claim 2, wherein, In the step S22, the enhancer sequence and the promoter sequence are extracted by the following method: In the human genome, find the human gene chromosome corresponding to the chromosome on which the enhancer-promoter pair is located, and obtain the center positions of the enhancer and the promoter from the human gene chromosome. Extract the enhancer sequence by extending the first preset length before and after the enhancer center position on the human gene chromosome respectively, and extract the promoter sequence by extending the second preset length before and after the promoter center position on the human gene chromosome respectively.
4. The method of claim 3, wherein, The first preset length is 3000bp, and the second preset length is 2000bp.
5. The method of claim 2, wherein, In the step S23, the k-mer sequence features of the enhancer sequence and the promoter sequence are obtained by the following method: Convert the percentage frequency of each k-mer in the enhancer sequence and the promoter sequence into a k-mer feature vector respectively, and all the k-mer feature vectors in each sequence form a multi-dimensional feature vector corresponding to the sequence, thereby obtaining the sequence features corresponding to the enhancer sequence and the promoter sequence respectively.
6. The method of claim 2, wherein, In the step S25, the chromatin openness features of the enhancer are calculated by the following method: wherein, and represent the end and start position of the enhancer, respectively, represents the signal of the enhancer region, represents the number of signals within the enhancer region on the chromosome; and the chromatin openness features of the promoter are calculated by the following method: wherein, and denote the end and start position of the promoter, respectively, denotes the signal of the promoter region, denotes the number of signals within the promoter region on the chromosome.
7. The method of claim 2, wherein, In the step S3, the CatBoost method is used to train the symmetric decision tree.
8. The method of claim 7, wherein, In the step S3, the parameters of the decision tree are updated by the following loss: wherein, is the true label of the th sample in the training set, is the probability that the th sample is predicted as a positive sample.
9. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program can be executed by a processor to implement the steps of the method of any one of claims 1 to 8.
10. An electronic device, comprising: Comprise: One or more processors; And A memory, wherein the memory is used to store executable instructions; The one or more processors are configured to implement the steps of the method of any one of claims 1 to 8 by executing the executable instructions.
Citation Information
Patent Citations
Enhancer-promoter interaction prediction method based on DNA sequence and genome signal characteristics
CN117542413A
Methods, systems and devices comprising support vector machine for regulatory sequence features
WO2016183348A1