Automatic Extraction and Identification Method of P450 Genes Based on Omics Data
Through an automated method based on omics data, combined with amino acid word frequency and gene evolution clustering information, efficient and accurate identification of P450 gene is achieved, solving the problems of low efficiency and large errors of traditional methods.
Patent Information
- Application Number
- CN202510316427.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-18
AI Technical Summary
The traditional P450 gene extraction and identification method is inefficient and prone to artificial errors when processing massive amounts of omics data, affecting the accuracy and reliability of the results.
An automated method based on omics data is adopted to learn the characteristics of P450 genes of different species through amino acid word frequency, and assist in rapid identification based on gene evolution clustering information, and accurately identify it in combination with multi-model prediction results.
It significantly improves the accuracy and efficiency of P450 gene identification, reduces the error caused by human factors, is suitable for large-scale biological data processing, and has good clinical application prospects.
Smart Images

Figure CN119851772B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gene identification, and particularly to an automatic extraction and identification method for P450 genes based on omics data. Background Art
[0002] The identification of P450 genes has far-reaching significance in the field of biomedicine. P450 genes are not only involved in the detoxification of drugs in the human liver and bile acid metabolism, but also participate in the biosynthesis of anti-tumor drug paclitaxel, anti-malaria drug artemisinin, antibacterial drug penicillin, and lipid-lowering statin drugs. Identifying P450 genes in the transcriptomes of different human states can help us better understand the metabolic differences of individuals, predict the responses of individuals to the environment, thereby guiding clinical medication and improving the safety and effectiveness of drug treatment; identifying P450 genes in the genomes of different species can help us better understand the distribution characteristics of P450 biological groups, explore the structural and functional diversity of P450, reveal the pathways and mechanisms involved in its metabolism, and provide new ideas and methods for the screening of innovative biological drugs, disease treatment and prevention.
[0003] In animal genomes, the number of members of the P450 gene family is large and widely distributed. For example, there are hundreds of different P450 sequences in mammals; there are thousands of P450s in wheat. At the same time, the sequence similarity of P450s between different species is relatively low, usually less than 40%. When conducting cross-species comparisons or whole-genome identifications, it is difficult to find all relevant P450 sequences through simple sequence alignment or search methods. Traditional gene identification methods mainly rely on manual operations, which are particularly time-consuming and laborious when dealing with massive data. More importantly, human factors may lead to errors and inconsistencies in the data processing process, thus seriously affecting the accuracy and reliability of subsequent analysis results.
[0004] To overcome the above limitations, it is particularly important to develop an efficient and accurate automated gene identification processing flow. The automated gene identification process can make full use of the advantages of computer technology and algorithms, quickly process massive omics data, and accurately identify, extract and analyze relevant genes, improve the efficiency and accuracy of data processing, and reduce errors and inconsistencies caused by human factors. Summary of the Invention
[0005] The purpose of the present invention is to provide an automatic extraction and identification method for P450 genes based on omics data. By comprehensively considering the differences in various types of omics data, fully learning the characteristics of P450 genes in different species through amino acid word frequencies, and at the same time assisting the rapid identification of P450 based on gene evolution clustering information, and finally accurately identifying P450 genes by integrating the prediction results of multiple models, so as to overcome the limitations of traditional P450 gene extraction and identification.
[0006] To achieve the above object, the present invention provides an automatic extraction and identification method for P450 genes based on omics data, and the method includes:
[0007] Perform normalization processing on different omics amino acid sequence data to obtain multiple gene homology groups;
[0008] Obtain P450 and non-P450 amino acid sequence data and summarize them into a data set;
[0009] Based on gene evolution clustering, split the data set into different groups to obtain grouping information, use the grouping information as an auxiliary feature for classification and identification, and add it to the data set;
[0010] Divide the data set into training data and test data, train the classification model with the training data, and evaluate the performance of the trained classification model with the test data;
[0011] Input multiple gene homology groups into the trained classification model and output the P450 gene identification result.
[0012] Further, the performing normalization processing on different omics amino acid sequence data to obtain multiple gene homology groups specifically includes:
[0013] Judge the attributes of different omics amino acid sequence data to distinguish whether the amino acid sequence data is protein data or nucleotide data;
[0014] For nucleotide data, judge whether it is annotated nucleotide data, translate the annotated nucleotide data into protein data and perform normalization;
[0015] For unannotated nucleotide data, further judge whether it comes from eukaryotes or prokaryotes;
[0016] Extract similar sequence fragments from unannotated nucleotide data from eukaryotes and prokaryotes respectively, perform gene structure prediction on the similar sequence fragments, and obtain potential P450 protein sequences.
[0017] Further, identify, replace and format the uncommon amino acids in the potential P450 protein sequences and the non-underlined characters in the fasta tags, and the uncommon amino acids are amino acids that do not belong to the standard 20 amino acids.
[0018] Further, the obtaining P450 and non-P450 amino acid sequence data and summarizing them into a data set specifically includes:
[0019] Convert the obtained P450 and non-P450 amino acid sequence data into text content, and use the CountVectorizer tool for text feature extraction in Scikit-learn to extract phrases of single amino acids, consecutive two amino acids, and consecutive three amino acids according to ngram_range=(1, 3);
[0020] Use regular expressions to identify each amino acid phrase so that only amino acid phrases will be used as the construction units of n-grams;
[0021] Generate a phrase-document matrix to record the occurrences of each n-gram in the text, and convert the amino acid phrases into word frequency representations in the encoding method of the bag-of-words model.
[0022] Furthermore, before converting the P450 and non-P450 amino acid sequence data into text content, perform preliminary cleaning and formatting on it.
[0023] Furthermore, the dataset is split into different groups based on gene evolution clustering to obtain grouping information. Specifically: process the dataset through the orthofinder software to obtain clustering information based on sequence evolution, and designate the data belonging to the same cluster in the clustering information as the same auxiliary group to obtain grouping information.
[0024] Furthermore, the grouping information is represented in the one-hot encoding method.
[0025] Furthermore, add the grouping information to the dataset. Specifically, splice the one-hot encoding of the grouping information with the bag-of-words model encoding of the amino acid word frequencies in the dataset to form a feature vector.
[0026] Furthermore, the performance of the trained classification model is evaluated using test data, and the performance includes precision, recall, F1 score, and accuracy.
[0027] Furthermore, input multiple gene homology groups into the trained classification model to output P450 gene identification results. Specifically, use linear regression models, Bayesian models, extreme gradient boosting models, support vector machine models, and random forest models based on natural language processing methods for amino acid sequence features as base learners, train each base learner separately, and at the same time construct an extreme gradient boosting model as a meta-learner. Input multiple gene homology groups into multiple base learners respectively, use the outputs of the base learners as the inputs of the meta-learner, and perform final prediction to output P450 gene identification results.
[0028] Compared with the prior art, the beneficial effects of the present invention are:
[0029] 1. The method provided by the present invention combines bioinformatics data processing with machine learning techniques for constructing a bag of words by combining amino acid word frequencies and gene evolution clustering grouping information, and is used for P450 gene identification. This method can assist machine learning algorithms in deeply and accurately analyzing gene sequences, thereby significantly improving the accuracy and efficiency of gene identification.
[0030] 2. Compared with traditional gene extraction methods, the method provided by the present invention has a simpler and clearer process, can reduce the time and labor costs of gene identification, making the method more economical and practical and easier to promote on a large scale.
[0031] 3. The method identifies P450 sequences based on automated bioinformatics processing and rapid machine learning, and can efficiently process large-scale biological data sets. This method is not only applicable to basic scientific research, but also has good clinical application prospects. At the same time, this method can also be applied to the identification of other genes in addition to P450 genes, providing strong support for disease diagnosis, drug development, genetic counseling, precision medicine and personalized treatment. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the drawings in the following description are only the preferred embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0033] Figure 1 is a schematic diagram of the overall process of a method for automatically extracting and identifying P450 genes based on omics data provided by an embodiment of the present invention.
[0034] Figure 2 is a schematic diagram of confusion matrix evaluation provided by an embodiment of the present invention.
[0035] Figure 3 is a schematic diagram of the ROC curve provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The following describes the principles and features of the present invention with reference to the accompanying drawings. The listed embodiments are only used to explain the present invention and are not used to limit the scope of the present invention.
[0037] Referring to Figure 1 , this embodiment provides a method for automatically extracting and identifying P450 genes based on omics data, and the method includes:
[0038] S101. Perform normalization processing on different omics amino acid sequence data to obtain multiple gene homologous groups.
[0039] S102. Obtain P450 and non-P450 amino acid sequence data and summarize them into a data set.
[0040] Exemplarily, in this step, P450 amino acid sequence data is collected through manual screening, and an equal amount of non-P450 amino acid sequence data is collected from the uniport public database.
[0041] S103. Based on gene evolution clustering, split the data set into different groups, obtain grouping information, use the grouping information as an auxiliary feature for classification and identification, and add it to the data set.
[0042] S104. Divide the data set into training data and test data, train the classification model with the training data, and evaluate the performance of the trained classification model with the test data.
[0043] S105. Input multiple gene homologous groups into the trained classification model and output the P450 gene identification result.
[0044] As a possible implementation manner, the homogenization process of different omics amino acid sequence data to obtain multiple gene homologous groups specifically includes:
[0045] S201. Judge the attributes of different omics amino acid sequence data to distinguish whether the amino acid sequence data is protein data or nucleotide data.
[0046] S202. For nucleotide data, judge whether it is annotated nucleotide data, translate the annotated nucleotide data into protein data and perform homogenization.
[0047] S203. For unannotated nucleotide data, further judge whether it is from eukaryotes or prokaryotes.
[0048] S204. Respectively extract similar sequence fragments from unannotated nucleotide data from eukaryotes and prokaryotes, perform gene structure prediction on the similar sequence fragments, and obtain potential P450 protein sequences.
[0049] In this implementation manner, based on sequence similarity and homology alignment of the P450 data set, extract genomic similar sequence fragments. The P450 data set used for similar fragment alignment is a manually screened and summarized P450 sequence of the whole biological group.
[0050] In this implementation manner, for the similar sequence fragments of unannotated nucleotide data from eukaryotes, the Augustus software can be used for gene structure prediction; for the similar sequence fragments of unannotated nucleotide data from prokaryotes, the Prokka software can be used for gene structure prediction.
[0051] In this embodiment, after obtaining the potential P450 protein sequence, the uncommon amino acids in the potential P450 protein sequence and the non-underlined characters in the fasta label are identified, replaced, and formatted to ensure the uniformity and accuracy in subsequent data analysis. The uncommon amino acids are those that do not belong to the standard 20 amino acids. Exemplarily, the replacement can be replacing the uncommon amino acids with X and replacing the non-underlined characters with underlines.
[0052] In this embodiment, similar sequence fragments are first extracted from the species genomic data, and then gene structure prediction is performed, which can reduce the time consumption caused by prediction.
[0053] As a possible embodiment, the operations of obtaining the P450 and non-P450 amino acid sequence data and summarizing them into a data set specifically include the following steps:
[0054] S301. Convert the obtained P450 amino acid sequence data and non-P450 amino acid sequence data into text content, and use the tool CountVectorizer for text feature extraction in Scikit-learn to extract phrases of single amino acids, consecutive two amino acids, and consecutive three amino acids according to ngram_range=(1,3).
[0055] S302. Use regular expressions to identify each amino acid phrase so that only the amino acid phrases will be used as the construction units of n-grams.
[0056] S303. Generate a phrase-document matrix to record the occurrences of each n-gram in the text, and convert the amino acid phrases into a term frequency representation in the encoding mode of the bag-of-words model.
[0057] In this embodiment, before converting the P450 and non-P450 amino acid sequence data into text content, it is preliminarily cleaned and formatted to prepare for subsequent model training. Exemplarily, the process of the preliminary cleaning includes reading the data file and then performing data cleaning, such as deleting the rows with empty content or labels.
[0058] In this embodiment, by manually screening the P450 database protein sequences from multiple categories and multiple species sources, the extensiveness of the P450 training data can be ensured, and the sequence characteristics of different species can be fully demonstrated. At the same time, sequence feature extraction is performed based on the amino acid term frequency, enabling the model to fully learn the P450 characteristics of different species.
[0059] As another possible implementation, the dataset is split into different groups based on gene evolution clustering to obtain grouping information. Specifically, the dataset is processed by the orthofinder software to obtain clustering information based on sequence evolution. The data belonging to the same cluster in the clustering information is designated as the same auxiliary group to obtain the grouping information.
[0060] In this implementation, auxiliary grouping is achieved through gene evolution clustering. This grouping can better display the homologous information of genes compared to sequence similarity grouping. The combination of training with the above multi-group data, extraction of word frequency features, and evolutionary grouping assistance can further improve the accuracy of the model for P450 identification.
[0061] In this implementation, the grouping information is represented in the one-hot encoding method.
[0062] The above implementation is illustrated by the example shown in Table 1 below.
[0063] Table 1 Dataset Example
[0064]
[0065] The amino acid word frequency is represented in the bag-of-words model encoding method:
[0066] P450 =
[0067] [1, 1, 1, 1, 1, 0, 0, 0] # Sequence 1
[0068] [……] # Other P450 sequences
[0070] UnP450 =
[0071] [1, 0, 0, 0, 1, 1, 1, 1] # Sequence 2
[0072] [……] # Other non-P450 sequences
[0074] For the grouping information, it is represented in the one-hot encoding method:
[0075] P450 =
[0076] [1, 0] # Sequence 1 # Auxiliary group is 1
[0077] [……] # Other P450 sequences
[0079] UnP450 =
[0080] [0, 1] # Sequence 2 # Auxiliary group is 2
[0081] [……]# Other non-P450 sequences
[0083] Use FeatureUnion to merge the representation results by column to form a matrix for subsequent use by the model. The matrix can be represented as:
[0084] P450 =
[0085] [1, 1, 1, 1, 1, 0, 0, 0, 1, 0] # Sequence 1
[0086] [……]# Other P450 sequences
[0088] UnP450 =
[0089] [1, 0, 0, 0, 1, 1, 1, 1, 0, 1] # Sequence 2
[0090] [……]# Other non-P450 sequences
[0092] In another alternative embodiment, add grouping information to the dataset. Specifically, concatenate the one-hot encoding of the grouping information with the bag-of-words model encoding of the amino acid word frequencies in the dataset to form a feature vector. Based on this method, concatenate all the bag-of-words model encodings in the dataset to form a feature vector, and on this basis, divide the dataset into training data and test data. Exemplarily, the ratio of training data to test data can be 8:2.
[0093] In this embodiment, the performance of the trained classification model is evaluated using the test data. The performance includes precision, recall, F1-score, and accuracy. Specifically, to confirm the accuracy of the model, use the classification model to predict the test data, generate prediction labels, calculate and output the classification results, including precision, recall, F1-score, and accuracy, calculate and output the confusion matrix to show the accuracy and error distribution of the model prediction results.
[0094] As another possible embodiment, input multiple gene homology groups into the trained classification model to output the P450 gene identification results. Specifically, use the linear regression model, Bayesian model, extreme gradient boosting model, support vector machine model, and random forest model based on the natural language processing method of amino acid sequence features as base learners, train each of the base learners separately, and at the same time construct an extreme gradient boosting model as the meta-learner. Input the multiple gene homology groups into the multiple base learners respectively, use the output of the base learners as the input of the meta-learner for final prediction, and output the P450 gene identification results.
[0095] In this embodiment, the meta-learner uses Grid Search to try all possible parameter combinations to find the optimal solution for the training and optimization of the final ensemble model.
[0096] In an experimental example of the present invention, the dataset for testing includes multi-group P450 data collected manually and non-P450 data randomly extracted from the uniport database, with a total of 114,992 protein sequences. 2 / 10 of them are randomly selected to test the reliability and accuracy of the results. For the ensemble model composed of the base learner and the meta-learner, the classification accuracy evaluation results are shown in Table 2.
[0097] Table 2 Classification accuracy evaluation results of the ensemble model
[0098]
[0099] The accuracy of the model refers to the proportion of correct predictions made by the model on all samples. The 0.98 in Table 2 means that approximately 98% of the samples are correctly predicted in the corresponding total number of samples.
[0100] As Figure 2 shown, looking at the numbers on the diagonal, the ensemble model achieves accurate prediction for approximately 98% of the sequences in both categories.
[0101] AUC (Area Under the Curve) is the area under the ROC curve (Receiver Operating Characteristic Curve). The ROC curve represents the performance of a classifier by plotting the true positive rate (TPR, also known as recall or sensitivity) against the false positive rate (FPR) at different thresholds. As Figure 3 shown, the AUC values for both the UnP450 and P450 labels are 0.98, indicating that the model provided in this embodiment can distinguish all positive and negative classes.
[0102] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for automatic extraction and identification of P450 genes based on omics data, characterized in that: The method comprises: The amino acid sequence data of different omics were homogenized to obtain multiple gene homology groups; Obtain P450 and non-P450 amino acid sequence data and summarize them into a data set; Splitting the data set into different groups based on gene evolution clustering, obtaining group information, using the group information as an auxiliary feature for classification and identification, and adding it to the data set; Divide the data set into training data and test data, train the classification model with the training data, and evaluate the performance of the trained classification model with the test data; Multiple gene homology groups are input into the trained classification model to output the P450 gene identification results; The obtaining of P450 and non-P450 amino acid sequence data and summarizing them into a data set specifically includes: The acquired P450 and non-P450 amino acid sequence data were converted into text content, and the CountVectorizer tool for text feature extraction in Scikit-learn was used to extract phrases of single amino acid, two consecutive amino acids, and three consecutive amino acids according to ngram_range=(1, 3); Using regular expressions to identify each amino acid phrase, so that only amino acid phrases will be used as building blocks of n-grams; Generate a phrase-document matrix to record the occurrence of each n-gram in the text, and convert the amino acid phrase into a frequency representation using the bag-of-words encoding method; The method of dividing the data set into different groups based on gene evolution clustering to obtain grouping information is specifically as follows: processing the data set by orthofinder software to obtain clustering information based on sequence evolution, and designating data belonging to the same cluster in the clustering information as the same auxiliary group to obtain grouping information; The grouping information is represented in a hot single encoding manner; The grouping information is added to the data set, specifically, the hot unique encoding of the grouping information and the bag-of-words model encoding of the amino acid word frequency in the data set are concatenated into a feature vector.
2. The method for automatic extraction and identification of P450 genes based on omics data according to claim 1, characterized in that: The homogenization of different omics amino acid sequence data to obtain multiple gene homology groups specifically includes: Determine the properties of amino acid sequence data from different omics to distinguish whether the amino acid sequence data is protein data or nucleotide data; For nucleotide data, determine whether it is annotated nucleotide data, translate the annotated nucleotide data into protein data and perform normalization; For unannotated nucleotide data, it is further determined whether it is derived from eukaryotes or prokaryotes; Similar sequence fragments were extracted from unannotated nucleotide data from eukaryotes and prokaryotes, and gene structures were predicted for the similar sequence fragments to obtain potential P450 protein sequences.
3. The method for automatic extraction and identification of P450 genes based on omics data according to claim 2, characterized in that: Uncommon amino acids in potential P450 protein sequences and non-underline characters in fasta tags are identified, replaced and formatted. The uncommon amino acids are amino acids that do not belong to the standard 20 amino acids.
4. The method for automatic extraction and identification of P450 genes based on omics data according to claim 1, characterized in that: The P450 and non-P450 amino acid sequence data were initially cleaned and formatted before being converted into text content.
5. The method for automatic extraction and identification of P450 genes based on omics data according to claim 1, characterized in that: The performance of the trained classification model is evaluated by testing data, and the performance includes precision, recall, F1 score and accuracy.
6. The method for automatic extraction and identification of P450 genes based on omics data according to claim 1, characterized in that: Multiple gene homology groups are input into a trained classification model, and the P450 gene identification results are output, specifically including: using a linear regression model, a Bayesian model, an extreme gradient boosting model, a support vector machine model and a random forest model of a natural language processing method based on amino acid sequence features as base learners, training multiple base learners respectively, and constructing an extreme gradient boosting model as a meta-learner at the same time, inputting multiple gene homology groups into multiple base learners respectively, using the output of the base learner as the input of the meta-learner, making a final prediction, and outputting the P450 gene identification results.
Citation Information
Patent Citations
Construction method and device of promoter recognition system
CN104834834A
Method for predicting outer membrane proteins at bacterial whole genome level
CN105930687A
Acute myelogenous leukemia drug sensitivity related gene classifier constructed by machine learning algorithm
CN113555070A
P450 enzyme gene identification method and device, electronic equipment and storage medium
CN118335184A