Method for identifying horizontal transfer gene based on model error
The method enhances HGT gene detection by using multi-dimensional feature vectors and model prediction errors to automate the process, addressing accuracy and efficiency issues in existing HGT detection technologies.
Patent Information
- Application Number
- CN202510765331.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-10
AI Technical Summary
The prior art has insufficient accuracy, low efficiency and cumbersome preparation of training data sets in horizontal gene transfer detection, making it difficult to meet the needs of high-throughput screening.
By obtaining taxa protein data sets from public databases, performing data cleaning and structural prediction, extracting multi-dimensional feature vectors, constructing a classification prediction model, and using model prediction errors to identify potential horizontally transferred genes.
It significantly improves the detection accuracy and efficiency of horizontally transferred genes, simplifies the collection process of training data, reduces costs, and is suitable for large-scale applications.
Smart Images

Figure CN120319310A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gene identification, and particularly to a method for identifying horizontally transferred genes based on model error identification. Background Art
[0002] Horizontal gene transfer (HGT), also known as lateral gene transfer, refers to the process by which organisms acquire genes from other species through non-reproductive means. Different from traditional vertical gene transfer (genetic transmission from parent to offspring), HGT breaks the species barrier and enables genes to flow among distantly related taxa such as bacteria, archaea, and eukaryotes, and is particularly widespread in the microbial world. Its main mechanisms include transformation (direct uptake of environmental free DNA), conjugation (cell contact-mediated transfer of plasmids or chromosomal DNA), and transduction (virus-mediated gene transfer), which endow organisms with the ability to rapidly acquire new functional genes, thereby adapting to environmental changes or evolving new traits.
[0003] With the development of genomics technology, research methods for HGT have been continuously enriched. Early studies relied on genomic sequence alignment to infer HGT events by identifying highly similar genes between different species. However, this method is limited by problems such as misjudgment of short sequence conserved regions, reduced similarity caused by gene loss or rapid evolution, incomplete or misannotated databases, and ambiguity in the identification of genes shared by multiple species, and its accuracy needs to be improved. Phylogenetic analysis detects HGT by comparing the topological differences between gene trees and species trees. Although it has high accuracy, it requires the construction of complex phylogenetic trees, involving a large number of sequence alignments, tree shape optimization, and statistical tests (such as bootstrap resampling), with high computational complexity and time-consuming, making it difficult to handle large-scale genomic data. In addition, methods such as GC content and codon preference analysis, and genomic island detection focus on sequence feature differences or specific genomic regions, but single-feature analysis is easily affected by host genomic variation and requires comprehensive judgment by combining multi-dimensional evidence.
[0004] HGT research has both scientific significance and application value. In evolutionary biology, it reveals the key role of horizontal gene flow in species adaptive evolution and diversity formation in addition to vertical inheritance in biological evolution. For example, bacteria acquire antibiotic resistance genes through HGT to adapt to the drug environment. In the field of ecology, HGT helps microorganisms adapt to extreme environments (such as archaea acquiring heat-resistant genes) and participate in ecological function regulation (such as soil microorganisms degrading pollutants). In medicine, HGT is the core mechanism for the transmission of antibiotic resistance genes by multidrug-resistant bacteria (such as MRSA) and the acquisition of virulence factors by pathogenic bacteria (such as Vibrio cholerae), posing a threat to public health, and analyzing its mechanism can provide targets for developing anti-infection strategies. In biotechnology, HGT mechanisms are used for the directed transfer of foreign genes in genetic engineering and the functional enhancement of environmental remediation microorganisms.
[0005] In recent years, the introduction of machine learning technology has brought new breakthroughs to HGT research. It can integrate multi-dimensional features such as sequences, structures, and evolution, and automatically process massive data to improve detection efficiency. However, traditional machine learning methods rely on pre-annotated HGT and non-HGT gene datasets, and there are three core problems: 1. The accuracy of identification methods based on sequence similarity is insufficient: Factors such as short sequence conserved regions, gene evolutionary variations, database defects, and algorithm threshold selection are prone to lead to misjudgments, and manual verification or cross-comparison of multiple methods is required.
[0006] 2. Phylogenetic analysis is time-consuming and inefficient: The complex calculation process and high resource requirements limit its application in large-scale data and are difficult to meet the needs of high-throughput screening.
[0007] 3. The preparation of training datasets is cumbersome and subjective: The experimental verification cost of HGT genes is high, and data annotation depends on the experience of domain experts, which is prone to introduce biases. Moreover, data imbalance and noise problems affect the performance of the model.
[0008] In summary, the existing technologies have bottlenecks in the accuracy, efficiency, and data reliability of HGT detection. There is an urgent need for a new method that integrates multi-dimensional features, reduces manual intervention, and improves the degree of automation to break through the limitations of traditional technologies and promote the development of HGT research in basic science and practical applications. Summary of the Invention
[0009] In view of the above-mentioned existing technologies, the present invention aims to provide a method for identifying horizontally transferred genes based on model error, mainly to solve the technical problems existing in the above-mentioned background technology.
[0010] To achieve the above object, the technical solution of the embodiment of the present invention is realized as follows: A method for identifying horizontally transferred genes based on model error, the method includes the following steps: Obtain a taxonomic protein dataset from a public database and perform data cleaning on the taxonomic protein dataset, where the taxonomic protein dataset contains multiple protein sequences of different taxa; Perform structure prediction on each protein in the cleaned taxonomic protein dataset, and extract multi-dimensional feature vectors based on the prediction results; Construct a classification prediction model and train it using the above multi-dimensional feature vectors, and the output of the classification prediction model is the taxonomic label probability; Predict the protein structure in the protein sequence to be detected, extract multi-dimensional feature vectors of the same dimension as in the training stage based on the prediction results, and input them into the trained classification prediction model, and identify potential HGT genes according to the prediction error of the classification prediction model.
[0011] Optionally, data cleaning is performed on the group protein dataset, specifically including: Performing similarity comparison between the protein sequence to be compared and the group protein dataset after removing the protein sequence to be compared; Selecting the top 500 protein sequences with the highest similarity in the comparison results, and counting the number of relative groups derived from the protein sequence to be compared; If the proportion of the number of protein sequences from the relative group exceeds 50%, then the protein sequence to be compared is removed from the group protein dataset. The relative group is another group different from the group to which the protein sequence to be compared belongs, and the two groups are distantly related species groups; Continuing to perform similarity comparison on the remaining protein sequences in the group protein dataset, and removing the protein sequences that do not meet the requirements based on the similarity comparison results until all protein sequences have completed similarity comparison, generating a cleaned group protein dataset.
[0012] Optionally, for each protein sequence in the group protein dataset after data cleaning, a protein three-dimensional structure prediction software is used to predict the protein three-dimensional structure.
[0013] Optionally, for the structure prediction results, screening is performed based on the structure confidence score, which specifically includes: Obtaining the structure confidence score of each protein from the protein three-dimensional structure prediction software; Comparing the structure confidence score with a preset threshold, and the preset threshold is set according to the structure prediction accuracy requirements; Removing the structure data with a structure confidence score lower than the preset threshold, and retaining the valid structure data with a structure confidence score not lower than the preset threshold to form a high-confidence three-dimensional structure dataset.
[0014] Optionally, a multi-dimensional feature vector is extracted from the high-confidence three-dimensional structure dataset. The multi-dimensional feature vector includes primary structure features, secondary structure features, and tertiary structure features. The primary structure features include: the number of residues extracted, molecular weight, amino acid frequency, and the charge properties of amino acids. The secondary structure features include: the number and distribution of hydrogen bonds, the number, length, and position of α-helices, β-sheets, and random coils, and the backbone dihedral angles. The tertiary structure features include: the hydrophobicity score of residues, contact maps, the Euclidean distance matrix between residues, backbone curvature and torsion angles, the area of residues exposed to the solvent on the surface, the geometric shape of the protein, and the symmetry characteristics of the protein structure.
[0015] Optionally, construct a classification prediction model and train it using the above multi-dimensional feature vectors. Specifically, divide the multi-dimensional feature vectors into a training set and a test set, and input the training set into the constructed classification prediction model for training. The output result of the classification prediction model is the target taxon or the relative taxon.
[0016] Optionally, for the trained classification prediction model, verify the accuracy of the model using the protein structures of horizontally transferred genes selected manually. Specifically, it includes: Obtain the protein sequence of the HGT gene, and use protein three-dimensional structure prediction software to predict the three-dimensional structure of the protein sequence of the HGT gene; Extract multi-dimensional feature vectors including the primary structure features, secondary structure features, and tertiary structure features from the above prediction results; Input the extracted multi-dimensional feature vectors into the trained high-accuracy classification prediction model. If the classification prediction model makes a wrong prediction, it indicates that the classification model can identify HGT genes.
[0017] Identify potential HGT genes according to the prediction error of the classification prediction model. Specifically, for the protein sequence to be detected, determine its true taxon, and use the classification prediction model to obtain the predicted taxon. If the true taxon and the predicted taxon are inconsistent, the gene corresponding to the protein sequence to be detected is a potential HGT gene.
[0018] The beneficial effects of the present invention are as follows: By integrating protein atoms, one-dimensional, two-dimensional, and three-dimensional features, a 115-dimensional vector matrix is constructed to create a highly reliable classification model. When identifying potential HGT genes, innovatively using the model prediction error as the screening basis breaks through the traditional thinking. Compared with the sequence-based method, the model accuracy is significantly improved; compared with the phylogeny-based method, the calculation time is greatly shortened; compared with the traditional machine learning identification method, the dependence on HGT and non-HGT source genes is avoided, and the collection and application process of training data is greatly simplified. Finally, relying on the deep learning model to construct a classification prediction model not only greatly improves the accuracy and efficiency of HGT gene identification, but also lays a foundation for large-scale popularization and application with a simple process flow and low cost advantage. This method abandons the traditional complex procedures, effectively saves time and resource costs, and while ensuring high scientificity and practicality, also has economy and operability, providing strong technical support for gene science research and related industrial development. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic flowchart in an embodiment of the present application; Figure 2 It is a schematic principle diagram in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The technical solution of the present invention will be further described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. In the following description, the expression "some embodiments" describes a subset of all possible embodiments, but it should be understood that "some embodiments" may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0021] In the following description, a large number of specific details are given to provide a more thorough understanding of the present invention. However, it is obvious to those skilled in the art that the present invention can be implemented without one or more of these details. In other examples, in order to avoid confusion with the present invention, some technical features well known in the art are not described.
[0022] It should be understood that the present invention can be implemented in different forms and should not be construed as limited to the embodiments presented herein. On the contrary, providing these embodiments will make the disclosure thorough and complete and will fully convey the scope of the present invention to those skilled in the art. And the purpose of the terms used herein is only to describe specific embodiments and not to limit the present invention. When used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "comprising" and / or "including", when used in this specification, determine the presence of the described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups. When used herein, the term "and / or" includes any and all combinations of the related listed items.
[0023] Optional embodiments of the present invention are described in detail below. However, in addition to these detailed descriptions, the present invention may also have other embodiments.
[0024] Please refer to the attached Figures 1 to 2 , the present application provides a method for identifying horizontally transferred genes based on model error, and the method includes the following steps: S1. Obtain a taxonomic protein dataset from a public database and perform data cleaning on the taxonomic protein dataset, where the taxonomic protein dataset contains multiple protein sequences of different taxa; S2. Perform structure prediction on each protein in the cleaned taxonomic protein dataset and extract multi-dimensional feature vectors based on the prediction results; S3. Construct a classification prediction model and train it using the above-mentioned multi-dimensional feature vectors. The output of the classification prediction model is the probability of the taxonomic group label. S4. Predict the protein structure in the protein sequence to be detected. Based on the prediction result, extract the multi-dimensional feature vectors with the same dimension as in the training stage, and input them into the trained classification prediction model. Identify potential HGT genes according to the prediction error of the classification prediction model.
[0025] Specifically, for a method for identifying horizontally transferred genes based on model error provided by the present application, after obtaining a taxonomic group protein dataset containing multiple protein sequences of different taxonomic groups from a public database, data cleaning is used to exclude potential horizontally transferred genes and contaminated sequences to ensure that the data used for model training is pure non-horizontal transferred genes. For each protein sequence in the cleaned taxonomic group protein dataset, its three-dimensional structure is predicted by means of a protein three-dimensional structure prediction tool to generate structure data in a format such as PDB, and the data with a structure confidence level lower than a specific threshold is removed. Based on the predicted protein structure, protein atom, one-dimensional, two-dimensional, and three-dimensional features are extracted from multiple dimensions, and a multi-dimensional feature vector is formed by a 115-dimensional vector matrix. The constructed classification prediction model is trained using this multi-dimensional feature vector. During the training process, the model learns the characteristic differences of inherent genes of different taxonomic groups to construct a high-resolution classification boundary that can accurately distinguish each taxonomic group, so that the probability of the taxonomic group label output by the model can reliably reflect the taxonomic group to which the protein sequence belongs. In the actual application stage, for the protein sequence to be detected, first predict its protein structure, then extract the multi-dimensional feature vectors with the same dimension in the same method as in the training stage, input the feature vector into the trained classification prediction model. The model outputs the probability that the protein sequence to be detected belongs to each taxonomic group according to the taxonomic group characteristic pattern learned during the training process and gives the predicted taxonomic group label. By comparing the true taxonomic group label of the protein sequence to be detected, if the two are inconsistent, that is, a prediction error occurs, then the gene corresponding to the protein is identified as a potential horizontally transferred gene.
[0026] In some embodiments, the taxonomic group protein dataset contains heat shock protein sequences of eukaryotes and prokaryotes. Heat shock proteins are a class of molecular chaperone proteins that are highly conserved in evolution and are widely present in prokaryotes (such as bacteria and archaea) and eukaryotes (such as plants and animals), and participate in key physiological processes such as protein folding and stress response. Its cross-taxonomic group distribution characteristics make it an ideal object for HGT detection.
[0027] Furthermore, the classification prediction model constructed by the present application includes, but is not limited to, discrimination techniques based on traditional machine learning such as linear regression models, Bayesian models, extreme gradient boosting (XGBoost) models, support vector machine (SVM) models, and random forest models, and deep learning models based on these multi-dimensional amino acid features.
[0028] In some embodiments, data cleaning is performed on the group protein dataset, which specifically includes: Performing a similarity comparison between the protein sequence to be aligned and the group protein dataset after removing the protein sequence to be aligned; Selecting the top 500 protein sequences in terms of similarity ranking from the comparison results, and counting the number of relative groups from the protein sequence to be aligned; If the proportion of the number of protein sequences from the relative group exceeds 50%, then the protein sequence to be aligned is removed from the group protein dataset. The relative group is another group different from the group to which the protein sequence to be aligned belongs, and the two groups are distantly related species groups; Continuing to perform similarity comparison on the remaining protein sequences in the group protein dataset, and removing the protein sequences that do not meet the requirements based on the similarity comparison results until all protein sequences have completed similarity comparison, generating a cleaned group protein dataset.
[0029] Specifically, for each protein sequence to be aligned in the group protein dataset, perform a similarity comparison between it and the group protein dataset after removing itself. The similarity comparison can exclude self-comparison interference and focus on the true similarity distribution of the sequence with other group proteins. Through alignment tools such as BLAST, obtain the top 500 matching sequences according to the similarity level, and select a sufficient number of highly similar sequences to cover the main distribution of homologous genes, avoiding accidental misjudgment caused by a single match.
[0030] Perform a group source statistics on the top 500 matching sequences, identify the number of sequences belonging to the relative group of the protein sequence to be aligned. If the proportion of the relative group sequences exceeds 50%, it indicates that the similarity distribution of the protein sequence to be aligned shows a systematic cross-group bias, violating the dominant feature of the same group. It may be a horizontally transferred gene or a contaminated sequence, so it is removed from the dataset. For example, if the protein sequence to be aligned is a protein sequence of a eukaryote, but among the top 500 matching sequences obtained by comparison, the proportion of the protein sequences belonging to prokaryotes exceeds 50%, then the protein sequence to be aligned is removed from the group protein dataset.
[0031] After completing the screening of a single protein sequence, repeat the above comparison, statistics, and removal process for the remaining protein sequences in the group protein dataset to form a traversal screening mechanism. By traversing each sequence one by one and dynamically updating the dataset, ensure that the dataset for each comparison is the remaining data after removing the current protein sequence to be aligned, avoiding interference from the removed sequences on subsequent screening. Continuously loop and process until all protein sequences have completed similarity comparison and screening determination, and finally generate a cleaned group protein dataset.
[0032] In some embodiments, for each protein sequence in the clustered protein dataset after data cleaning, a protein three-dimensional structure prediction software is used to predict the protein three-dimensional structure. It should be noted that the selection of the protein three-dimensional structure prediction software and how to select it belong to the common knowledge of those skilled in the art, and will not be specifically elaborated in this embodiment.
[0033] In some embodiments, for the structure prediction results, screening is performed based on the structure confidence score, which specifically includes: Obtain the structure confidence scores of each protein from the protein three-dimensional structure prediction software; Compare the structure confidence score with a preset threshold, and the preset threshold is set according to the structure prediction accuracy requirement; Eliminate the structure data with a structure confidence score lower than the preset threshold, and retain the valid structure data with a structure confidence score not lower than the preset threshold to form a high-confidence three-dimensional structure dataset.
[0034] Specifically, extract the structure confidence scores of each protein from protein three-dimensional structure prediction software such as AlphaFold and EsmFold. This score is a quantitative evaluation of the reliability of the software's own structure prediction. For example, the structure confidence score (pLDDT value) output by AlphaFold ranges from 0 to 100 and reflects the structure confidence at the residue level. The structure confidence score can be achieved by parsing the metadata generated by the prediction software, such as PDB file annotations or independent confidence reports.
[0035] Compare the extracted structure confidence score with the preset threshold, which is set according to the structure prediction accuracy requirement of the specific application scenario. For example, in model training that requires high-precision structure features, the threshold can be set to 70 or 80. If the structure confidence score of a certain protein is lower than the preset threshold, it indicates that there are large uncertainties in its three-dimensional structure prediction results, such as domain folding errors, loop region conformational ambiguities, etc. Such data may carry incorrect spatial structure information and need to be eliminated from the dataset. By screening the confidence of each protein structure one by one, a dataset containing only highly reliable structure data is finally formed.
[0036] In some embodiments, a multi-dimensional feature vector is extracted from the high-confidence three-dimensional structure dataset. The multi-dimensional feature vector includes primary structure features, secondary structure features, and tertiary structure features. The primary structure features include: the number of residues extracted, molecular weight, amino acid frequency, and the charge property of amino acids. The secondary structure features include: the number and distribution of hydrogen bonds, the number, length, and position of α-helices, β-sheets, and random coils, and the backbone dihedral angles. The tertiary structure features include: the hydrophobicity score of residues, contact maps, the Euclidean distance matrix between residues, backbone curvature and torsion angles, the area of residue surfaces exposed to the solvent, the geometry of the protein, and the symmetry characteristics of the protein structure. The constructed multi-dimensional feature vector is a 115-dimensional feature vector, as shown in Table 1 below. Table 1
[0037] In some embodiments, a classification prediction model is constructed and trained using the above multi-dimensional feature vector. Specifically, the multi-dimensional feature vector is divided into a training set and a test set. The training set is input into the constructed classification prediction model for training. The output result of the classification prediction model is the target taxon or relative taxon.
[0038] Specifically, during the training process of inputting the training set into the constructed classification prediction model for training, each feature vector serves as an input sample, and the model outputs the prediction result of whether the sample belongs to the eukaryotic or prokaryotic taxon. For example, the output form is usually the probability values of the two taxa, such as eukaryotic probability 0.85 and prokaryotic probability 0.15, or the directly determined taxon label. The finally trained model can accurately map to the eukaryotic or prokaryotic taxon label based on the input multi-dimensional feature vector.
[0039] Further, after training is completed, the test set is used to evaluate the model, generating a confusion matrix and calculating metrics such as precision, recall, and F1-score. Exemplarily, the obtained confusion matrix is as follows:
[0040] The obtained metric results are shown in Table 2 below: Table 2
[0041] From the confusion matrix, among eukaryotic HSP genes, 2064 genes are actually eukaryotic and are correctly predicted, and only 24 genes are misjudged as prokaryotic. Among prokaryotic HSP genes, 1750 genes are actually prokaryotic and are correctly predicted, and 24 genes are misjudged as eukaryotic. Further analyzing the classification evaluation metrics, the precision, recall rate, and F1-score of eukaryotic and prokaryotic HSP genes all reach 0.99, the overall accuracy rate is 0.99, and the macro-average and weighted average also remain at 0.99, indicating that the classification effect of the model is excellent.
[0042] In some embodiments, for the trained classification prediction model, the accuracy of the model is verified by using the protein structures of horizontally transferred genes selected manually, specifically including: Obtain the protein sequences of HGT genes, and use protein three-dimensional structure prediction software to predict the three-dimensional structures of the protein sequences of HGT genes; Extract multi-dimensional feature vectors including the primary structure features, secondary structure features, and tertiary structure features from the aforementioned prediction results; Input the extracted multi-dimensional feature vectors into the trained classification prediction model. If the classification prediction model makes a wrong prediction, it indicates that the classification model can identify HGT genes.
[0043] Specifically, when verifying the accuracy of the trained classification prediction model, first, it is necessary to obtain the protein sequences of horizontally transferred genes (HGT genes) screened manually. Such sequences are usually strictly identified through phylogenetic analysis and experimental verification, such as detecting transposable elements or functional phenotype experiments, etc., to ensure that there is clear evidence of cross-species transfer between their source taxa and host taxa. Then, use protein three-dimensional structure prediction software to predict the three-dimensional structures of the protein sequences of the above HGT genes. For the structure prediction results, extract multi-dimensional feature vectors according to exactly the same rules as in the training stage, and input the extracted multi-dimensional feature vectors into the trained classification prediction model. The classification prediction model makes predictions based on the eukaryotic / prokaryotic inherent gene feature patterns learned in the training stage. Since all potential HGT genes have been removed from the training dataset during the cleaning stage, the model should have high accuracy in predicting non-HGT genes. However, the protein features of HGT genes belong to exogenous taxa, and their feature vectors will deviate from the typical distribution of the host taxa, resulting in inconsistent output results of the model with the true source taxa, that is, wrong predictions. If a large number of manually verified HGT genes all show wrong predictions and the misjudgment rate reaches the preset standard, it indicates that the accuracy of the trained classification prediction model meets the requirements.
[0044] Identifying potential HGT genes based on the prediction error of the classification prediction model, specifically including: for the protein sequence to be detected, determining its true taxon, and using the classification prediction model to obtain the predicted taxon. If the true taxon and the predicted taxon are inconsistent, the gene corresponding to the protein sequence to be detected is a potential HGT gene.
[0045] Specifically, for the protein sequence to be detected, following the same process as in the training stage, predicting the three-dimensional structure of the protein sequence to be detected and extracting multi-dimensional feature vectors, and inputting them into the trained classification prediction model. Based on the learned characteristic patterns of eukaryotic / prokaryotic native genes, the model outputs the predicted taxon label of the protein sequence. Then, when it is necessary to further determine the true taxon of the protein sequence to be detected, it can be achieved through two complementary methods: If the species to which the gene belongs is known, its true taxon can be directly determined based on the systematics of the species. For example, the protein sequences derived from archaea such as Escherichia coli and Bacillus subtilis have their true taxon belonging to "prokaryote"; the protein sequences derived from organisms such as humans, Saccharomyces cerevisiae, and Arabidopsis thaliana have their true taxon belonging to "eukaryote".
[0046] If the species information is unknown or further verification is required, similarity analysis is performed with known gene sequences in the public database through the homology alignment method. Using alignment tools such as BLAST, the protein sequence to be detected is aligned with the full-length sequences with taxon labels annotated in the database. If there is a gene with 100% homology in the alignment results, the gene corresponding to the protein sequence to be detected is directly classified into the taxon to which the homologous gene belongs. For example, if a protein sequence to be detected is 100% matched with the Escherichia coli protein sequence annotated as "prokaryote" in the UniProt database, its true taxon is determined to be "prokaryote"; if it is exactly the same as the human protein sequence, it belongs to "eukaryote".
[0047] Finally, comparing the predicted taxon with the true taxon: if the two are consistent, it indicates that the characteristics of the protein to be detected conform to the inherent pattern of its host taxon, and the gene corresponding to this protein is probably the host's own gene, that is, a non-HGT gene. If the two are inconsistent, that is, in the case of prediction error, it indicates that there are significant differences between the protein characteristics and the typical pattern of the host taxon, and it is more similar to the characteristic pattern of another taxon. Since potential HGT genes have been strictly removed from the training dataset during construction, the model has a very high prediction accuracy for non-HGT genes. Therefore, the prediction error can be reasonably inferred as an abnormal cross-taxon of gene characteristics. The most likely biological explanation for this abnormality is that the gene has acquired characteristics from an exogenous taxon through horizontal transfer, such as an HGT gene of prokaryotic origin showing the structural characteristics of prokaryotic proteins in a eukaryotic host, thus realizing the automatic identification of potential HGT genes.
[0048] In actual use, among the 21 HSP protein structure verification models selected manually, 20 were successfully determined to be HGT genes, and only 1 was not correctly determined. This indicates that the model can effectively identify HGT genes that deviate from the characteristics of eukaryotic / prokaryotic native genes, and has high reliability and practicability in practical applications. Although there are a small number of misjudgments, the overall determination ability is worthy of affirmation.
[0049] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered within the protection scope of the present invention. The protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for identifying horizontally transferred genes based on model error, characterized in that, The method includes the following steps: Obtain a taxon protein dataset from a public database and perform data cleaning on the taxon protein dataset, where the taxon protein dataset contains multiple protein sequences of different taxa; Perform structure prediction on each protein in the cleaned taxon protein dataset and extract multi-dimensional feature vectors based on the prediction results; Construct a classification prediction model and train it using the above multi-dimensional feature vectors, where the output of the classification prediction model is the taxon label probability; Predict the protein structure in the protein sequence to be detected, extract multi-dimensional feature vectors with the same dimension as in the training stage based on the prediction results, and input them into the trained classification prediction model, and identify potential HGT genes according to the prediction error of the classification prediction model.
2. The method for identifying horizontally transferred genes based on model error according to claim 1, wherein Perform data cleaning on the taxon protein dataset, specifically including: Perform similarity comparison between the protein sequence to be aligned and the taxon protein dataset after removing the protein sequence to be aligned; Select the top 500 protein sequences with the highest similarity in the comparison results and count the number of protein sequences from the relative taxon of the protein sequence to be aligned; If the proportion of the number of protein sequences from the relative taxon exceeds 50%, then remove the protein sequence to be aligned from the taxon protein dataset, where the relative taxon is another taxon different from the taxon to which the protein sequence to be aligned belongs, and the two taxa are distantly related species taxa; Continue to perform similarity comparison on the remaining protein sequences in the taxon protein dataset, and remove the protein sequences that do not meet the requirements based on the similarity comparison results until all protein sequences have completed the similarity comparison, generating a cleaned taxon protein dataset.
3. The method for identifying horizontally transferred genes based on model error according to claim 2, wherein For each protein sequence in the taxon protein dataset after data cleaning, use a protein three-dimensional structure prediction software to predict the three-dimensional structure of the protein.
4. A method for identifying horizontally transferred genes based on model error according to claim 3, characterized in that For the structure prediction results, perform screening based on the structure confidence score, which specifically includes: Obtain the structure confidence scores of each protein from the protein three-dimensional structure prediction software; Compare the structure confidence scores with a preset threshold, where the preset threshold is set according to the requirements of structure prediction accuracy; Remove the structure data with a structure confidence score lower than the preset threshold, and retain the valid structure data with a structure confidence score not lower than the preset threshold to form a high-confidence three-dimensional structure dataset.
5. A method for identifying horizontally transferred genes based on model error according to claim 4, characterized in that Extract multi-dimensional feature vectors from the high-confidence three-dimensional structure dataset, where the multi-dimensional feature vectors include primary structure features, secondary structure features, and tertiary structure features. The primary structure features include: the number of residues, molecular weight, amino acid frequency, and charge properties of amino acids. The secondary structure features include: the number and distribution of hydrogen bonds, the number, length, and position of α-helices, β-sheets, and random coils, and the backbone dihedral angles. The tertiary structure features include: the hydrophobicity score of residues, contact map, Euclidean distance matrix between residues, backbone curvature and torsion angles, the area of residues exposed to the solvent on the surface, the geometric shape of the protein, and the symmetry characteristics of the protein structure.
6. The method for identifying horizontally transferred genes based on model error according to claim 5, wherein, Construct a classification prediction model and train it using the above-mentioned multi-dimensional feature vectors. Specifically, divide the multi-dimensional feature vectors into a training set and a test set, input the training set into the constructed classification prediction model for training, and the output result of the classification prediction model is the target taxon or the relative taxon.
7. A method for identifying horizontally transferred genes based on model error according to claim 6, wherein, For the trained classification prediction model, verify the accuracy of the model using the protein structures of horizontally transferred genes selected manually. Specifically, it includes: Obtain the protein sequence of the HGT gene, and use protein three-dimensional structure prediction software to predict the three-dimensional structure of the protein sequence of the HGT gene; Extract the multi-dimensional feature vectors including the primary structure features, the secondary structure features, and the tertiary structure features from the above prediction results; Input the extracted multi-dimensional feature vectors into the trained classification prediction model. If the classification prediction model makes a wrong prediction, it indicates that the classification model can identify the HGT gene.
8. A method for identifying horizontally transferred genes based on model error according to claim 7, characterized in that, Identify potential HGT genes according to the prediction error of the classification prediction model. Specifically, for the protein sequence to be detected, determine its true taxon, and use the classification prediction model to obtain the predicted taxon. If the true taxon and the predicted taxon are inconsistent, the gene corresponding to the protein sequence to be detected is a potential HGT gene.
Citation Information
Patent Citations
Construction method of drug interaction prediction model, prediction method and device
CN114974408A
Hybrid prediction method for virulence factor and antibiotic resistance gene
CN115171792A
Track geometry small sample abnormal value classification method and device
CN118606851A
Genomic prediction of discordant gene-phenotype relationships between species
WO2025090647A1
KR20220060976A
Cited By
A method for screening of anti-mrsa active molecules based on xgboost algorithm, storage medium and equipment
CN122157783B