A method for identifying horizontally transferred genes based on model error

By obtaining taxon protein data sets from public databases, performing data cleaning and three-dimensional structure prediction, constructing a classification prediction model, and using model prediction errors to identify potential horizontal transfer genes, solving the problems of insufficient accuracy and inefficiency in the existing technology, and achieving efficient and economical HGT gene identification.

CN120319310BActive Publication Date: 2025-08-15SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510765331.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-08-15
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The prior art has insufficient accuracy, low efficiency and cumbersome data preparation in horizontal gene transfer detection, making it difficult to meet the needs of high-throughput screening.

Method used

By obtaining taxa protein data sets from public databases, performing data cleaning and three-dimensional structure prediction, extracting multi-dimensional feature vectors, constructing a classification prediction model, and using model prediction errors to identify potential horizontally transferred genes.

Benefits of technology

It significantly improves the accuracy and efficiency of HGT gene identification, simplifies the training data collection process, reduces costs, and is suitable for large-scale applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120319310B_ABST
    Figure CN120319310B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of gene identification, and in particular to a method for identifying horizontally transferred genes based on model error. The method comprises the following steps: obtaining a group protein dataset from a public database and performing data cleaning on the group protein dataset, wherein the group protein dataset contains multiple protein sequences from different groups; performing structure prediction on each protein in the cleaned group protein dataset, and extracting a multidimensional feature vector based on the prediction result; constructing a classification prediction model and training it using the multidimensional feature vector, wherein the output of the classification prediction model is a group label probability; predicting the protein structure in the protein sequence to be detected, extracting a multidimensional feature vector with the same dimension as that in the training stage based on the prediction result, and inputting the extracted multidimensional feature vector into the trained classification prediction model, and identifying potential HGT genes based on the prediction error of the classification prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gene identification, and in particular to a method for identifying horizontally transferred genes based on model errors. Background Art

[0002] Horizontal gene transfer (HGT), also known as lateral gene transfer, refers to the process by which organisms acquire genes from other species through non-reproductive means. Unlike traditional vertical gene transfer (genetic transmission from parent to offspring), HGT breaks down species boundaries, enabling gene flow between distantly related groups such as bacteria, archaea, and eukaryotes. It is particularly widespread in the microbial kingdom. Its main mechanisms include transformation (direct uptake of free environmental DNA), conjugation (transfer of plasmid or chromosomal DNA through cell-to-cell contact), and transduction (viral-mediated gene transfer). These mechanisms enable organisms to rapidly acquire new functional genes, enabling them to adapt to environmental changes or evolve new traits.

[0003] With the development of genomic technologies, HGT research methods have continued to expand. Early studies relied on genome sequence alignment to infer HGT events by identifying highly similar genes across species. However, this approach is limited by misidentification of short conserved regions, loss of similarity due to gene loss or rapid evolution, incomplete or incorrectly annotated databases, and ambiguous identification of genes shared by multiple species, leaving room for improvement in accuracy. Phylogenetic analysis detects HGT by comparing topological differences between gene trees and species trees. While highly accurate, it requires the construction of complex phylogenetic trees involving extensive sequence alignments, tree optimization, and statistical tests (such as bootstrap resampling). This is computationally complex and time-consuming, making it difficult to scale with large-scale genomic data. Furthermore, methods such as GC content and codon bias analysis and genomic island detection focus on differences in sequence features or specific genomic regions. However, single-feature analysis is susceptible to host genomic variation and requires comprehensive judgment based on multidimensional evidence.

[0004] HGT research has both scientific significance and practical value. In evolutionary biology, it reveals that in addition to vertical inheritance, horizontal gene flow plays a key role in the adaptive evolution and diversity formation of species. For example, bacteria acquire antibiotic resistance genes through HGT to adapt to the drug environment. In the field of ecology, HGT helps microorganisms adapt to extreme environments (such as archaea acquiring heat-resistant genes) and participates in the regulation of ecological functions (such as soil microorganisms degrading pollutants). In medicine, HGT is the core mechanism for the spread of antibiotic resistance genes by multidrug-resistant bacteria (such as MRSA) and the acquisition of virulence factors by pathogens (such as Vibrio cholerae), posing a threat to public health. Analyzing its mechanism can provide targets for the development of anti-infection strategies. In biotechnology, the HGT mechanism is used for the directed transfer of exogenous genes in genetic engineering and for the functional enhancement of environmental remediation microorganisms.

[0005] In recent years, the introduction of machine learning technology has brought new breakthroughs to HGT research. It can integrate multi-dimensional features such as sequence, structure, and evolution, and automatically process massive amounts of data to improve detection efficiency. However, traditional machine learning methods rely on pre-annotated HGT and non-HGT gene datasets, which have three core problems:

[0006] 1. The identification method based on sequence similarity is not accurate enough: factors such as short sequence conserved regions, gene evolutionary variation, database defects and algorithm threshold selection can easily lead to misjudgment, requiring manual verification or multi-method cross-comparison.

[0007] 2. Phylogenetic analysis is time-consuming and inefficient: complex computational processes and high resource requirements limit its application in large-scale data, making it difficult to meet high-throughput screening needs.

[0008] 3. Preparation of training datasets is cumbersome and subjective: Experimental verification of HGT genes is costly, data annotation relies on the experience of domain experts, and is prone to introducing biases. In addition, data imbalance and noise problems affect model performance.

[0009] In summary, existing technologies have bottlenecks in the accuracy, efficiency, and data reliability of HGT detection. A new method that integrates multi-dimensional features, reduces manual intervention, and improves the degree of automation is urgently needed to break through the limitations of traditional technologies and promote the development of HGT research in basic science and practical applications. Summary of the Invention

[0010] In view of the above-mentioned prior art, the present invention provides a method for identifying horizontally transferred genes based on model error, which mainly solves the technical problems existing in the above-mentioned background technology.

[0011] To achieve the above-mentioned purpose, the technical solution of the embodiment of the present invention is implemented as follows:

[0012] A method for identifying horizontally transferred genes based on model error, the method comprising the following steps:

[0013] Obtaining a group protein dataset from a public database and performing data cleaning on the group protein dataset, wherein the group protein dataset contains multiple protein sequences from different groups;

[0014] Perform structural prediction on each protein in the cleaned cluster protein dataset and extract multidimensional feature vectors based on the prediction results;

[0015] Constructing a classification prediction model and using the multidimensional feature vector for training, wherein the output of the classification prediction model is the probability of the class group label;

[0016] Predict the protein structure in the protein sequence to be tested, extract a multidimensional feature vector with the same dimension as the training phase based on the prediction results, and input it into the trained classification prediction model. Identify potential HGT genes based on the prediction error of the classification prediction model.

[0017] Optionally, data cleaning is performed on the protein dataset of the group, specifically including:

[0018] Performing a similarity comparison between the protein sequence to be compared and the protein dataset of the group after removing the protein sequence to be compared;

[0019] The top 500 protein sequences ranked by similarity were selected from the comparison results, and the number of relative groups derived from the protein sequences to be compared was counted;

[0020] If the number of protein sequences from the relative group accounts for more than 50%, the protein sequence to be compared will be removed from the protein dataset of the group. The relative group is another group different from the group to which the protein sequence to be compared belongs, and the two groups are distantly related species.

[0021] The remaining protein sequences in the cluster protein dataset are further compared for similarity. Based on the similarity comparison results, protein sequences that do not meet the requirements are eliminated until all protein sequences have completed the similarity comparison to generate a cleaned cluster protein dataset.

[0022] Optionally, for each protein sequence in the group protein data set after data cleaning, protein three-dimensional structure prediction software is used to predict the protein three-dimensional structure.

[0023] Optionally, the structure prediction results are screened based on the structure confidence score, which specifically includes:

[0024] Obtain the structural confidence score of each protein from the protein three-dimensional structure prediction software;

[0025] Comparing the structure confidence score with a preset threshold, wherein the preset threshold is set according to the structure prediction accuracy requirement;

[0026] Structural data with a structure confidence score lower than a preset threshold are eliminated, and valid structural data with a structure confidence score not lower than the preset threshold are retained to form a high-confidence three-dimensional structure dataset.

[0027] Optionally, a multidimensional feature vector is extracted from the high-confidence three-dimensional structure data set, wherein the multidimensional feature vector includes primary structure features, secondary structure features, and tertiary structure features, wherein the primary structure features include: the number of extracted residues, molecular weight, amino acid frequency, and charge properties of amino acids; the secondary structure features include: the number and distribution of hydrogen bonds, the number, length and position of α-helices, β-folds, and random coils, and the main chain dihedral angles; the tertiary structure features include: the hydrophobicity score of the residues, the contact map, the Euclidean distance matrix between residues, the main chain curvature and torsion angle, the area of the residue surface exposed to the solvent, the geometric shape of the protein, and the symmetry characteristics of the protein structure.

[0028] Optionally, a classification prediction model is constructed and trained using the above-mentioned multidimensional feature vector, specifically including dividing the multidimensional feature vector into a training set and a test set, inputting the training set into the constructed classification prediction model for training, and the output result of the classification prediction model is a target group or a relative group.

[0029] Optionally, for the trained classification prediction model, the accuracy of the model is verified using the protein structure of manually screened horizontally transferred genes, specifically including:

[0030] Obtain the protein sequence of the HGT gene and use protein three-dimensional structure prediction software to predict the protein three-dimensional structure of the protein sequence of the HGT gene;

[0031] Extracting a multidimensional feature vector including the primary structure feature, the secondary structure feature, and the tertiary structure feature from the aforementioned prediction results;

[0032] The extracted multidimensional feature vector is input into the trained high-accuracy classification prediction model. If the classification prediction model makes an error in prediction, it indicates that the classification model can identify HGT genes.

[0033] Identifying potential HGT genes based on the prediction error of the classification prediction model specifically includes: for a protein sequence to be detected, determining its true group, and using the classification prediction model to obtain a predicted group; if the true group and the predicted group are inconsistent, the gene corresponding to the protein sequence to be detected is a potential HGT gene.

[0034] The beneficial effects of the present invention are: by integrating protein atoms, one-dimensional, two-dimensional and three-dimensional features, a 115-dimensional vector matrix is constructed to create a highly reliable classification model. When identifying potential HGT genes, the model prediction error is innovatively used as the screening basis, breaking through the traditional thinking. Compared with the sequence-based method, the model accuracy is significantly improved; compared with the system evolution-based method, the calculation time is greatly shortened; compared with the traditional machine learning identification method, it avoids the dependence on HGT and non-HGT source genes, and greatly simplifies the collection and application process of training data. Finally, relying on the deep learning model to construct a classification prediction model, not only greatly improves the accuracy and efficiency of HGT gene identification, but also lays the foundation for large-scale promotion and application with its simple process flow and low cost advantages. This method abandons traditional complex procedures, effectively saves time and resource costs, and while ensuring a high degree of scientificity and practicality, it is both economical and operational, providing strong technical support for genetic science research and related industry development. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 This is a schematic diagram of the process in the embodiment of the present application;

[0036] Figure 2 This is a schematic diagram of the principles in an embodiment of the present application. DETAILED DESCRIPTION

[0037] The technical solution of the present invention is further elaborated in detail below in conjunction with the drawings and specific embodiments of the specification. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. In the following description, reference is made to "some embodiments", which describes a subset of all possible embodiments, but it should be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0038] In the following description, numerous specific details are provided to provide a more thorough understanding of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced without one or more of these details. In other instances, certain technical features well known in the art are not described to avoid confusion with the present invention.

[0039] It should be understood that the present invention can be implemented in different forms and should not be interpreted as being limited to the embodiments proposed herein. On the contrary, providing these embodiments will make the disclosure thorough and complete, and will fully convey the scope of the present invention to those skilled in the art. And the purpose of the terms used herein is only to describe specific embodiments and is not intended to limit the present invention. When used herein, the singular forms "one", "an" and "said / the" are also intended to include plural forms, unless the context clearly indicates another way. It should also be understood that the terms "comprising" and / or "comprising" when used in this specification determine the presence of the features, integers, steps, operations, elements and / or parts, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, parts and / or groups. When used herein, the term "and / or" includes any and all combinations of the relevant listed items.

[0040] Alternative embodiments of the present invention are described in detail below, however, in addition to these detailed descriptions, the present invention may also have other embodiments.

[0041] Please refer to the attached Figures 1 to 2 The present application provides a method for identifying horizontally transferred genes based on model error, the method comprising the following steps:

[0042] S1. Obtaining a group protein dataset from a public database and performing data cleaning on the group protein dataset, wherein the group protein dataset contains multiple protein sequences from different groups;

[0043] S2. Predict the structure of each protein in the cleaned cluster protein dataset and extract a multidimensional feature vector based on the prediction results;

[0044] S3, constructing a classification prediction model and using the above-mentioned multidimensional feature vector for training, wherein the output of the classification prediction model is the probability of the class group label;

[0045] S4. Predict the protein structure in the protein sequence to be tested, extract a multidimensional feature vector with the same dimension as the training phase based on the prediction results, and input it into the trained classification prediction model. Identify potential HGT genes based on the prediction error of the classification prediction model.

[0046] Specifically, the present application provides a method for identifying horizontally transferred genes based on model errors. After obtaining a group protein data set containing multiple protein sequences of different groups from a public database, data cleaning is used to exclude potential horizontally transferred genes and contaminating sequences to ensure that the data used for model training are pure non-horizontally transferred genes. For each protein sequence in the cleaned group protein data set, its three-dimensional structure is predicted with the help of a protein three-dimensional structure prediction tool to generate structural data in the PDB format, and data with a structural confidence lower than a specific threshold are removed. Based on the predicted protein structure, protein atoms, one-dimensional, two-dimensional and three-dimensional features are extracted from multiple dimensions, and a total of 115-dimensional vector matrices are formed into a multidimensional feature vector, and the multidimensional feature vector is used to train the constructed classification prediction model. During the training process, the model learns the characteristic differences of the inherent genes of different groups to construct a high-resolution classification boundary that can accurately distinguish each classification group, so that the group label probability output by the model can reliably reflect the group to which the protein sequence belongs;

[0047] In the actual application stage, for the protein sequence to be tested, its protein structure is first predicted, and then a multidimensional feature vector of the same dimension is extracted according to the same method as the training stage. The feature vector is input into the trained classification prediction model. The model outputs the probability that the protein sequence to be tested belongs to each classification group and gives a predicted group label based on the group feature pattern learned during the training process. By comparing the actual group label of the protein sequence to be tested, if the two are inconsistent, that is, a prediction error occurs, the gene corresponding to the protein is identified as a potential horizontal transfer gene.

[0048] In some embodiments, the cluster protein dataset includes sequences of heat shock proteins from both eukaryotic and prokaryotic organisms. Heat shock proteins are a class of evolutionarily highly conserved molecular chaperone proteins, widely present in prokaryotes (e.g., bacteria and archaea) and eukaryotes (e.g., plants and animals), involved in key physiological processes such as protein folding and stress response. Their cross-cluster distribution makes them ideal targets for HGT detection.

[0049] Furthermore, the classification prediction model constructed in this application includes but is not limited to identification techniques based on traditional machine learning such as linear regression models, Bayesian models, extreme gradient boosting (XGBoost) models, support vector machine (SVM) models and random forest models, and deep learning models based on the multidimensional features of these amino acids.

[0050] In some embodiments, data cleaning is performed on the protein dataset, specifically comprising:

[0051] Performing a similarity comparison between the protein sequence to be compared and the protein dataset of the group after removing the protein sequence to be compared;

[0052] The top 500 protein sequences ranked by similarity were selected from the comparison results, and the number of relative groups derived from the protein sequences to be compared was counted;

[0053] If the number of protein sequences from the relative group accounts for more than 50%, the protein sequence to be compared will be removed from the protein dataset of the group. The relative group is another group different from the group to which the protein sequence to be compared belongs, and the two groups are distantly related species.

[0054] The remaining protein sequences in the cluster protein dataset are further compared for similarity. Based on the similarity comparison results, protein sequences that do not meet the requirements are eliminated until all protein sequences have completed the similarity comparison to generate a cleaned cluster protein dataset.

[0055] Specifically, for each protein sequence to be compared in the cluster protein dataset, a similarity comparison is performed between it and the cluster protein dataset after removing itself. The similarity comparison can eliminate the interference of self-comparison and focus on the true similarity distribution between the sequence and other cluster proteins. Through comparison tools such as BLAST, the top 500 matching sequences are obtained according to the level of similarity, and a sufficient number of highly similar sequences are selected to cover the main distribution of homologous genes to avoid accidental misjudgment caused by a single match.

[0056] The group origin statistics are performed on the first 500 matching sequences to identify the number of sequences belonging to the relative group of the protein sequence to be compared. If the number of sequences in the relative group accounts for more than 50%, it indicates that the similarity distribution of the sequence to be compared shows a systematic cross-group bias, which violates the dominant characteristics of the same group. It may be a horizontally transferred gene or a contaminating sequence, so it is removed from the data set. For example, the protein sequence to be compared is a eukaryotic protein sequence, but among the first 500 matching sequences, the number of protein sequences belonging to prokaryotes accounts for more than 50%, then the protein sequence to be compared is removed from the group protein data set.

[0057] After completing the screening of a single protein sequence, the above alignment, statistics, and elimination process is repeated for the remaining protein sequences in the cluster protein dataset, forming a traversal screening mechanism. By traversing each sequence one by one and dynamically updating the dataset, it ensures that the dataset for each alignment is the remaining data after the current sequence to be aligned is eliminated, thus preventing interference from the eliminated sequences in subsequent screening. This loop process continues until all protein sequences have completed similarity alignment and screening, ultimately generating a cleaned cluster protein dataset.

[0058] In some embodiments, protein three-dimensional structure prediction software is used to predict the protein three-dimensional structure of each protein sequence in the group protein dataset after data cleaning. It should be noted that the selection of protein three-dimensional structure prediction software and how to select the software are common knowledge among those skilled in the art and are not specifically described in this example.

[0059] In some embodiments, the structure prediction results are screened based on the structure confidence score, which specifically includes:

[0060] Obtain the structural confidence score of each protein from the protein three-dimensional structure prediction software;

[0061] Comparing the structure confidence score with a preset threshold, wherein the preset threshold is set according to the structure prediction accuracy requirement;

[0062] Structural data with a structure confidence score lower than a preset threshold are eliminated, and valid structural data with a structure confidence score not lower than the preset threshold are retained to form a high-confidence three-dimensional structure dataset.

[0063] Specifically, the structural confidence score of each protein is extracted from protein three-dimensional structure prediction software such as AlphaFold and EsmFold. This score is a quantitative assessment of the reliability of the prediction software's own structure prediction. For example, the structural confidence score (pLDDT value) output by AlphaFold ranges from 0 to 100 and reflects the structural confidence at the residue level. The structural confidence score can be achieved by parsing the metadata generated by the prediction software, such as PDB file annotations or independent confidence reports.

[0064] The extracted structural confidence score is compared with a preset threshold. The preset threshold is set according to the structural prediction accuracy requirements of the specific application scenario. For example, in model training that requires high-precision structural features, the threshold can be set to 70 or 80. If the structural confidence score of a protein is lower than the preset threshold, it indicates that there is a large uncertainty in the three-dimensional structure prediction results, such as domain folding errors, ambiguous loop conformation, etc. Such data may carry erroneous spatial structural information and need to be eliminated from the dataset. By performing confidence screening on each protein structure one by one, a dataset containing only high-reliability structural data is finally formed.

[0065] In some embodiments, a multidimensional feature vector is extracted from the high-confidence three-dimensional structure dataset, wherein the multidimensional feature vector includes primary structure features, secondary structure features, and tertiary structure features. The primary structure features include: the number of extracted residues, molecular weight, amino acid frequency, and amino acid charge properties; the secondary structure features include: the number and distribution of hydrogen bonds, the number, length, and position of α-helices, β-sheets, and random coils, and main chain dihedral angles; the tertiary structure features include: hydrophobicity scores of residues, contact maps, Euclidean distance matrices between residues, main chain curvature and torsion angles, the area of the residue surface exposed to the solvent, the geometric shape of the protein, and the symmetry properties of the protein structure. The constructed multidimensional feature vector is a 115-dimensional feature vector, which is specifically shown in Table 1 below.

[0066] Table 1

[0067]

[0068] In some embodiments, a classification prediction model is constructed and trained using the above-mentioned multidimensional feature vector, specifically including dividing the multidimensional feature vector into a training set and a test set, inputting the training set into the constructed classification prediction model for training, and the output result of the classification prediction model is a target group or a relative group.

[0069] Specifically, during the training process of inputting the training set into the constructed classification prediction model for training, each feature vector is used as an input sample, and the model outputs the prediction result that the sample belongs to the eukaryotic or prokaryotic group. For example, the output form is usually the probability value of the two groups, such as the eukaryotic probability of 0.85 and the prokaryotic probability of 0.15, or the group label directly determined. The final trained model can accurately map to the eukaryotic or prokaryotic group label based on the input multi-dimensional feature vector.

[0070] Furthermore, after training is completed, the test set is used to evaluate the model, generate a confusion matrix, and calculate indicators such as precision, recall, and F1-score. For example, the obtained confusion matrix is as follows:

[0071]

[0072] The obtained index results are shown in Table 2 below:

[0073] Table 2

[0074]

[0075] The confusion matrix shows that 2,064 eukaryotic HSP genes were correctly predicted to be eukaryotic, with only 24 misclassified as prokaryotic. Similarly, 1,750 prokaryotic HSP genes were correctly predicted to be prokaryotic, with 24 misclassified as eukaryotic. Further analysis of the classification evaluation metrics revealed that the precision, recall, and F1-score for both eukaryotic and prokaryotic HSP genes reached 0.99, with an overall accuracy of 0.99. Both the macro-average and weighted average scores also remained at 0.99, indicating excellent classification performance.

[0076] In some embodiments, the trained classification prediction model is validated using the protein structure of manually screened horizontally transferred genes to verify the accuracy of the model, specifically including:

[0077] Obtain the protein sequence of the HGT gene and use protein three-dimensional structure prediction software to predict the protein three-dimensional structure of the protein sequence of the HGT gene;

[0078] Extracting a multidimensional feature vector including the primary structure feature, the secondary structure feature, and the tertiary structure feature from the aforementioned prediction results;

[0079] The extracted multidimensional feature vector is input into the trained classification prediction model. If the classification prediction model makes an error in prediction, it indicates that the classification model can identify HGT genes.

[0080] Specifically, when verifying the accuracy of the trained classification prediction model, it is first necessary to obtain the protein sequences of manually screened horizontally transferred genes (HGT genes). Such sequences are usually strictly identified through phylogenetic analysis and experimental verification, such as detection of transposon elements or functional phenotypic experiments, to ensure that there is clear evidence of cross-species transfer between their source groups and host groups. Then, protein three-dimensional structure prediction software is used to predict the three-dimensional structure of the protein sequences of the above-mentioned HGT genes. For the structure prediction results, multidimensional feature vectors are extracted according to rules that are exactly the same as those in the training phase, and the extracted multidimensional feature vectors are input into the trained classification prediction model. The classification prediction model makes predictions based on the eukaryotic / prokaryotic inherent gene feature patterns learned in the training phase. Since all potential HGT genes have been removed from the training dataset during the cleaning phase, the model's predictions for non-HGT genes should be highly accurate. However, the protein characteristics of HGT genes belong to exogenous groups, and their feature vectors will deviate from the typical distribution of the host group, resulting in inconsistencies between the model output results and the true source group, that is, prediction errors. If a large number of manually verified HGT genes all have prediction errors and the error rate reaches the preset standard, it indicates that the accuracy of the trained classification prediction model meets the requirements.

[0081] Identifying potential HGT genes based on the prediction error of the classification prediction model specifically includes: for a protein sequence to be detected, determining its true group, and using the classification prediction model to obtain a predicted group; if the true group and the predicted group are inconsistent, the gene corresponding to the protein sequence to be detected is a potential HGT gene.

[0082] Specifically, the same process as the training phase is followed for the protein sequence to be tested: a three-dimensional structure prediction is performed and a multidimensional feature vector is extracted, which is then fed into the trained classification prediction model. Based on the learned characteristic patterns of eukaryotic / prokaryotic intrinsic genes, the model outputs a predicted cluster label for the protein sequence. Further determination of the true cluster of the protein sequence to be tested can be achieved through two complementary approaches:

[0083] If the species to which the gene belongs is known, its true group can be directly determined based on the systematic taxonomy of the species. For example, the protein sequence derived from archaea such as Escherichia coli and Bacillus subtilis has a true group classification of "prokaryotic"; the protein sequence derived from organisms such as humans, Saccharomyces cerevisiae, and Arabidopsis thaliana has a true group classification of "eukaryotic".

[0084] If species information is unknown or requires further verification, similarity analysis is performed with known gene sequences in public databases using homology comparison methods. Using alignment tools such as BLAST, the protein sequence to be tested is compared with full-length sequences in the database that have been annotated with group labels. If there are genes with 100% homology in the comparison results, the gene corresponding to the protein sequence to be tested is directly classified into the group to which the homologous gene belongs. For example, if a sequence to be tested matches 100% with an Escherichia coli protein sequence annotated as "prokaryotic" in the UniProt database, its true group is determined to be "prokaryotic"; if it is completely consistent with the human protein sequence, it is classified as "eukaryotic".

[0085] Finally, the predicted cluster is compared with the true cluster. If the two match, the characteristics of the protein being tested conform to the inherent pattern of its host cluster, and the gene corresponding to the protein is most likely a native host gene, that is, a non-HGT gene. If the two do not match, indicating a prediction error, the protein's characteristics differ significantly from the typical pattern of the host cluster and are closer to the characteristic pattern of another cluster. Because the training dataset was rigorously cleaned to remove potential HGT genes during construction, the model's prediction accuracy for non-HGT genes is extremely high. Therefore, the prediction error can be reasonably inferred to be a cross-cluster anomaly in the gene's characteristics. The most likely biological explanation for this anomaly is that the gene was acquired from an exogenous cluster through horizontal transfer, thus carrying the characteristic signal of the exogenous cluster. For example, a prokaryotic HGT gene exhibits the structural characteristics of a prokaryotic protein in a eukaryotic host, thus enabling automated identification of potential HGT genes.

[0086] In actual use, the model was validated using 21 manually screened HSP protein structures, successfully identifying 20 as HGT genes, with only one incorrectly identified. This demonstrates that the model can effectively identify HGT genes that deviate from intrinsic eukaryotic / prokaryotic gene characteristics, demonstrating high reliability and practicality in practical applications. Despite a small number of false positives, its overall identification capability is commendable.

[0087] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. The scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A method for identifying horizontally transferred genes based on model error, characterized in that: The method comprises the following steps: Obtaining a group protein dataset from a public database and performing data cleaning on the group protein dataset, wherein the group protein dataset contains multiple protein sequences from different groups; Perform structure prediction on each protein sequence in the cleaned cluster protein dataset and extract multidimensional feature vectors based on the prediction results; Constructing a classification prediction model and using the multidimensional feature vector for training, wherein the output of the classification prediction model is the probability of the class group label; Predict the protein structure in the protein sequence to be tested, extract the multidimensional feature vector with the same dimension as the training stage based on the prediction results, and input it into the trained classification prediction model. Identify potential horizontal transfer genes based on the prediction error of the classification prediction model.

2. The method for identifying horizontally transferred genes based on model error according to claim 1, characterized in that: Data cleaning is performed on the protein dataset of the group, specifically including: Performing a similarity comparison between the protein sequence to be compared and the protein dataset of the group after removing the protein sequence to be compared; The top 500 protein sequences ranked by similarity were selected from the comparison results, and the number of relative groups derived from the protein sequences to be compared was counted; If the number of protein sequences from the relative group accounts for more than 50%, the protein sequence to be compared will be removed from the protein dataset of the group. The relative group is another group different from the group to which the protein sequence to be compared belongs, and the two groups are distantly related species. The remaining protein sequences in the cluster protein dataset are further compared for similarity. Based on the similarity comparison results, protein sequences that do not meet the requirements are eliminated until all protein sequences have completed the similarity comparison to generate a cleaned cluster protein dataset.

3. The method for identifying horizontally transferred genes based on model error according to claim 2, wherein: For each protein sequence in the group protein data set after data cleaning, the protein three-dimensional structure prediction software is used to predict the protein three-dimensional structure.

4. The method for identifying horizontally transferred genes based on model error according to claim 3, wherein: For the structure prediction results, screening is performed based on the structure confidence score, which specifically includes: Obtain the structural confidence score of each protein from the protein three-dimensional structure prediction software; Comparing the structure confidence score with a preset threshold, wherein the preset threshold is set according to the structure prediction accuracy requirement; Structural data with a structure confidence score lower than a preset threshold are eliminated, and valid structural data with a structure confidence score not lower than the preset threshold are retained to form a high-confidence three-dimensional structure dataset.

5. The method for identifying horizontally transferred genes based on model error according to claim 4, characterized in that: A multidimensional feature vector is extracted from the high-confidence three-dimensional structure data set, wherein the multidimensional feature vector includes primary structure features, secondary structure features, and tertiary structure features. The primary structure features include: the number of extracted residues, molecular weight, amino acid frequency, and charge properties of amino acids; the secondary structure features include: the number and distribution of hydrogen bonds, the number, length and position of α-helices, β-sheets, and random coils, and the main chain dihedral angles; the tertiary structure features include: the hydrophobicity score of the residues, the contact map, the Euclidean distance matrix between residues, the main chain curvature and torsion angle, the area of the residue surface exposed to the solvent, the geometric shape of the protein, and the symmetry properties of the protein structure.

6. The method for identifying horizontally transferred genes based on model error according to claim 5, characterized in that: Constructing a classification prediction model and using the above-mentioned multidimensional feature vector for training, specifically including dividing the multidimensional feature vector into a training set and a test set, inputting the training set into the constructed classification prediction model for training, and the output result of the classification prediction model is the target group or the relative group.

7. The method for identifying horizontally transferred genes based on model error according to claim 6, characterized in that: For the trained classification prediction model, the accuracy of the model is verified by using the protein structure of manually screened horizontally transferred genes, specifically including: Obtaining the protein sequence of the horizontally transferred gene, and using protein three-dimensional structure prediction software to predict the protein three-dimensional structure of the protein sequence of the horizontally transferred gene; Extracting a multidimensional feature vector including the primary structure feature, the secondary structure feature, and the tertiary structure feature from the aforementioned prediction results; The extracted multidimensional feature vector is input into the trained classification prediction model. If the classification prediction model makes an error in prediction, it indicates that the classification prediction model can identify horizontally transferred genes.

8. The method for identifying horizontally transferred genes based on model error according to claim 7, characterized in that: Identifying potential horizontal transfer genes based on the prediction error of the classification prediction model specifically includes: for the protein sequence to be detected, determining its true group, and using the classification prediction model to obtain the predicted group; if the true group and the predicted group are inconsistent, the gene corresponding to the protein sequence to be detected is a potential horizontal transfer gene.

Citation Information

Patent Citations

  • Construction method of drug interaction prediction model, prediction method and device

    CN114974408A

  • Hybrid prediction method for virulence factor and antibiotic resistance gene

    CN115171792A