Protein toxicity prediction method and system based on sequence information

By combining a graph attention-based neural network model and a dual-channel graph attention network with protein sequence embedding technology, we solved the problems of feature extraction relying on manual design and insufficient modeling of long-range dependency relationships in existing technologies, and achieved efficient and accurate prediction of protein toxicity.

CN119418780BActive Publication Date: 2025-10-03SUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510014140.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-10-03
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

The existing technologies for protein toxicity prediction have the following problems: feature extraction relies on manual design and is unable to automatically capture complex features. The model is not capable of modeling long-range sequence dependencies and has insufficient generalization capabilities for unknown protein sequences and mutants, resulting in insufficient prediction accuracy and limited generalization capabilities.

Method used

A graph attention-based neural network model is used to obtain multiple protein sequences from the protein database, calculate six categories of feature vectors, perform dimensionality reduction screening, combine pre-trained language models with protein sequence embedding technology, extract multi-level and deep features, and use a dual-channel graph attention network and gated recurrent unit (GRU) module to capture local and long-range dependencies for toxicity prediction.

Benefits of technology

It improves the accuracy and generalization ability of protein toxicity prediction and can adapt to different types of protein sequences. In particular, it can still make accurate predictions when faced with new sequences and variants, which is significantly better than existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119418780B_ABST
    Figure CN119418780B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for predicting protein toxicity based on sequence information. The method comprises obtaining multiple protein sequences from a protein database; performing feature calculation on all protein sequences to obtain protein feature vectors for six categories; concatenating the six feature vectors by column to obtain a first feature vector; performing dimensionality reduction screening on the first feature vector to obtain a second feature vector of the target dimension; using the second feature vector to train a graph attention-based neural network model, and using the trained neural network model as a protein toxicity prediction model; inputting a new protein sequence into the protein toxicity prediction model to obtain a toxicity prediction result output by the protein toxicity prediction model. The present invention is adaptable to different types of protein sequences and accurately predicts protein toxicity by learning common sequence patterns.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of protein biology, and in particular relates to a protein toxicity prediction method and system based on sequence information. Background Art

[0002] Traditional methods for predicting protein toxicity rely primarily on biological experiments and structural analysis, assessing toxicity by studying the physicochemical properties or three-dimensional structure of proteins. However, these methods have numerous drawbacks. They often require a large amount of experimental data and resources, resulting in high experimental costs and lengthy processing times. Furthermore, they are highly dependent on experimental conditions, making it difficult to quickly and comprehensively predict the toxicity of a large number of unknown proteins. Therefore, the efficient and accurate prediction of toxicity using computational methods has become a key research topic.

[0003] With the continuous expansion of biological sequence databases, researchers have gradually shifted their focus from structural information to the direct use of sequence information. Sequence-based analysis methods can break through the limitations of obtaining protein structural data and use the information contained in the sequence itself to infer the biological function and toxicity of proteins. Initially, sequence-based prediction methods mainly used sequence alignment tools, such as BLAST, which was based on the principle of predicting toxicity by looking for sequence similarities with known toxic proteins. However, this sequence alignment method can only find similar sequences in known databases. Once faced with new sequences, variants, or proteins that have not been studied, its predictive ability will drop significantly.

[0004] With the development of deep learning technology, particularly the successful application of convolutional neural networks (CNNs) and recurrent neural networks (RNNs), researchers have begun exploring how to use deep learning models to automatically extract potential features from protein sequences. CNNs are widely used in image processing and excel at capturing local patterns. When applied to protein sequences, they can effectively identify local patterns within amino acid sequences, such as the signatures of specific sequence segments or motifs that influence toxicity. RNNs, particularly their variants such as long short-term memory (LSTM) and gated recurrent units (GRU), have been applied to protein sequence toxicity prediction due to their superior processing capabilities for sequence data. They can capture long-range dependencies within sequences, thereby better understanding the complex relationship between sequence information and toxicity.

[0005] Despite this, current methods using classical machine learning and traditional deep learning still have some shortcomings in protein toxicity prediction:

[0006] Feature extraction relies on manual design and is unable to automatically capture complex features: In classic machine learning methods, feature extraction often relies on manually designed features such as amino acid composition, polarity, and hydrophobicity. These manually extracted features often only reflect surface information about protein sequences and struggle to capture deeper sequence characteristics. In particular, protein toxicity may be determined by subtle changes in the sequence, local structure, or long-range interactions. Due to the inability to automatically learn complex features from data, these methods exhibit limitations when faced with underlying patterns in sequences, resulting in insufficient prediction accuracy. Feature design often requires considerable expertise and time, lacking flexibility and versatility.

[0007] The model's ability to model long-range sequence dependencies is insufficient: Traditional deep learning methods, such as convolutional neural networks (CNNs), while advantageous in capturing local features of protein sequences, are inadequate when it comes to handling long-range dependencies within sequences. Protein toxicity often depends not only on local regions within the sequence but also on interactions between long-range amino acid residues. Because CNNs can only process a limited receptive field and are unable to effectively model the global information of the entire sequence, they struggle to fully capture the long-range dependency patterns within the sequence that influence toxicity, hindering accurate predictions for complex proteins.

[0008] Insufficient generalization for unknown protein sequences and mutants: Classical machine learning and traditional deep learning methods typically rely on existing protein databases during training, and the model's generalization ability is often limited to the scope of the training data. When faced with new sequences or mutated proteins not seen in the training set, the model's prediction performance is often unsatisfactory. This is because traditional models cannot fully learn the universal patterns implicit in protein sequences, making them difficult to handle novel sequences or rare mutants outside the database, and unable to provide accurate toxicity predictions for these sequences. This lack of generalization limits the model's widespread use in practical biological research and applications. Summary of the Invention

[0009] In view of the deficiencies in the prior art, the present invention provides a method and system for predicting protein toxicity based on sequence information.

[0010] In a first aspect, the present invention provides a method for predicting protein toxicity based on sequence information, characterized by comprising:

[0011] Acquire multiple protein sequences from a protein database; wherein the multiple protein sequences include toxic and non-toxic protein sequences;

[0012] The features of all protein sequences were calculated to obtain the feature vectors of six major categories of proteins;

[0013] Concatenate the six categories of eigenvectors by column to obtain the first eigenvector;

[0014] Perform dimensionality reduction screening on the first eigenvector to obtain the second eigenvector of the target dimension;

[0015] The second eigenvector is used to train a graph attention-based neural network model, and the trained neural network model is used as a protein toxicity prediction model;

[0016] The new protein sequence is input into the protein toxicity prediction model to obtain the toxicity prediction result output by the protein toxicity prediction model.

[0017] Optionally, the feature calculation is performed on all protein sequences to obtain six categories of protein feature vectors, including:

[0018] Determine the feature vector of protein R;

[0019] determining eigenvectors of a position-specific scoring matrix based on evolutionary features;

[0020] Determine the feature vector of amino acid index based on physicochemical properties;

[0021] Determine the eigenvector of amino acid group frequency counts;

[0022] A feature vector that determines the length of the protein sequence;

[0023] Determine the feature vector of the common protein segment sequence.

[0024] Optionally, determining the characteristic vector of protein R includes:

[0025] The expression for constructing the characteristic vector AAC of amino acid composition is:

[0026] ;

[0027] Among them, p(a i ) is the i-th amino acid a in the sequence i The frequency of occurrence of f(a i ) is the i-th amino acid a in the sequence i The number of occurrences; i=1,2,...,20; N is the total length of the sequence;

[0028] And / or, construct the expression of the dipeptide composition feature vector DPC:

[0029] ;

[0030] Among them, p(d i,j ) is the dipeptide d in the sequence i,j The frequency of occurrence of f(d i,j) is the dipeptide d in the sequence i,j The number of occurrences of; j = 1, 2, ..., 20;

[0031] And / or, construct an expression for the eigenvector MB of the Moreau-Broto autocorrelation:

[0032] ;

[0033] Among them, R d is the Moreau-Broto autocorrelation value of the sequence with delay d; d = 1, 2, ..., D; D represents the maximum delay; p i is the physicochemical property value of the i-th amino acid in the sequence; p i+d is the physicochemical property value of the i+dth amino acid in the sequence;

[0034] And / or, construct the expression for the eigenvector MORAN of the Moran autocorrelation:

[0035] ;

[0036] Among them, M d is the Moran autocorrelation value of the sequence with delay d; is the mean of all physicochemical property values;

[0037] And / or, construct the expression for the eigenvector GEARY of the Geary autocorrelation:

[0038] ;

[0039] Among them, G d is the Geary autocorrelation value of the sequence with delay d;

[0040] And / or, construct the expression of the characteristic vector CTD of the combined transformation distribution:

[0041] ;

[0042] Among them, count(A i' ) is the i'th type amino acid A i' The number of occurrences of Indicates the i'th amino acid A in the sequence i' and the j'th amino acid A j' The number of adjacencies; Representation captures the spatial distribution characteristics of each amino acid class; An operator representing the concatenation vector;

[0043] And / or, construct the expression for the eigenvector CTRIAD of the joint triples:

[0044] ;

[0045] in, Represents a triple The normalized frequency of occurrence of Represents a triple in a sequence The number of occurrences of g i is the group number of the i-th amino acid in the sequence; g j is the group number of the jth amino acid in the sequence; is the group number of the k'th amino acid in the sequence;

[0046] And / or, construct the expression for the eigenvector SOCN of the sequence order coupling number:

[0047] ;

[0048] in, represents the number of sequential couplings of delay d in the sequence on the physicochemical property k''; is the value of the physicochemical property k'' of the i-th amino acid; Represents The amino acid property value of delay d; Represents the global variance of attribute values, used for normalization; is the mean value of the physicochemical property k''; K is the total number of physicochemical properties;

[0049] And / or, construct the expression of the eigenvector QSO of the quasi-sequential order descriptor:

[0050] ;

[0051] in, represents the coupling characteristics of the delay d in the sequence on the physicochemical property k'';

[0052] And / or, construct the expression of the pseudo amino acid composition feature vector PAAC:

[0053] ;

[0054] in, represents the value of the physicochemical property k'' representing the autocorrelation part of the series with delay d.

[0055] Optionally, determining the eigenvector of the position-specific scoring matrix based on the evolutionary feature comprises:

[0056] The corrected probability of amino acid a at position i'' in the sequence is calculated according to the following formula :

[0057] ;

[0058] in, represents the frequency of occurrence of amino acid a at position i'' in the sequence; represents the frequency of occurrence of amino acid b at position i'' in the sequence; A is the set of natural amino acids;

[0059] The score of position i'' and amino acid a in the sequence is calculated according to the following formula :

[0060] ;

[0061] Where λ is a scaling factor used to adjust the scale of the score; q(a) is the average frequency of amino acid a in nature;

[0062] The average score p[a] of amino acid a was calculated according to the following formula:

[0063] ;

[0064] Where N is the total length of the sequence.

[0065] Optionally, performing dimensionality reduction screening on the first eigenvector to obtain a second eigenvector of a target dimension includes:

[0066] Delete the columns with a standard deviation of 0 in the first eigenvector to obtain the third eigenvector;

[0067] Determine the Pearson correlation coefficient of each feature with the other features in the third eigenvector;

[0068] Randomly delete one of the two features whose Pearson correlation coefficient is greater than the correlation coefficient threshold to obtain the fourth eigenvector;

[0069] constructing a co-selection relationship graph between features in the fourth eigenvector to identify features associated with protein toxicity;

[0070] The features with the largest target number of weights are selected as representatives and passed to recursive feature elimination based on cross-validation to obtain the second feature vector of the target dimension.

[0071] In a second aspect, the present invention provides a protein toxicity prediction system based on sequence information, comprising:

[0072] A sequence acquisition module is used to acquire multiple protein sequences from a protein database; wherein the multiple protein sequences include toxic and non-toxic protein sequences;

[0073] Feature calculation module, used to calculate the features of all protein sequences and obtain the feature vectors of six major categories of proteins;

[0074] The vector concatenation module is used to concatenate the six categories of eigenvectors by column to obtain the first eigenvector;

[0075] A dimensionality reduction and screening module is used to perform dimensionality reduction and screening on the first eigenvector to obtain a second eigenvector of the target dimension;

[0076] A model training module, for training a graph attention-based neural network model using the second eigenvector, and using the trained neural network model as a protein toxicity prediction model;

[0077] The toxicity prediction module is used to input a new protein sequence into the protein toxicity prediction model to obtain a toxicity prediction result output by the protein toxicity prediction model.

[0078] Optionally, the feature calculation module includes:

[0079] A first determining unit is used to determine a characteristic vector of protein R;

[0080] a second determining unit, configured to determine a feature vector of a position-specific scoring matrix based on evolutionary features;

[0081] a third determining unit, configured to determine a feature vector of an amino acid index based on physicochemical properties;

[0082] a fourth determining unit, configured to determine a characteristic vector of the amino acid group frequency count;

[0083] a fifth determining unit, for determining a feature vector of a protein sequence length;

[0084] The sixth determining unit is used to determine the feature vector of the protein common fragment sequence.

[0085] Optionally, the first determining unit includes:

[0086] The first construction means is used to construct the expression of the amino acid composition feature vector AAC:

[0087] ;

[0088] Among them, p(a i ) is the i-th amino acid a in the sequence i The frequency of occurrence of f(a i ) is the i-th amino acid a in the sequence i The number of occurrences; i=1,2,...,20; N is the total length of the sequence;

[0089] And / or, a second constructing means for constructing an expression of a dipeptide composition characteristic vector DPC:

[0090] ;

[0091] Among them, p(d i,j ) is the dipeptide d in the sequence i,j The frequency of occurrence of f(d i,j ) is the dipeptide d in the sequence i,j The number of occurrences of; j = 1, 2, ..., 20;

[0092] And / or, a third constructing means for constructing an expression for the eigenvector MB of the Moreau-Broto autocorrelation:

[0093] ;

[0094] Among them, R d is the Moreau-Broto autocorrelation value of the sequence with delay d; d = 1, 2, ..., D; D represents the maximum delay; p i is the physicochemical property value of the i-th amino acid in the sequence; p i+d is the physicochemical property value of the i+dth amino acid in the sequence;

[0095] And / or, the fourth constructing means is used to construct an expression of the eigenvector MORAN of the Moran autocorrelation:

[0096] ;

[0097] Among them, M d is the Moran autocorrelation value of the sequence with delay d; is the mean of all physicochemical property values;

[0098] And / or, a fifth constructing means for constructing an expression for the eigenvector GEARY of the Geary autocorrelation:

[0099] ;

[0100] Among them, G d is the Geary autocorrelation value of the sequence with delay d;

[0101] And / or, a sixth constructing means for constructing an expression of a characteristic vector CTD of a combined transformation distribution:

[0102] ;

[0103] Among them, count(A i' ) is the i'th type amino acid A i' The number of occurrences of Indicates the i'th amino acid A in the sequence i' and the j'th amino acid A j'The number of adjacencies; Representation captures the spatial distribution characteristics of each amino acid class; An operator representing the concatenation vector;

[0104] And / or, a seventh constructing means for constructing an expression of a characteristic vector CTRIAD of a joint triple:

[0105] ;

[0106] in, Represents a triple The normalized frequency of occurrence of Represents a triple in a sequence The number of occurrences of g i is the group number of the i-th amino acid in the sequence; g j is the group number of the jth amino acid in the sequence; is the group number of the k'th amino acid in the sequence;

[0107] And / or, an eighth constructing means for constructing an expression of a characteristic vector SOCN of a sequence order coupling number:

[0108] ;

[0109] in, represents the number of sequential couplings of delay d in the sequence on the physicochemical property k''; is the value of the physicochemical property k'' of the i-th amino acid; Represents The amino acid property value of delay d; Represents the global variance of attribute values, used for normalization; is the mean value of the physicochemical property k''; K is the total number of physicochemical properties;

[0110] And / or, the ninth constructing means is used to construct the expression of the characteristic vector QSO of the quasi-sequential order descriptor:

[0111] ;

[0112] in, represents the coupling characteristics of the delay d in the sequence on the physicochemical property k'';

[0113] And / or, a tenth constructing means for constructing an expression of a pseudo amino acid composition feature vector PAAC:

[0114] ;

[0115] in, represents the value of the physicochemical property k'' representing the autocorrelation part of the series with delay d.

[0116] Optionally, the second determining unit includes:

[0117] The first calculation means is used to calculate the corrected probability of amino acid a at position i'' in the sequence according to the following formula :

[0118] ;

[0119] in, represents the frequency of occurrence of amino acid a at position i'' in the sequence; represents the frequency of occurrence of amino acid b at position i'' in the sequence; A is the set of natural amino acids;

[0120] The second calculation device is used to calculate the score of position i'' and amino acid a in the sequence according to the following formula :

[0121] ;

[0122] Where λ is a scaling factor used to adjust the scale of the score; q(a) is the average frequency of amino acid a in nature;

[0123] The third calculating means is used to calculate the average score p[a] of amino acid a according to the following formula:

[0124] ;

[0125] Where N is the total length of the sequence.

[0126] Optionally, the dimensionality reduction screening module includes:

[0127] A first deletion unit is used to delete the columns with a standard deviation of 0 in the first eigenvector to obtain a third eigenvector;

[0128] a determining unit, configured to determine a Pearson correlation coefficient between each feature in the third feature vector and other features;

[0129] The second deletion unit is used to randomly delete one of the two features whose Pearson correlation coefficient is greater than the correlation coefficient threshold to obtain a fourth eigenvector;

[0130] a construction unit for constructing a co-selection relationship graph between features in the fourth feature vector to identify features associated with protein toxicity;

[0131] The selection unit is used to select the target number of features with the largest weight as representatives, pass them to the recursive feature elimination based on cross-validation, and obtain the second feature vector of the target dimension.

[0132] The present invention provides a method and system for predicting protein toxicity based on sequence information. The method introduces a deep learning model (such as a neural network model based on a self-attention mechanism) that can extract multi-level and deep features from the original sequence, especially capturing features that are difficult to design or cannot be discovered manually, thereby improving the prediction accuracy of the model, reducing dependence on manual feature design, and being adaptable to different types of protein sequences.

[0133] This method combines a pretrained language model with protein sequence embedding technology. By pretraining on large-scale, unlabeled protein sequence data, it learns rich sequence representations, providing a powerful foundation for subsequent toxicity prediction. This improves the model's predictive power for novel protein sequences. Even for previously unseen variants or rare proteins, the method can accurately predict these variants based on learned common sequence patterns. BRIEF DESCRIPTION OF THE DRAWINGS

[0134] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0135] Figure 1 A schematic diagram of a process for predicting protein toxicity based on sequence information provided in an embodiment of the present invention;

[0136] Figure 2 A schematic diagram of the structure of a deep learning model provided by an embodiment of the present invention;

[0137] Figure 3 A graph comparing the classification performance of the embodiment of the present invention with other classifiers on a test set;

[0138] Figure 4 A schematic diagram of the structure of a protein toxicity prediction system based on sequence information provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0139] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention. Example 1

[0140] like Figure 1As shown, this embodiment provides a method for predicting protein toxicity based on sequence information, comprising:

[0141] Step 101: Acquire multiple protein sequences from a protein database; wherein the multiple protein sequences include toxic and non-toxic protein sequences.

[0142] In this step, a search was performed on the Protein Data Bank website using the keywords "KW-0800 AND Reviewed" within the complete database. The keyword "KW-0800" represents naturally occurring toxic proteins that damage or kill other cells. These toxic proteins are produced by venomous animals (primarily snakes, scorpions, spiders, sea anemones, and cone snails), plants, fungi, and pathogenic bacteria. "Reviewed" represents experimentally verified data that has undergone manual review. The protein sequences obtained using the keywords "KW-0800 AND Reviewed" are considered the set of toxic positive samples (also called positive samples).

[0143] Then, the keywords "NOT KW-0800 NOT KW-0020 AND Reviewed" were used to retrieve protein sequences that were non-toxic, non-allergenic, and manually reviewed. The resulting protein sequences were used as a set of non-toxic negative samples (also known as negative samples).

[0144] The two sets of protein sequence samples are stored in two separate FASTA files. FASTA is a text file format widely used in bioinformatics, primarily for storing nucleic acid sequences (such as DNA and RNA) or protein sequences and their associated annotation information.

[0145] The acquired protein sequences require rigorous screening to ensure the reliability and accuracy of subsequent analyses. First, the length of all sequences in the positive and negative samples (i.e., the toxic and non-toxic sample sets) is checked, and sequences shorter than 35 amino acids are removed. Excessively short sequences are often fragmented or incomplete peptide sequences (short-chain proteins are often called peptides) and cannot provide sufficient information to support analytical tasks such as function prediction or sequence alignment. Therefore, protein sequences that are too short are excluded.

[0146] This example also filters for ambiguous amino acids in the sequence. Ambiguous amino acids are typically represented by the symbols B (indicating an unknown state for aspartic acid or asparagine), J (indicating an unknown state for leucine or isoleucine), O (indicating pyrrolysine, a rare amino acid often introduced into proteins as a result of post-translational modification), U (indicating selenocysteine, introduced via a specific signal and very rare), X (an unknown amino acid), and Z (indicating an ambiguous state for glutamate or glutamine). These ambiguous symbols may interfere with the accuracy of subsequent analysis, so all protein sequences containing these symbols were removed from the dataset. After this data cleaning step, the retained protein sequences both met the length requirements and did not contain any ambiguous amino acids.

[0147] Step 102: perform feature calculation on all protein sequences to obtain feature vectors of six major categories of proteins.

[0148] Illustratively, this step includes determining a feature vector of protein R.

[0149] Determine eigenvectors of a position-specific scoring matrix based on evolutionary signatures.

[0150] Determine the feature vectors for amino acid indices based on physicochemical properties.

[0151] Determine the feature vector of frequency counts for the amino acid group.

[0152] A feature vector that determines the length of a protein sequence.

[0153] Determine the feature vector of the common protein segment sequence.

[0154] Among them, the characteristic vector for determining protein R includes:

[0155] The expression for constructing the characteristic vector AAC of amino acid composition is:

[0156] .

[0157] Among them, p(a i ) is the i-th amino acid a in the sequence i The frequency of occurrence of f(a i ) is the i-th amino acid a in the sequence i The number of occurrences of i=1, 2, ..., 20; N is the total length of the sequence; in this embodiment, only the 20 common amino acids in nature are considered, so the length of the vector AAC is fixed to 20, and each value represents the relative frequency of the corresponding amino acid.

[0158] And / or, construct the expression of the dipeptide composition feature vector DPC:

[0159] .

[0160] Among them, p(d i,j ) is the dipeptide d in the sequence i,j The frequency of occurrence of f(d i,j ) is the dipeptide d in the sequence i,j The number of occurrences of; j = 1, 2, ..., 20; a sequence of length N has N-1 adjacent amino acid pairs, so N-1 represents the total number of dipeptides that may be formed in the sequence; the length of the vector is fixed to 400, and each value represents the relative frequency of the corresponding dipeptide.

[0161] And / or, construct an expression for the eigenvector MB of the Moreau-Broto autocorrelation:

[0162] .

[0163] Among them, R d is the Moreau-Broto autocorrelation value of the sequence with a delay of d; d represents the number of amino acids in the autocorrelation calculation, which is called the delay; d = 1, 2, ..., D; D represents the maximum delay. In this embodiment, D is a fixed value of 240; p i is the physicochemical property value of the i-th amino acid in the sequence; p i+d is the physicochemical property value of the i+dth amino acid in the sequence.

[0164] And / or, construct the expression for the eigenvector MORAN of the Moran autocorrelation:

[0165] .

[0166] Among them, M d is the Moran autocorrelation value of the sequence with delay d; is the mean of all physicochemical property values;

[0167] And / or, construct the expression for the eigenvector GEARY of the Geary autocorrelation:

[0168] .

[0169] Among them, G d is the Geary autocorrelation value of the sequence with delay d;

[0170] And / or, construct the expression of the characteristic vector CTD of the combined transformation distribution:

[0171] .

[0172] Among them, count(A i' ) is the i'th type amino acid A i' The number of occurrences of Indicates the i'th amino acid A in the sequence i' and the j'th amino acid A j' In this embodiment, considering the rare 21st amino acid, the values ​​of i' and j' range from 1 to 21; Representation captures the spatial distribution characteristics of each amino acid class; An operator representing a splicing vector; position encoding is performed at specific positions in the sequence (the first occurrence, the 25th percentile, the 50th percentile, the 75th percentile, and the last position), with five distribution positions calculated for each category. A 21st rare amino acid is also considered here, so the CTD feature vector dimension is 21 + 21 + 105 = 147.

[0173] And / or, construct the expression for the eigenvector CTRIAD of the joint triples:

[0174] .

[0175] Among them, the grouping mapping of the physicochemical properties of amino acids is G={g1,g2,…,g N};g i =1,2,…,7; g i Indicates the group number of the i-th amino acid; each component of CTRIAD represents a specific triplet The normalized frequency of ; Represents a triple in a sequence The number of occurrences of g i is the group number of the i-th amino acid in the sequence; g j is the group number of the jth amino acid in the sequence; is the group number of the k'th amino acid in the sequence.

[0176] And / or, construct the expression for the eigenvector SOCN of the sequence order coupling number:

[0177] .

[0178] in, represents the number of sequential couplings of delay d in the sequence on the physicochemical property k''; is the value of the physicochemical property k'' of the i-th amino acid; Represents The amino acid property value of delay d; Represents the global variance of attribute values, used for normalization; is the mean value of the physical and chemical properties k''; K is the total number of physical and chemical properties; in this embodiment, the number of physical and chemical properties considered K is 2; the maximum delay considered D is 30, so the length of the SOCN feature vector is K×D=60.

[0179] And / or, construct the expression of the eigenvector QSO of the quasi-sequential order descriptor:

[0180] .

[0181] in, represents the coupling characteristics of the delay d in the sequence on the physical and chemical property k''; in the embodiment, QSO is a complete feature vector with a total length of 100.

[0182] And / or, construct the expression of the pseudo amino acid composition feature vector PAAC:

[0183] .

[0184] in, represents the value of the physicochemical property k'' representing the autocorrelation part of the series with delay d.

[0185] In this example, the feature vector for pseudo-amino acid compositions (K=3, D=10) only considers limited-delay autocorrelation information about primary sequence characteristics and physicochemical properties. The feature vector for amphipathic pseudo-amino acid compositions (APAAC) (i.e., K=6, D=10 in PAAC) is expanded to include richer physicochemical properties and higher-level autocorrelation information.

[0186] Determine the eigenvectors of the position-specific scoring matrix based on evolutionary features, including:

[0187] The tool Psi-BLAST was used to calculate a position-specific scoring matrix (PSSM) for each protein sample. The search database was set to a vetted uniprot database file, with an evalue threshold of 0.001 and a number of iterations of 3. First, Psi-BLAST begins with a query sequence and performs a standard BLAST search to find sequences that are significantly similar to the query sequence. Next, the homologous sequences found (typically filtered by an evalue threshold) are aligned with the query sequence to generate a multiple sequence alignment. Next, a position-specific frequency matrix was calculated.

[0188] The corrected probability of amino acid a at position i'' in the sequence is calculated according to the following formula :

[0189] .

[0190] in, represents the frequency of occurrence of amino acid a at position i'' in the sequence; represents the frequency of occurrence of amino acid b at position i'' in the sequence; A is a set of natural amino acids (20 natural amino acids).

[0191] The score of position i'' and amino acid a in the sequence is calculated according to the following formula , which are the elements of PSSM:

[0192] .

[0193] Among them, λ is the scaling factor used to adjust the scale of the score; q(a) is the average frequency of amino acid a in nature, that is, the background probability, which is a value based on the statistics of the reference database.

[0194] Psi-BLAST uses the updated PSSM to re-search the database to find new homologous sequences, repeating the above process until the results converge. Finally, the global average pooling algorithm is used.

[0195] The average score p[a] of amino acid a was calculated according to the following formula:

[0196] .

[0197] Where N is the total length of the sequence. In this example, P[a] is the average score for amino acid a in the pooled global representation. This operation is independent of input length and is well suited for unified processing of variable-length inputs. The feature vector length is 20 dimensions.

[0198] Amino acid indexing based on physicochemical properties. The physicochemical properties of 617 amino acids, including volume, folding propensity, isoelectric point, flexibility, and folding free energy, were selected from the AAindex database. Physicochemical property statistics were generated for each protein sequence, with a feature vector dimension of 617.

[0199] Amino acid group frequency counts. Amino acids are categorized into six groups: hydrophobic, positively charged, negatively charged, polar, conformational, and other. The amino acids in each sequence are grouped and the number of amino acids in each group is counted. The resulting feature vector is 6-dimensional.

[0200] Protein sequence length statistics: Only the number of amino acids needs to be counted, and the feature vector length is 1 dimension.

[0201] Common fragment sequence search. The negative and positive samples were fed into the MERCI tool separately. Possible common fragment sequences were mined from the positive samples, while the specificity of the common fragment sequences was verified using the negative samples. The core goal of MERCI was to find common fragment sequences that were significantly enriched in the positive samples but almost absent in the negative samples. The number of common fragment sequences was set to 50, and one-hot encoding was used to generate a 50-dimensional feature vector from the common fragment sequences.

[0202] Step 103: Concatenate the six categories of eigenvectors by column to obtain a first eigenvector.

[0203] All feature vectors obtained in step 102 are concatenated together according to dimension 1 to obtain the first feature vector Features:

[0204] .

[0205] Step 104: perform dimensionality reduction screening on the first eigenvector to obtain a second eigenvector of the target dimension.

[0206] Exemplarily, this step includes:

[0207] Delete the columns with a standard deviation of 0 in the first eigenvector to obtain the third eigenvector.

[0208] Determine the Pearson correlation coefficient of each feature with every other feature in the third eigenvector.

[0209] If the correlation coefficient of two features is greater than or equal to 0.8, then the two features are strongly correlated. Therefore, one of the two features whose Pearson correlation coefficient is greater than the correlation coefficient threshold (0.8 in this embodiment) is randomly deleted to obtain the fourth eigenvector.

[0210] A co-selection relationship graph between features in the fourth eigenvector was constructed to identify features associated with protein toxicity.

[0211] The features with the largest target number of weights are selected as representatives and passed to recursive feature elimination based on cross-validation to obtain the second feature vector of the target dimension.

[0212] By running multiple machine learning algorithms (such as decision trees, support vector machines, and gradient boosting machines), we construct a co-selection relationship graph between features and identify key features highly correlated with protein toxicity. We then select the 100 most heavily weighted features as representatives and pass them through recursive feature elimination based on cross-validation. The goal is to find an optimal feature subset that maximizes model performance given a feature matrix and a target. The evaluation function is:

[0213] .

[0214] Among them, K f is the number of cross-validation folds, which is selected as 20 folds; is the feature vector of the hth feature of the training; is the feature vector of the hth feature to be verified; y (train) is the training label; y (val) The inference result label for verification. Recursively eliminate features until the optimal one is found.

[0215] Step 105: Use the second eigenvector to train a graph attention-based neural network model, and use the trained neural network model as a protein toxicity prediction model.

[0216] A dual-channel graph attention network is used to extract protein feature matrix. The model architecture is as follows Figure 2 As shown in Figure 2, the feature data first undergoes preprocessing and a one-dimensional convolution operation to generate a preliminary feature representation. The one-dimensional convolution layer effectively captures local relationships between adjacent amino acids in the sequence, helping to extract local patterns in the protein (such as toxic fragments). The output of this convolution layer is used in subsequent feature extraction modules. The model employs two parallel dual-channel graph attention network modules. One is a feature-oriented attention module, which receives the feature representation and establishes associations between feature nodes using a graph attention mechanism to output feature-driven embeddings. This module focuses on interactions within protein features, identifying important toxicity-related fragments in the sequence, such as specific amino acid combinations or structures. The other is a sequence-oriented graph attention module, which constructs a sequence-oriented graph structure to capture long-range dependencies between different positions in the sequence and output sequence-driven embeddings. This mechanism helps the model understand complex interactions within protein sequences, particularly relationships between long-range amino acids, and aids in identifying complex toxic fragments. After feature fusion, the concatenated feature vector is further processed by a gated recurrent unit (GRU) module to capture the complex dependencies within the protein sequence. GRUs can extract deeper dependencies and patterns within the sequence, helping the model better identify toxicity-related feature combinations. The output vector of the GRU serves as input to the subsequent classification module. The classifier module generates an inference score between 0 and 1, which is used for the final toxicity prediction decision. If the inference score is greater than or equal to a certain threshold (set to 0.5), the protein is considered toxic; otherwise, it is considered non-toxic. The model training step saves the trained model as a protein toxicity prediction model for subsequent inference prediction of new samples.

[0217] Step 106: input the new protein sequence into the protein toxicity prediction model to obtain a toxicity prediction result output by the protein toxicity prediction model.

[0218] Input a new protein sequence file in FASTA format to the protein toxicity prediction model and directly perform inference prediction. The model will generate a corresponding comma-delimited file, annotating the toxicity of the input protein.

[0219] In order to verify the prediction performance of the toxicity prediction method provided in this example, we tested it on the uniprotKB dataset and compared it with other prediction tools, including ToxinPred2, ToxIBTL, Toxify, CSM-Toxin, VISH-Pred, and ToxinPred3. The performance comparison results are shown in Figure 2. Figure 3 shown.

[0220] Figure 3 In the figure, the best is shown in bold, and the second-best is shown in underline. Experimental results show that our method outperforms other existing classification predictors in multiple metrics, including accuracy, recall, F1 value, area under the ROC curve (AUC), and Matthews correlation coefficient. These performance advantages fully demonstrate that our method has significant classification effectiveness and reliability in protein toxicity prediction tasks.

[0221] The accuracy of the toxicity prediction method provided in this embodiment reached 0.948, which is significantly higher than other predictors, especially about 5.5% higher than the current optimal VISH-Pred, showing its significant advantage in overall classification accuracy. In terms of recall rate, the recall rate of the toxicity prediction method provided in this embodiment reached 0.953, which is significantly improved compared with other methods, indicating that the toxicity prediction method provided in this embodiment can more effectively capture positive samples in protein toxicity and improve the sensitivity of the model. In terms of the comprehensive indicator of F1 value, the F1 value of the toxicity prediction method provided in this embodiment is 0.932, which is much higher than other models, indicating that the model can accurately predict toxic proteins while also reducing the occurrence of false positives and maintaining good balanced performance. In terms of AUC, the AUC value of the toxicity prediction method provided in this embodiment reached 0.949, which is nearly 7.8% higher than VISH-Pred's 0.871, showing a stronger ability to distinguish between positive and negative samples under different thresholds, thereby enhancing the robustness of the model. In terms of MCC, the MCC value of the toxicity prediction method provided in this embodiment reached 0.891, further proving that the model still has excellent prediction performance when the sample categories are unbalanced.

[0222] In summary, this embodiment provides a method for predicting protein toxicity based on sequence information. This method automatically extracts features from protein sequences by introducing deep learning models (such as convolutional neural networks, recurrent neural networks, and Transformer models based on self-attention mechanisms). Unlike traditional methods, this embodiment leverages the powerful feature learning capabilities of deep learning models to extract multi-layered, deep features from raw sequences, particularly capturing features that are difficult or impossible to discover manually. This automated feature extraction significantly improves the model's prediction accuracy while reducing reliance on manual feature design, making this embodiment more versatile and adaptable to different types of protein sequences.

[0223] Modeling global and local dependencies in sequences: Existing deep learning models, such as convolutional neural networks (CNNs), although they have certain advantages in capturing local features of protein sequences, have difficulty processing long-range dependencies in sequences. The toxicity of proteins is often not only related to local sequence fragments, but is also affected by the complex interactions between long-range amino acid residues. To overcome this limitation, this embodiment adopts more advanced deep learning architectures, such as Transformer or improved RNN models, which can process the entire sequence through the self-attention mechanism, thereby effectively modeling long-range dependencies in protein sequences. At the same time, this embodiment introduces a multi-scale feature fusion mechanism in the model, which can simultaneously capture local and global feature information, ensuring that toxicity prediction not only takes into account local patterns in the sequence, but also integrates the interactions between long-range amino acids, thereby improving the comprehensiveness and accuracy of the prediction.

[0224] Enhance the generalization ability of unknown protein sequences and variants: Existing technologies often lack sufficient generalization ability for the prediction of new protein sequences or variants. The models rely too much on known databases during training, and the amount of data in these databases is limited and the coverage is not comprehensive. When faced with sequences that have not appeared in the training set, the prediction performance of traditional models decreases significantly. To solve this problem, this embodiment combines pre-trained language models with protein sequence embedding technology. These pre-trained models learn rich sequence representations by pre-training on large-scale unlabeled protein sequence data, thereby providing a powerful basic model for subsequent toxicity prediction. This method greatly improves the model's predictive ability on new protein sequences. Even for variants or rare proteins that have never been seen before, the method of this embodiment can make accurate predictions through the learned general sequence patterns. In addition, this embodiment also uses data enhancement techniques, such as random sequence perturbations and mutation simulations, to further enhance the adaptability of the model when processing variant sequences. Example 2

[0225] Based on the same inventive concept as Example 1, this example provides a protein toxicity prediction system based on sequence information. Since the principle of solving the problem by this system is similar to the aforementioned protein toxicity prediction method based on sequence information, the implementation of this system can refer to the implementation of the protein toxicity prediction method based on sequence information.

[0226] like Figure 4 As shown, the protein toxicity prediction system based on sequence information includes:

[0227] The sequence acquisition module 10 is used to acquire multiple protein sequences from a protein database; wherein the multiple protein sequences include toxic and non-toxic protein sequences.

[0228] The feature calculation module 20 is used to perform feature calculation on all protein sequences to obtain feature vectors of six major categories of proteins.

[0229] The vector concatenation module 30 is used to concatenate the six types of feature vectors by column to obtain a first feature vector.

[0230] The dimensionality reduction and screening module 40 is used to perform dimensionality reduction and screening on the first eigenvector to obtain a second eigenvector of a target dimension.

[0231] The model training module 50 is used to train a graph attention-based neural network model using the second eigenvector, and use the trained neural network model as a protein toxicity prediction model.

[0232] The toxicity prediction module 60 is used to input a new protein sequence into the protein toxicity prediction model to obtain a toxicity prediction result output by the protein toxicity prediction model.

[0233] Exemplarily, the feature calculation module includes:

[0234] The first determining unit is used to determine the characteristic vector of protein R.

[0235] The second determining unit is configured to determine a feature vector of a position-specific scoring matrix based on evolutionary features.

[0236] The third determining unit is used to determine a feature vector of the amino acid index based on physicochemical properties.

[0237] The fourth determining unit is used to determine a characteristic vector of the amino acid group frequency count.

[0238] The fifth determining unit is used to determine the feature vector of the protein sequence length.

[0239] The sixth determining unit is used to determine the feature vector of the protein common fragment sequence.

[0240] Exemplarily, the first determining unit includes:

[0241] The first construction means is used to construct the expression of the amino acid composition feature vector AAC:

[0242] .

[0243] Among them, p(a i ) is the i-th amino acid a in the sequence i The frequency of occurrence of f(a i ) is the i-th amino acid a in the sequence i The number of occurrences of; i=1,2,...,20; N is the total length of the sequence.

[0244] And / or, a second constructing means for constructing an expression of a dipeptide composition characteristic vector DPC:

[0245] .

[0246] Among them, p(d i,j ) is the dipeptide d in the sequence i,j The frequency of occurrence of f(d i,j ) is the dipeptide d in the sequence i,j The number of occurrences of ; j=1,2,...,20.

[0247] And / or, a third constructing means for constructing an expression for the eigenvector MB of the Moreau-Broto autocorrelation:

[0248] .

[0249] Among them, R d is the Moreau-Broto autocorrelation value of the sequence with delay d; d = 1, 2, ..., D; D represents the maximum delay; p i is the physicochemical property value of the i-th amino acid in the sequence; p i+d is the physicochemical property value of the i+dth amino acid in the sequence.

[0250] And / or, the fourth constructing means is used to construct an expression of the eigenvector MORAN of the Moran autocorrelation:

[0251] .

[0252] Among them, M d is the Moran autocorrelation value of the sequence with delay d; is the mean of all physicochemical property values.

[0253] And / or, a fifth constructing means for constructing an expression for the eigenvector GEARY of the Geary autocorrelation:

[0254] .

[0255] Among them, G d is the Geary autocorrelation value of the sequence with delay d.

[0256] And / or, a sixth constructing means for constructing an expression of a characteristic vector CTD of a combined transformation distribution:

[0257] .

[0258] Among them, count(A i' ) is the i'th type amino acid A i' The number of occurrences of Indicates the i'th amino acid A in the sequence i' and the j'th amino acid A j' The number of adjacencies; Representation captures the spatial distribution characteristics of each amino acid class; An operator representing concatenation vectors.

[0259] And / or, a seventh constructing means for constructing an expression of a characteristic vector CTRIAD of a joint triple:

[0260] .

[0261] in, Represents a triple The normalized frequency of occurrence of Represents a triple in a sequence The number of occurrences of g i is the group number of the i-th amino acid in the sequence; g j is the group number of the jth amino acid in the sequence; is the group number of the k'th amino acid in the sequence.

[0262] And / or, an eighth constructing means for constructing an expression of a characteristic vector SOCN of a sequence order coupling number:

[0263] .

[0264] in, represents the number of sequential couplings of delay d in the sequence on the physicochemical property k''; is the value of the physicochemical property k'' of the i-th amino acid; Represents The amino acid property value of delay d; Represents the global variance of attribute values, used for normalization; is the mean value of the physicochemical properties k''; K is the total number of physicochemical properties.

[0265] And / or, the ninth constructing means is used to construct the expression of the characteristic vector QSO of the quasi-sequential order descriptor:

[0266] .

[0267] in, represents the coupling characteristics of the delay d in the sequence on the physicochemical property k'';

[0268] And / or, a tenth constructing means for constructing an expression of a pseudo amino acid composition feature vector PAAC:

[0269] .

[0270] in, represents the value of the physicochemical property k'' representing the autocorrelation part of the series with delay d.

[0271] Exemplarily, the second determining unit includes:

[0272] The first calculation means is used to calculate the corrected probability of amino acid a at position i'' in the sequence according to the following formula :

[0273] .

[0274] in, represents the frequency of occurrence of amino acid a at position i'' in the sequence; represents the frequency of occurrence of amino acid b at position i'' in the sequence; A is a set of natural amino acids.

[0275] The second calculation device is used to calculate the score of position i'' and amino acid a in the sequence according to the following formula :

[0276] .

[0277] Where λ is a scaling factor used to adjust the scale of the score; q(a) is the average frequency of amino acid a in nature.

[0278] The third calculating means is used to calculate the average score p[a] of amino acid a according to the following formula:

[0279] .

[0280] Where N is the total length of the sequence.

[0281] Exemplarily, the dimensionality reduction screening module includes:

[0282] The first deletion unit is used to delete the columns with a standard deviation of 0 in the first eigenvector to obtain a third eigenvector.

[0283] The determining unit is used to determine the Pearson correlation coefficient between each feature in the third feature vector and other features.

[0284] The second deletion unit is used to randomly delete one of the two features whose Pearson correlation coefficient is greater than the correlation coefficient threshold to obtain a fourth eigenvector;

[0285] A construction unit is used to construct a co-selection relationship graph between features in the fourth feature vector to identify features associated with protein toxicity.

[0286] The selection unit is used to select the target number of features with the largest weight as representatives, pass them to the recursive feature elimination based on cross-validation, and obtain the second feature vector of the target dimension.

[0287] For more specific working processes of the above modules, please refer to the corresponding content disclosed in Example 1, which will not be repeated here. Example 3

[0288] This embodiment provides a computer device, including a processor and a memory; wherein, when the processor executes a computer program stored in the memory, the steps of the protein toxicity prediction method based on sequence information described in Example 1 are implemented.

[0289] For more specific details about the above method, please refer to the corresponding content disclosed in Example 1, which will not be repeated here. Example 4

[0290] This embodiment provides a computer-readable storage medium for storing a computer program; when the computer program is executed by a processor, the steps of the protein toxicity prediction method based on sequence information described in Example 1 are implemented.

[0291] For more specific details about the above method, please refer to the corresponding content disclosed in Example 1, which will not be repeated here. Example 5

[0292] This embodiment provides a computer program product, including computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the steps of the protein toxicity prediction method based on sequence information described in Example 1 are implemented.

[0293] For more specific details about the above method, please refer to the corresponding content disclosed in Example 1, which will not be repeated here.

[0294] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. References to the same or similar parts between the various embodiments will be sufficient. The systems, devices, storage media, and computer program products disclosed in the embodiments correspond to the methods disclosed in the embodiments, so their descriptions are relatively simplified. For relevant details, refer to the method descriptions.

[0295] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software and a necessary general-purpose hardware platform. Based on this understanding, the technical solutions in the embodiments of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, or an optical disk, and includes instructions for enabling a computer device (such as a personal computer, server, or network device) to execute the methods described in various embodiments of the present invention, or portions thereof.

[0296] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0297] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0298] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0299] The present invention has been described in detail above with reference to specific embodiments and exemplary examples. However, these descriptions should not be construed as limiting the present invention. Those skilled in the art will appreciate that various equivalent substitutions, modifications, or improvements may be made to the technical solutions and implementations of the present invention without departing from the spirit and scope of the present invention, all of which fall within the scope of the present invention. The scope of protection of the present invention shall be determined by the appended claims.

Claims

1. A method for predicting protein toxicity based on sequence information, characterized in that: include: Acquire multiple protein sequences from a protein database; wherein the multiple protein sequences include toxic and non-toxic protein sequences; The features of all protein sequences were calculated to obtain the feature vectors of six major categories of proteins; Concatenate the six categories of eigenvectors by column to obtain the first eigenvector; Perform dimensionality reduction screening on the first eigenvector to obtain a second eigenvector of the target dimension; this includes: deleting some columns in the first eigenvector to obtain a third eigenvector; determining the Pearson correlation coefficient between each feature in the third eigenvector and other features to obtain a fourth eigenvector; constructing a co-selection relationship diagram between the features in the fourth eigenvector, selecting the target number of features with the largest weights as representatives, and passing them to recursive feature elimination based on cross-validation to obtain the second eigenvector of the target dimension; The second eigenvector is used to train a graph attention-based neural network model, and the trained neural network model is used as a protein toxicity prediction model; The graph attention-based neural network model includes the following steps: feature data is first preprocessed and subjected to a one-dimensional convolution operation to generate a preliminary feature representation; two parallel dual-channel graph attention network modules are used: a feature-oriented attention module that outputs a feature-oriented embedding; and a sequence-oriented graph attention module that outputs a sequence-oriented embedding. After feature fusion, the concatenated feature vector is processed by a gated recurrent unit. The output vector of the gated recurrent unit serves as the input to the subsequent classification module, which generates an inference score between 0 and 1. Inputting the new protein sequence into the protein toxicity prediction model to obtain a toxicity prediction result output by the protein toxicity prediction model; The feature calculation is performed on all protein sequences to obtain six categories of protein feature vectors, including: determining the feature vector of protein R, determining the feature vector of the position-specific scoring matrix based on evolutionary characteristics, determining the feature vector of the amino acid index based on physicochemical properties, determining the feature vector of the amino acid group frequency count, determining the feature vector of the protein sequence length, and determining the feature vector of the protein common fragment sequence; The step of determining the characteristic vector of protein R comprises: Construct the feature vector AAC of amino acid composition and the feature vector DPC of dipeptide composition; Construct the expression of the eigenvector MB of the Moreau-Broto autocorrelation: Among them, R d is the Moreau-Broto autocorrelation value of the sequence with delay d; d = 1, 2, ..., D; D is the maximum delay; p i is the physicochemical property value of the i-th amino acid in the sequence; p i+d is the physicochemical property value of the i+dth amino acid in the sequence; the expression of the eigenvector GEARY of the Geary autocorrelation is constructed: Among them, G d is the Geary autocorrelation value of the sequence with delay d; The expression for constructing the characteristic vector CTD of the combined transformation distribution is: Among them, count(A i′ ) is the i'th type amino acid A i' The number of occurrences of count(A i′ →A j′ ) represents the i'th amino acid A in the sequence i' and the j'th amino acid A j' The number of adjacencies; D′ k Representation captures the spatial distribution characteristics of each amino acid class; An operator representing the concatenation vector; The expression for constructing the eigenvector CTRIAD of the joint triple is: Among them, p(T ijk′ ) represents the triple T ijk′ The normalized frequency of occurrence of count(T ijk′ ) represents the triple T in the sequence ijk′ The number of occurrences of g i is the group number of the i-th amino acid in the sequence; g j is the group number of the jth amino acid in the sequence; g k′ is the group number of the k'th amino acid in the sequence; Construct the expression of the characteristic vector SOCN of the sequence order coupling number: Among them, Θ k″ (d) represents the sequential coupling number of the delay d in the sequence on the physicochemical property k'; p k″ (a i ) is the value of the physicochemical property k" of the i-th amino acid; p k″ (a i+d ) indicates that k″ (a i ) delay d amino acid property value; Represents the global variance of attribute values, used for normalization; is the mean value of the physicochemical property k”; K is the total number of physicochemical properties; Construct the eigenvector QSO of the quasi-sequential order descriptor, Pseudo amino acid composition feature vector PAAC.

2. The protein toxicity prediction method according to claim 1, characterized in that: The step of determining the eigenvector of the position-specific scoring matrix based on the evolutionary features comprises: The corrected probability p(i″,a) of amino acid a at position i″ in the sequence is calculated according to the following formula: Where F(i″,a) represents the frequency of occurrence of amino acid a at position i″ of the sequence; F(i″,b) represents the frequency of occurrence of amino acid b at position i″ of the sequence; A is the set of natural amino acids; The score M[i″][a] for position i″ in the sequence and amino acid a is calculated according to the following formula: Where λ is a scaling factor used to adjust the scale of the score; q(a) is the average frequency of amino acid a in nature; The average score p[a] of amino acid a was calculated according to the following formula: Where N is the total length of the sequence.

3. The protein toxicity prediction method according to claim 1, characterized in that: The step of deleting some columns in the first eigenvector to obtain the third eigenvector is specifically performed by deleting columns with a standard deviation of 0 in the first eigenvector to obtain the third eigenvector; the step of determining the Pearson correlation coefficient between each feature in the third eigenvector and other features to obtain the fourth eigenvector is specifically performed by determining the Pearson correlation coefficient between each feature in the third eigenvector and other features, and randomly deleting one of two features whose Pearson correlation coefficient is greater than a correlation coefficient threshold to obtain the fourth eigenvector; The co-selection relationship graph between the features in the fourth feature vector is constructed to identify features associated with protein toxicity.

4. A protein toxicity prediction system based on sequence information, characterized in that: include: A sequence acquisition module is used to acquire multiple protein sequences from a protein database; wherein the multiple protein sequences include toxic and non-toxic protein sequences; Feature calculation module, used to calculate the features of all protein sequences and obtain the feature vectors of six major categories of proteins; The vector concatenation module is used to concatenate the six categories of eigenvectors by column to obtain the first eigenvector; A dimensionality reduction and screening module is used to perform dimensionality reduction screening on the first eigenvector to obtain a second eigenvector of the target dimension. The module includes: deleting some columns in the first eigenvector to obtain a third eigenvector; determining the Pearson correlation coefficient between each feature in the third eigenvector and other features to obtain a fourth eigenvector; constructing a co-selection relationship diagram between the features in the fourth eigenvector, selecting the target number of features with the largest weights as representatives, and passing them to recursive feature elimination based on cross-validation to obtain the second eigenvector of the target dimension. A model training module is used to train a graph attention-based neural network model using the second eigenvector and use the trained neural network model as a protein toxicity prediction model. The graph attention-based neural network model includes: feature data is first preprocessed and subjected to a one-dimensional convolution operation to generate a preliminary feature representation; two parallel dual-channel graph attention network modules are used: a feature-oriented attention module outputs a feature-oriented embedding; and a sequence-oriented graph attention module outputs a sequence-oriented embedding. After feature fusion, the concatenated feature vector is processed by a gated recurrent unit. The output vector of the gated recurrent unit serves as the input of a subsequent classification module, which generates an inference score between 0 and 1. A toxicity prediction module is used to input a new protein sequence into a protein toxicity prediction model to obtain a toxicity prediction result output by the protein toxicity prediction model; The feature calculation module includes: a first determination unit for determining a feature vector of protein R; a second determination unit for determining a feature vector of a position-specific scoring matrix based on evolutionary characteristics; a third determination unit for determining a feature vector of an amino acid index based on physicochemical properties; a fourth determination unit for determining a feature vector of an amino acid group frequency count; a fifth determination unit for determining a feature vector of a protein sequence length; and a sixth determination unit for determining a feature vector of a protein common segment sequence. The first determining unit includes: The first construction device is used to construct an amino acid composition feature vector AAC; The second construction device is used to construct a characteristic vector DPC of the dipeptide composition; The third constructing means is used to construct the expression of the characteristic vector MB of the Moreau-Broto autocorrelation: Among them, R d is the Moreau-Broto autocorrelation value of the sequence with delay d; d = 1, 2, ..., D; D is the maximum delay; p i is the physicochemical property value of the i-th amino acid in the sequence; p i+d is the physicochemical property value of the i+dth amino acid in the sequence; The fifth constructing means is used to construct the expression of the characteristic vector GEARY of the Geary autocorrelation: Among them, G d is the Geary autocorrelation value of the sequence with delay d; The sixth constructing means is used to construct an expression of the characteristic vector CTD of the combined transformation distribution: Among them, count(A i' ) is the i'th type amino acid A i' The number of occurrences of count(A i′ →A j′ ) represents the i'th amino acid A in the sequence i' and the j'th amino acid A j' The number of adjacencies; D′ k Representation captures the spatial distribution characteristics of each amino acid class; An operator representing the concatenation vector; The seventh construction means is used to construct the expression of the characteristic vector CTRIAD of the joint triple: Among them, p(T ijk′ ) represents the triple T ijk′ The normalized frequency of occurrence of count(T ijk′ ) represents the triple T in the sequence ijk′ The number of occurrences of g i is the group number of the i-th amino acid in the sequence; g j is the group number of the jth amino acid in the sequence; g k′ is the group number of the k'th amino acid in the sequence; The eighth constructing means is used to construct an expression for the characteristic vector SOCN of the sequence order coupling number: Among them, Θ k″ (d) represents the sequential coupling number of the delay d in the sequence on the physicochemical property k'; p k″ (a i ) is the value of the physicochemical property k" of the i-th amino acid; p k″ (a i+d ) indicates that k″ (a i ) delay d amino acid property value; Represents the global variance of attribute values, used for normalization; is the mean value of the physicochemical property k”; K is the total number of physicochemical properties; A ninth constructing means for constructing a characteristic vector QSO of a quasi-sequence order descriptor; The tenth construction device is used to construct the characteristic vector PAAC composed of pseudo amino acids.

5. The protein toxicity prediction system according to claim 4, characterized in that: The second determining unit includes: The first calculation means is used to calculate the corrected probability p(i″,a) of amino acid a at position i″ in the sequence according to the following formula: Where F(i″,a) represents the frequency of occurrence of amino acid a at position i″ of the sequence; F(i″,b) represents the frequency of occurrence of amino acid b at position i″ of the sequence; A is the set of natural amino acids; The second calculation means is used to calculate the score M[i″][a] between position i″ and amino acid a in the sequence according to the following formula: Where λ is a scaling factor used to adjust the scale of the score; q(a) is the average frequency of amino acid a in nature; The third calculating means is used to calculate the average score p[a] of amino acid a according to the following formula: Where N is the total length of the sequence.

6. The protein toxicity prediction system according to claim 4, characterized in that: The dimension reduction screening module includes: A first deletion unit is used to delete the columns with a standard deviation of 0 in the first eigenvector to obtain a third eigenvector; a determining unit, configured to determine a Pearson correlation coefficient between each feature in the third feature vector and other features; The second deletion unit is used to randomly delete one of the two features whose Pearson correlation coefficient is greater than the correlation coefficient threshold to obtain a fourth eigenvector; a construction unit for constructing a co-selection relationship graph between features in the fourth feature vector to identify features associated with protein toxicity; The selection unit is used to select the target number of features with the largest weight as representatives, pass them to the recursive feature elimination based on cross-validation, and obtain the second feature vector of the target dimension.