Biological sequence function prediction method and system

By applying the TF-IDF method to the characterization of biological sequences and combining artificial intelligence models, the problems of insufficient predictive capabilities of biological sequence functions and sparse data in the prior art are solved, and efficient and accurate prediction of biological sequence functions are achieved, with wide application prospects.

CN120048337APending Publication Date: 2025-05-27WEST CHINA HOSPITAL SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510119257.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When the existing biological sequence function prediction methods deal with new genes or proteins that have not been included in the database, their prediction capabilities are limited, which can easily lead to a high false negative rate. They can easily lead to data sparse problems when processing high-dimensional biological sequences, which has a large calculation overhead and affects the accuracy of prediction.

Method used

The biological sequence is characterized by the TF-IDF method. By combining the biological sequence to be predicted with a fixed vocabulary, the TF-IDF vector representation is obtained after processing using the TF-IDF method, and inputting it into the artificial intelligence model for prediction.

Benefits of technology

It realizes efficient and accurate prediction of biological sequences in massive sequence data, which is significantly better than the prediction performance of the prior art, and realizes alignment-free functional prediction, which is suitable for various sequences and functional categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048337A_ABST
    Figure CN120048337A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of biological research, and particularly relates to a biological sequence function prediction method and system. The system comprises: an input module for inputting a sequence of at least one biological sequence to be predicted; the data preprocessing module is configured to combine a biological sequence to be predicted with a fixed vocabulary comprising inverse document frequency (IDF) values of lexical items, and process the biological sequence to be predicted through a TF-IDF method to obtain TF-IDF vector representation of the biological sequence to be predicted; and the prediction module is used for inputting the TF-IDF vector representation into an artificial intelligence model to obtain a prediction result about whether the biological sequence to be predicted has a specific function or not. The invention also provides a method for predicting the function of the biological sequence by using the system. According to the method, the TF-IDF method is applied to characterization of the biological sequence for the first time, the characterization result and the machine learning technology are combined to be used for predicting the relation between the specific sequence and the biological function, and the method has good application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of biological research, and particularly relates to a method and system for predicting the function of biological sequences. Background Art

[0002] Biological sequences refer to the linear arrangements of biological molecules such as DNA (deoxyribonucleic acid), RNA (ribonucleic acid), and proteins, which carry the genetic information and functional instructions of organisms and are the core basis of life activities. In modern biological research, exploring the relationship between biological sequences and their functions is an important research direction.

[0003] With the rapid development of high-throughput sequencing technology, researchers can now obtain various biological sequences at an affordable cost and in a relatively short time. This technological breakthrough has made it possible to predict the function based on biological sequences, directly inferring their functional attributes by analyzing sequence features and significantly improving the efficiency of related research. However, how to efficiently and accurately predict functions in the vast amount of sequence data remains a scientific and technical problem that urgently needs to be solved.

[0004] Alignment methods based on sequence similarity (such as BLAST) are traditional tools for predicting the function of biological sequences. These methods determine the function by aligning the sequence to be tested with the sequences in the reference database and presetting a relatively high similarity threshold. Although simple and effective, they have limited ability to predict new genes or proteins not included in the database, which is likely to result in a relatively high false negative rate.

[0005] In recent years, significant progress has been made in the application of machine learning technology in the field of biology, which can learn potential features from a large amount of data. Taking the prediction of pathogen drug resistance as an example, the DeepARG model developed by Arango Argoty et al. (2018) makes predictions based on the bit scores of the alignment, and this method still inherits the limitations of the alignment tools. Li (2024) et al. published a drug resistance prediction method (patent application number CN118571309A), which also requires combining the bit scores of the alignment and fails to completely get rid of the deficiencies of the alignment method. In contrast, ARGNet developed by Pei (2024) et al. represents the original sequence through one-hot encoding and directly feeds it into the model for prediction, achieving certain results. However, one-hot encoding is prone to data sparsity problems when dealing with high-dimensional biological sequences, increasing the computational overhead and affecting the prediction accuracy of downstream tasks.

[0006] Therefore, developing an artificial intelligence prediction method that can directly predict the function based on biological sequences remains an important research topic in this field. In artificial intelligence prediction tasks, how to extract and represent the features of biological sequences is one of the key factors determining the prediction performance of the model.

[0007] TF-IDF (Term Frequency-Inverse Document Frequency) is a commonly used weighting technique for information retrieval and text mining. It reflects the importance of a word for a document set or a single document in a corpus. The TF-IDF value increases proportionally with the frequency of the word in the document, but decreases inversely with the frequency of the word in the corpus. This means that TF-IDF tends to filter out common words and retain important words.

[0008] Currently, the TF-IDF method has been applied in some research in the biological field, such as processing gene expression matrices in single-cell omics and measuring the specific expression level of genes in each cell. However, there are no literature reports on using TF-IDF for biological sequence function prediction. It is also unclear whether the data obtained after characterizing biological sequences using TF-IDF can be combined with artificial intelligence technology to achieve effective biological sequence function prediction, and further relevant research is urgently needed. Summary of the Invention

[0009] In view of the problems of the prior art, the present invention provides a method and system for predicting biological sequence functions.

[0010] A biological sequence function prediction system includes:

[0011] An input module, configured to input at least one biological sequence to be predicted;

[0012] A data preprocessing module, configured to process the biological sequence to be predicted by the TF-IDF method in combination with a fixed vocabulary including the inverse document frequency (IDF) values of terms, to obtain a TF-IDF vector representation of the biological sequence to be predicted;

[0013] A prediction module, configured to input the TF-IDF vector representation into an artificial intelligence model to obtain a prediction result on whether the biological sequence to be predicted has a specific biological function.

[0014] Preferably, the biological sequence is selected from DNA, RNA, DNA-RNA hybrid sequences, proteins, and polypeptides.

[0015] Preferably, the TF-IDF method is implemented by a TF-IDF vectorizer;

[0016] And / or, the fixed vocabulary is constructed by the following steps: extracting terms from the training data, screening terms that meet the frequency conditions, and calculating the document frequency and inverse document frequency of the terms.

[0017] Preferably, the TF-IDF method includes the following steps:

[0018] Step 1: Split the sequence into consecutive subsequences, and denote the subsequences as k-mers;

[0019] Step 2: Calculate the term frequency TF of each k-mer in its corresponding sequence;

[0020] Step 3: Calculate the inverse document frequency IDF of each k-mer in the training set;

[0021] Step 4: Calculate the TF-IDF weight of each k-mer;

[0022] Step 5: Construct the TF-IDF vector representation through the TF-IDF weights of each k-mer.

[0023] Preferably, in Step 1, the k-mer is represented as:

[0024] kmer = {S[i:i + k - 1] | 1 ≤ i ≤ n - k + 1}

[0025] where S is a biological sequence of length n, S[i:i + k - 1] represents the k-mer subsequence starting from the i-th position in sequence S, and k is the length of the k-mer;

[0026] After splitting, the sequence S is converted into a feature sequence S K , that is

[0027] S K = (kmer 1 , …, kmer i , …, kmer K )

[0028] where kmer i represents the i-th kmer;

[0029] The length of the subsequence is selected from 3 - 21, preferably an odd number;

[0030] In Step 2, the calculation formula for the term frequency TF is:

[0031]

[0032] where count(kmer i , S) is the number of times kmer i appears in sequence S K ;

[0033] In Step 3, the calculation formula for the inverse document frequency IDF is:

[0034]

[0035] Among them, N is the number of sequences in the training set, and DF(kmer i ) is the number of sequences containing kmer i in the training set;

[0036] In step 4, the calculation formula of the TF-IDF weight is:

[0037] TF-IDF(kmer i , S K ) = TF(kmer i , S) · IDF(kmer i )

[0038] Among them, TF-IDF(kmer i , S K ) is the TF-IDF weight of kmer K in sequence S i .

[0039] Preferably, the algorithm of the artificial intelligence model is selected from random forest, support vector machine, convolutional neural network, Transformer, BERT, gradient boosting machine, extremely randomized tree, deep neural network, long short-term memory network, gated recurrent unit, deep residual network, recurrent neural network, variational autoencoder, generative adversarial network.

[0040] Preferably, the biological sequence is derived from an animal, a plant or a microorganism, and the specific function is selected from antibacterial drug resistance, biochemical function, immune response or metabolic pathway.

[0041] Preferably, the biological sequence is DNA and protein from a pathogenic bacterium, and the specific function is antibacterial drug resistance. The training set for training the artificial intelligence model is constructed according to the following method:

[0042] The drug resistance gene data comes from the CARD database, and the non-drug resistance gene data comes from the DEG database;

[0043] Use the CD-HIT software to perform redundancy removal on the sequences in the DEG database;

[0044] Compare the sequences in the DEG database after redundancy removal with the sequences in the CARD database, and remove the entries with high similarity to the sequences in the CARD database;

[0045] Randomly select several sequences from the remaining sequences in the DEG database as the sequences of non-drug resistance genes, and merge them with the drug resistance gene sequences in the CARD database to construct the final training set.

[0046] The present invention also provides a method for predicting the function of a biological sequence, which uses the above biological sequence function prediction system to perform the following steps:

[0047] Input the sequences of at least one biological sequence to be predicted;

[0048] Combine the biological sequence to be predicted with a fixed vocabulary including the inverse document frequency (IDF) values of terms, and process the biological sequence to be predicted through the TF-IDF method to obtain the TF-IDF vector representation of the biological sequence to be predicted;

[0049] Input the TF-IDF vector representation into an artificial intelligence model to obtain a prediction result on whether the biological sequence to be predicted has a specific biological function.

[0050] The present invention also provides a computer device, including a processor, an input device, an output device, and a memory. The processor, input device, output device, and memory are interconnected. Among them, the memory is used to store a computer program, and the computer program includes program instructions. The processor is configured to call the program instructions to execute the above method.

[0051] The present invention also provides a computer-readable storage medium, on which there is stored: a computer program for implementing the above biological sequence function prediction system, or a computer program for implementing the above biological sequence function prediction method.

[0052] In the present invention, the "fixed vocabulary" is a table containing a set of target terms and their global statistical information. One way to construct the fixed vocabulary is to generate it based on training data. Specifically, it includes the following steps: extract terms from the training data, screen the terms that meet the frequency conditions, and calculate their document frequency (DF) and inverse document frequency (IDF). As a preferred method, the fixed vocabulary remains unchanged after generation to ensure the consistency of the subsequent feature extraction process.

[0053] The present invention first applies the TF-IDF method to the characterization of biological sequences and combines the characterization results with machine learning techniques, which can efficiently and accurately predict the relationship between a specific sequence and a biological function (such as a certain gene phenotype or the biological function of a protein). Experimental results show that the method of the present invention has significantly better prediction performance than the prior art for specific prediction tasks (such as the prediction of pathogenic bacteria drug resistance genes). At the same time, the method of the present invention achieves true alignment-free, and has no restrictions on sequence and phenotype categories, and has broad application prospects.

[0054] Obviously, based on the above content of the present invention, and according to the common general technical knowledge and customary means in the art, without departing from the above basic technical idea of the present invention, various other forms of modifications, substitutions or changes can also be made.

[0055] The following is a further detailed description of the above content of the present invention through specific embodiments in the form of examples. However, this should not be construed as limiting the scope of the above subject matter of the present invention to the following examples. All technologies implemented based on the above content of the present invention fall within the scope of the present invention. Description of the Drawings

[0056] Figure 1 It is a schematic flowchart of the method for Example 1.

[0057] Figure 2 It is a schematic diagram of the construction of k-mer in Example 1. Detailed Embodiments

[0058] It should be particularly noted that the algorithms for steps such as data acquisition, transmission, storage, and processing that are not specifically described in the embodiments, as well as the hardware structures and circuit connections that are not specifically described, can all be implemented through the content disclosed in the prior art.

[0059] Biological Sequence Function Prediction Method and System of Example 1

[0060] The system provided in this embodiment includes:

[0061] An input module, configured to input the sequences of at least one biological sequence to be predicted;

[0062] A data preprocessing module, configured to process the biological sequence to be predicted by the TF-IDF method in combination with a fixed vocabulary including the inverse document frequency (IDF) values of terms, to obtain a TF-IDF vector representation of the biological sequence to be predicted;

[0063] A prediction module, configured to input the TF-IDF vector representation into an artificial intelligence model to obtain a prediction result on whether the biological sequence to be predicted has a specific function.

[0064] Through this system, a method for predicting the function of a biological sequence can be executed (as shown in Figure 1 ).

[0065] Among them, the biological sequence is selected from DNA, RNA, DNA-RNA hybrid sequences, proteins, and polypeptides.

[0066] The fixed vocabulary is a table containing a set of target terms and their global statistical information. In this embodiment, the fixed vocabulary is constructed through the following steps: extracting terms from the training data, screening the terms that meet the frequency conditions, and calculating their document frequency (DF) and inverse document frequency (IDF). As a preferred method, the fixed vocabulary remains unchanged after generation to ensure the consistency of the subsequent feature extraction process.

[0067] As a preferred method, the TF-IDF vector representation can be generated using a TF-IDF vectorizer (TfidfVectorizer, which is based on the training set and stores the TF-IDF matrix of each term in the training set).

[0068] The specific steps of the TF-IDF method are as follows:

[0069] 1. Sequence segmentation: k-mer extraction

[0070] First, given a biological sequence S, the present invention performs feature extraction by splitting the sequence into k-mers (i.e., consecutive subsequences of length k, where k = 5 for nucleic acid sequences and k = 3 for protein sequences in this embodiment). Let S be a biological sequence of length n, and all k-mers can be represented as:

[0071] kmer = {S[i:i + k - 1] | 1 ≤ i ≤ n - k + 1}

[0072] Where:

[0073] S[i:i + k - 1] represents the k-mer subsequence starting from the i-th position in the sequence S. n is the length of the sequence S, and k is the length of the k-mer.

[0074] Through this k-mer operation, the biological sequence S is converted into a feature sequence S of length K = n - k + 1 K , that is

[0075] S K = (kmer 1 , …, kmer i , …, kmer K )

[0076] Where kmer i represents the i-th kmer.

[0077] An example of this step is as Figure 2 shown.

[0078] 2. Calculate TF (term frequency)

[0079] Given a sequence S and its corresponding feature sequence S K , for kmer i, whose term frequency (TF) in sequence S K is calculated as:

[0080]

[0081] Where:

[0082] count(kmer i , S) is the number of times kmer i appears in sequence S K . K is the length of sequence S K .

[0083] This step obtains the term frequency of each k-mer by calculating the frequency of each k-mer in the current sequence.

[0084] 3. Calculate IDF (Inverse Document Frequency)

[0085] To measure the importance of each k-mer in all sequences, we calculate the inverse document frequency (IDF). Given a set D = {S 1 , S 2 , …, S N} containing N sequences, for kmer i , the frequency DF(kmer i ) with which it appears in the training set D is the number of sequences containing kmer i . Then the inverse document frequency (IDF) of kmer i is calculated as:

[0086]

[0087] Where:

[0088] DF(kmer i ) is the number of sequences containing kmer i . Using 1 + DF(kmer i ) to avoid a zero denominator. log represents taking the logarithm of the obtained value.

[0089] This step is used to measure the importance of a k-mer in the entire dataset. The more common a k-mer is, the lower its IDF value; the less common a k-mer is, the higher its IDF value. Generally, k-mers with lower occurrence frequencies are considered to have higher information content.

[0090] 4. Calculate TF-IDF

[0091] Finally, for kmer i , its TF-IDF weight is calculated by multiplying TF and IDF:

[0092] TF-IDF(kmer i , S K ) = TF(kmer i , S) · IDF(kmer i )

[0093] It can be understood that the TF-IDF value of kmer i is proportional to the number of occurrences of kmer i in the current sequence S and inversely proportional to the number of occurrences of kmer i in all sequences D, which represents the relative importance of kmer i in sequence S and also takes into account its scarcity in all sequences D.

[0094] 5. TF-IDF Vector Representation of Sequences

[0095] Assume that all sequences D contain V different kmers, i.e., D V = {kmer 1 , …, kmer i , …, kmer V}. For the d-th sequence S d in D and its corresponding characteristic sequence S d,K , construct a TF-IDF vector with a dimension size of V, where each element represents the TF-IDF value of the corresponding k-mer, expressed as:

[0096]

[0097] where:

[0098] kmer i is the i-th k-mer in D V . If kmer i does not appear in sequence S d,K , its corresponding TF-IDF value is 0.

[0099] Once each sequence is TF-IDF vectorized, these vectors can be used as features and input into an artificial intelligence classification model for model training or sequence prediction. In this embodiment, the selection of the artificial intelligence model includes, but is not limited to, random forest, support vector machine, convolutional neural network, Transformer, BERT, gradient boosting machine, extremely randomized tree, deep neural network, long short-term memory network, gated recurrent unit, deep residual network, recurrent neural network, variational autoencoder, generative adversarial network, etc. As a preferred method, random forest is selected in this embodiment.

[0100] The technical solution of the present invention will be further described through experiments below.

[0101] Experimental Example 1 Verification and Comparison of the Performance of Biological Sequence Function Prediction

[0102] I. Experimental Method

[0103] In this experimental example, taking the prediction of antibacterial drug resistance of pathogenic bacteria as an example, the prediction performance of different prediction methods was investigated.

[0104] In this experimental example, the following three experimental groups were set up:

[0105] 1. Predict according to the system and method of Example 1;

[0106] 2. Deploy the latest version of DeepARG (v1.0.2) of the prior art for prediction;

[0107] 3. Deploy the latest version of ARGNet (v1.0.1) of the prior art for prediction.

[0108] Dataset Description:

[0109] The drug resistance gene data comes from the CARD database (version 3.3.0), which integrates the ARDB database, and a total of 4,840 drug resistance gene information is collected, and the corresponding nucleic acid and protein sequences are provided at the same time.

[0110] The non-drug resistance gene data comes from the DEG database (version 2020.9.1), which records the genes necessary for bacteria to maintain life activities, and a total of 26,619 essential gene information is included, and the corresponding nucleic acid and protein sequences are also provided.

[0111] We first used the CD-HIT software to perform redundancy removal on the sequences in the DEG database (similarity threshold set to 0.9) to reduce data redundancy. Subsequently, the redundant DEG sequences were compared with the sequences in the CARD database, and the entries with similarity higher than 0.8 to the sequences in the CARD database were removed. Finally, 4,840 were randomly selected from the remaining DEG sequences as the sequences of non-drug resistance genes and merged with the drug resistance gene sequences in the CARD database to construct the final dataset.

[0112] The above processes were performed independently on nucleic acid and protein sequences.

[0113] Evaluation of the Prediction Performance of the Method:

[0114] The model learns the patterns of the sequences through the training data and is used to predict new sequences.

[0115]

[0116] Among them are prediction labels, representing two types of sequences respectively.

[0117] Subsequently, stratified K-fold cross-validation (K = 10) is performed on the dataset D:

[0118]

[0119] Among them, represents the training set of the k-th fold, represents the test set of the k-th fold. During the training process of each fold, the training set is used to train the model, and the test set

[0120] is used to calculate the corresponding evaluation metrics, including Precision, Recall, F1-Score, and Accuracy:

[0121]

[0122] represents the performance metrics (such as Precision, Recall, etc.) on the test set of the k-th fold K represents the number of folds of cross-validation (10 in the embodiment). is the finally calculated performance metric (such as average Precision, average Recall, etc.).

[0123] To ensure the fairness of the results, the same 10-fold test set data is used in this experimental example to evaluate each method.

[0124] II. Experimental Results

[0125] The prediction performance data of three methods for predicting whether a nucleic acid (DNA) sequence and a protein sequence are drug-resistant genes (i.e., there is an association between the biological sequence to be predicted and antimicrobial drug resistance) are shown in Table 1 and Table 2.

[0126] Table 1 Comparison of Prediction Performance of Nucleic Acid Sequences of Drug-Resistant Genes

[0127]

[0128] Table 2 Comparison of Prediction Performance of Protein Sequences of Drug-Resistant Genes

[0129]

[0130] From the experimental results, it can be seen that the prediction performance (Accuracy, Precision, Recall, F1) data of the system and method of Example 1 are all comprehensively better than those of the existing technologies (DeepARG or ARGNet).

[0131] As can be seen from the above embodiments and experimental examples, the present invention for the first time applies the TF-IDF method to the characterization of biological sequences, and combines the characterization results with machine learning techniques to predict the relationship between a specific sequence and a biological function (such as a certain gene phenotype or the biological function of a protein). The present invention has better prediction performance than the prior art and has good application prospects.

Claims

1. A biological sequence function prediction system, characterized in that: include: An input module, configured to input at least one biological sequence to be predicted; A data preprocessing module is configured to combine the biological sequence to be predicted with a fixed vocabulary including inverse document frequency (IDF) values ​​of terms, and process the biological sequence to be predicted by a TF-IDF method to obtain a TF-IDF vector representation of the biological sequence to be predicted; The prediction module is configured to input the TF-IDF vector representation into an artificial intelligence model to obtain a prediction result of whether the biological sequence to be predicted has a specific biological function.

2. The biological sequence function prediction system according to claim 1, characterized in that: The biological sequence is selected from DNA, RNA, DNA-RNA hybrid sequence, protein, and polypeptide.

3. The biological sequence function prediction system according to claim 1, characterized in that: The TF-IDF method is implemented by a TF-IDF vectorizer; And / or, the fixed vocabulary is constructed by the following steps: extracting terms from training data, screening terms that meet frequency conditions, and calculating document frequencies and inverse document frequencies of the terms.

4. The biological sequence function prediction system according to claim 1, characterized in that: The TF-IDF method comprises the following steps: Step 1, dividing the sequence into continuous subsequences, and the subsequences are recorded as k-mers; Step 2, calculate the word frequency TF of each k-mer in its sequence; Step 3, calculate the inverse document frequency IDF of each k-mer in the training set; Step 4, calculate the TF-IDF weight of each k-mer; Step 5: construct the TF-IDF vector representation through the TF-IDF weight of each k-mer.

5. The biological sequence function prediction system according to claim 4, characterized in that: In step 1, k-mer is represented as: kmer={S[i:i+k-1]∣1≤i≤n-k+1} Where S is a biological sequence of length n, S[i:i+k-1] represents the k-mer subsequence starting from the i-th position in sequence S, and k is the length of the k-mer; After segmentation, the sequence S is converted into a feature sequence S with a length of K = n-k+1 K ,Right now S K =(khmer1,…,khmer i ,…,Khmer K ) Among them, kmer i represents the i-th kmer; The length of the subsequence is selected from 3-21, preferably an odd number; In step 2, the calculation formula of word frequency TF is: Among them, count(kmer i ,S) is kmer i In sequence S K The number of times it appears in In step 3, the calculation formula for inverse document frequency IDF is: Where N is the number of sequences in the training set, DF(kmer i ) is the kmer contained in the training set i number of sequences; In step 4, the calculation formula of TF-IDF weight is: TF-IDF(kmer i ,S K )=TF(kmer i ,S)·IDF(kmer i ) Among them, TF-IDF (kmer i ,S K ) is the sequence S K Middle kmer i TF-IDF weight.

6. The biological sequence function prediction system according to any one of claims 1 to 5, characterized in that: The algorithm of the artificial intelligence model is selected from random forest, support vector machine, convolutional neural network, Transformer, BERT, gradient boosting machine, extreme random tree, deep neural network, long short-term memory network, gated recurrent unit, deep residual network, recurrent neural network, variational autoencoder, and generative adversarial network.

7. The biological sequence function prediction system according to any one of claims 1 to 5, characterized in that: The biological sequence is derived from animals, plants or microorganisms, and the specific function is selected from antimicrobial resistance, biochemical function, immune response or metabolic pathway.

8. A method for predicting biological sequence function, characterized in that: The biological sequence function prediction system according to any one of claims 1 to 7 is used to perform the following steps: Input at least one sequence of a biological sequence to be predicted; Combine the biological sequence to be predicted with a fixed vocabulary including the inverse document frequency (IDF) values ​​of terms, and process the biological sequence to be predicted by the TF-IDF method to obtain a TF-IDF vector representation of the biological sequence to be predicted; The TF-IDF vector representation is input into an artificial intelligence model to obtain a prediction result of whether the biological sequence to be predicted has a specific biological function.

9. A computer device, characterized in that: The method comprises a processor, an input device, an output device and a memory, which are interconnected, wherein the memory is used to store a computer program, the computer program comprises program instructions, and the processor is configured to call the program instructions to execute the biological sequence function prediction method as described in claim 8.

10. A computer-readable storage medium, characterized in that: Stored thereon are: a computer program for implementing the biological sequence function prediction system described in any one of claims 1 to 7, or a computer program for implementing the biological sequence function prediction method described in claim 8.