A method for biological sequence processing and model training
Through a biological sequence processing and model training method, using deep learning models and reverse complementary networks to process virus gene sequences, the problems of high resource consumption and difficult classification of traditional methods when dealing with newly discovered viruses are solved, and efficient and accurate gene sequence classification is achieved.
Patent Information
- Application Number
- CN202210446243.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-04-26
AI Technical Summary
Traditional biological methods require a lot of resources when dealing with newly discovered viruses, and it is difficult to accurately classify viruses that are not recorded in the database, which is very expensive.
A biological sequence processing and model training method is adopted, including obtaining biological gene sequence data, preprocessing data, building balanced data sets and training models with reverse complementary networks. This method uses deep learning models to process virus gene sequences, mine potential information, and improve model performance using base complementarity relationships.
On the basis of maintaining the classification accuracy similar to the traditional method, it significantly saves time and costs, and can correctly predict genes that cannot be correctly classified by traditional methods, improving the efficiency and accuracy of gene sequence classification.
Smart Images

Figure CN114881131B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer processing of biological gene sequences, and particularly relates to a method for biological sequence processing and model training. Background Art
[0002] Since the pneumonia epidemic caused by the novel coronavirus has been threatening the health and safety of humans recently. In fact, the novel coronavirus is just a very common one among various viruses that have emerged in human history. Currently, the viruses still raging globally include influenza virus, AIDS virus, liver disease virus, etc. Viruses have always existed in the world. With the development of humans, they have also been constantly evolving and updating. It wasn't until the late 19th century that people first recognized this tiny pathogen. These invisible viruses are affecting people's physical health all the time.
[0003] For a newly discovered virus, to figure out its origin, using traditional biological methods requires consuming a large amount of resources. Since the results obtained by traditional methods are mainly based on comparison with the data in existing databases, for some special viruses, if there is no similar record in the database, there is no way to classify them accurately. Deep learning has a wide range of application scenarios, such as in the fields of computer vision, natural language processing, speech analysis, etc.
[0004] Viral gene data is similar to text data, which is a highly serialized string of data containing some potential feature information; and the DNA sequence of an organism is a double-stranded helical structure, which means that there is a special base complementary relationship between the two DNA strands, enabling the bases on the two strands to bind to each other. Using the method of natural language processing to process viral gene sequences, mining the potential information in the sequences, and at the same time being able to make good use of the complementary relationship between base pairs, so as to train a model to process the classification of viral genes. This is a feasible method and has unique advantages compared with traditional biological methods. In the traditional biological gene classification and recognition method, the most commonly used and most effective method is BLAST. Its working principle is based on a huge biological gene database. By comparing the gene to be classified and recognized with the data in the gene database, the gene with the highest similarity in the comparison result and the gene to be classified and recognized are classified into the same category. Because it is necessary to compare the gene to be classified and recognized with some genes in the database in detail, the time cost of the BLAST method for identifying and classifying genes is relatively large. Summary of the Invention
[0005] Aiming at the deficiencies in the existing technology, the first object of the present invention is to provide a method for biological sequence processing and model training. The method proposed by the present invention can save time on the basis of achieving an accuracy level similar to that of traditional gene classification and recognition methods, and can correctly predict some genes that cannot be correctly classified by traditional biological methods.
[0006] To solve the above technical problems, the present invention is realized through the following technical solutions:
[0007] A method for biological sequence processing and model training, characterized in that it includes the following steps:
[0008] S1. Obtain the data of biological gene sequences and integrate the data;
[0009] S2. Preprocess the data, traverse the read biological gene sequences, and filter out the biological gene sequences that meet the requirements;
[0010] S3. Construct a data set required for training the model, and fine-tune the data set according to the number of data in each category in the data set to ensure that the scale of data in each category in the data set is roughly equal;
[0011] S4. Perform processing on the balance of the number of data in each category in the data set and the balance of the length of gene data to obtain a training set;
[0012] S5. Use the training set to train a model with a reverse complementary network.
[0013] Further: In the step S1, obtain the GenBank ID of the biological gene sequence; the method for obtaining the GenBank ID of the biological gene sequence: set a file containing the GenBank IDs of several biological gene sequences, query and download the biological information corresponding to the index of this item in the public biological gene database according to the GenBank IDs of the biological gene sequences listed in the file, and store it in a FASTA file, or directly obtain a FASTA file containing biological gene sequence information.
[0014] Further: Read the FASTA file, organize the information contained in the FASTA file into a table in a specified format respectively, and arrange the biological information of the same gene in the same column of the table to obtain a local database containing all the required data. Remove duplicates from the data in the local database and save it as a CSV file.
[0015] Further: In the step S2, the data preprocessing method is as follows: Call the CSV file, traverse each piece of data in the CSV file, analyze the data, and replace the unconventional bases contained in the data with base N; When N appears continuously more than 20 times in a single gene sequence or N appears non - continuously but its quantity accounts for 5% of all bases in the whole gene sequence, this piece of gene sequence data is excluded.
[0016] Further: In the step S3, construct the dataset required for training the model according to the rule of class balance:
[0017] First, determine the categories to be predicted by the prediction task of the training model, and count the number of biological gene sequences contained in each category according to these categories; The number of sequences used for each category to form the dataset is roughly equal. Therefore, randomly extract the same number of biological gene sequences from other categories according to the number of biological gene sequences contained in the category with the least number of biological gene sequences; When the difference between the number of biological gene sequences in the category with the least number of biological gene sequences and the number of biological gene sequences in other categories is greater than the number of sequences in the category with the least number of biological gene sequences itself, then select another category with an appropriate number of sequences as the benchmark and extract data from other categories; When dividing the dataset, if the data volume is greater than a certain value, then divide it according to a certain ratio, otherwise divide the dataset according to the rule that the number of training sets > the number of test sets > the number of validation sets;
[0018] After the number of training sets, test sets, and validation sets are divided, randomly shuffle the order of the three, and then write the data and their corresponding categories into different CSV files respectively, and store them in the same folder for the model to call during training.
[0019] Further: In the step S4,
[0020] (a) Duplicate and fill genes with lengths less than the top 5% in the local database length ranking to the required length: Randomly select a certain base on a gene with a length less than the top 5% in the local database length ranking as the starting position of the self - replication fragment. The gene sequence segment from the starting position to the last base of this gene sequence is the gene sequence segment used for self - replication and filling. Then fill this gene segment to the end of the original gene sequence; Repeat the above operation until the gene length reaches the required length;
[0021] (b) Expand the dataset for the class with insufficient training set data: Duplicate a part of an existing biological gene sequence and regard it as an independent biological gene sequence that can represent this class, so as to achieve the effect of balancing the dataset;
[0022] (c) Splitting the full-length gene: Using the sliding window sampling method, a gene fragment is sampled at a certain interval length. This gene fragment serves as the input data during model training and is also called a subsequence of the gene sequence. When the interval length is small enough, sufficient gene fragments can be sampled.
[0023] Furthermore, in step S5:
[0024] S01: In the transformation of biological gene sequences into digital coding expressions, use one-hot encoding, Skip-Gram, CBOW, or Elmo models to pre-train biological genes and output the vector representations of each base of the biological gene, which are used as the input data of this method;
[0025] S02: Adopt sequence reverse complement processing. During model training, input the DNA strand and its complementary strand into the model for training. Use two independent branch network structures in parallel to process two different data respectively. The weight parameters are shared between each pair of the same network layers of the two networks. Before outputting the data from the last layer, merge the data of the two strands to output the final prediction result;
[0026] S03: When training the model, flexibly adjust the parameters and train different models according to different subsequence lengths, the interval length for extracting subsequences using the sliding window, and the types of deep learning networks used for training the model; During the training process, save the model parameters with the best performance for calling the best model during model testing.
[0027] The second object of the present invention is to provide an electronic device, which is characterized in that it includes:
[0028] One or more processors;
[0029] A storage device for storing one or more programs,
[0030] When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of the above.
[0031] The third object of the present invention is to provide a computer-readable medium, on which a computer program is stored, and it is characterized in that: when the program is executed by a processor, it implements the method as described in any one of the above.
[0032] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0033] The present invention combines biological sequence data processing, model training, and biological sequence classification. The augmented dataset can effectively alleviate the problem of too small a data volume of a certain type in the dataset, enabling the model to learn the features of each category in a relatively fair environment. The present invention uses subsequences of biological sequences instead of complete sequences to train the model, which can limit the length of the sequences input into the model while increasing the scale of the dataset. At the same time, the reverse complementary relationship of DNA sequences is used to improve the performance of the model, so that the model can fully extract the intronic information of gene sequences and better develop the potential of deep learning models. The present invention uses a deep learning model to replace the traditional biological gene sequence classification method, which can improve the classification accuracy while reducing the time cost. The model of the present invention can achieve good results when tested on datasets composed of different data, so it can also be easily extended to other biological sequence classification tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is the overall framework diagram of the present invention;
[0035] Figure 2 is the schematic diagram of data preprocessing in the present invention;
[0036] Figure 3 is the structural diagram of the reverse complementary network adopted by the present invention;
[0037] Figure 4 is the flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the preferred implementation embodiments of the present invention will be described below in conjunction with specific embodiments. However, it should be understood that the drawings are only for illustrative purposes and cannot be construed as a limitation of the present invention; for better illustration of this embodiment, some components in the drawings will be omitted, enlarged, or reduced, and do not represent the actual size of the product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted. The positional relationships described in the drawings are only for illustrative purposes and cannot be construed as a limitation of the present invention.
[0039] The present invention will be further described below in conjunction with the drawings and embodiments, but it shall not be used as a basis for limiting the present invention.
[0040] As Figures 1 to 4 shown, a biological sequence processing and model training method includes the following steps:
[0041] S1. Obtain the data of biological gene sequences and integrate the data;
[0042] S2. Preprocess the data by traversing the read biological gene sequences and filtering out the biological gene sequences that meet the requirements;
[0043] S3. Construct the dataset required for training the model, and fine-tune the dataset according to the number of data in each category in the dataset to ensure that the scales of various types of data in the dataset are roughly equal;
[0044] S4. Perform processing on the data in the dataset to balance the number of various types of data in the dataset and the length of gene data to obtain the training set;
[0045] S5. Use the training set to train a model with a reverse complementary network.
[0046] In step S1, obtain the GenBank ID of the biological gene sequence; the method for obtaining the GenBank ID of the biological gene sequence: set a file containing the GenBank IDs of several biological gene sequences, query and download the biological information corresponding to this index in the public biological gene database according to the GenBank IDs of the biological gene sequences listed in the file, and store it in a FASTA file, or directly obtain a FASTA file containing biological gene sequence information.
[0047] In step S1, download gene data according to the GenBank ID to be downloaded from a public database (GenBank is an open-access sequence database that collects and annotates all publicly available nucleotide sequences and their translated proteins. This database is part of the International Nucleotide Sequence Database Collaboration (INSDC)).
[0048] The download channels include: logging in to the NCBI website and downloading according to the steps prompted on the website; or using the API built in the BioPython package for downloading. Among them, when downloading biological gene sequences, the GenBank IDs of these biological sequences need to be written in a txt file in order, and only one GenBank ID can be written in each line, and the downloaded gene sequence data is stored in a FASTA file. In this embodiment, the dataset to be obtained includes AlphaVirus data, Flavivirus data, and COVID-19 data. In the present invention, the virus data is classified according to the host or transmission vector it infects, so AlphaVirus can be divided into the following categories: "Barmah Forest virus", "Chikungunya virus", "Eastern equine encephalitis virus", "Getah virus", "Madariaga virus", "Mayaro virus", "Sindbis virus", "Venezuelan equine encephalitis virus", and "Western equine encephalitis virus". Flavivirus and COVID-19 can also be classified in this way.
[0049] Read the FASTA file, organize the information contained in the FASTA file into a table according to the specified format, and arrange the biological information of the same gene in the same column of the table to obtain a local database containing all the required data. The local database is used for subsequent model training and testing; the data in the local database is deduplicated and saved as a CSV file for future calls and searches. In addition, the CSV file can also be used as a backup of the training model data to facilitate the modification of the database when modifying the model in the future.
[0050] In the step S2, the data preprocessing method is as follows: calling the CSV file, traversing each data in the CSV file, analyzing the data, if the data contains unconventional bases, that is, bases other than A, T, C, G, U and unknown base N, the unconventional bases contained in the data are replaced with base N; when N appears more than 20 times continuously in a single gene sequence or N appears more than 20 times non-continuously but its number occupies 5% of all bases in the entire gene sequence, it is considered that the data is not suitable for model training. If the impact of deleting the gene on the size of the data set is within 1%, the data in which N appears more than 20 times continuously in a single gene sequence or N appears more than 20 times non-continuously but its number occupies 5% of all bases in the entire gene sequence can be eliminated, thereby optimizing the training effect of the model.
[0051] In step S3, the data set required for the training model is constructed according to the class balance rule:
[0052] First, determine the categories that need to be predicted for the prediction task of the training model, and count the number of biological gene sequences contained in each category according to these categories; the number of sequences used to form the data set in each category is roughly equal, so the same number of biological gene sequences are randomly extracted from other categories according to the number of biological gene sequences contained in the class with the least number of biological gene sequences; when the difference between the number of biological gene sequences in the class with the least number of biological gene sequences and the number of sequence biological genes in other classes is greater than the number of sequences in the class with the least number of biological gene sequences itself, another class with a suitable number of sequences is selected as a benchmark to extract data from other classes; when dividing the data set, when the amount of data is greater than a certain value, generally the threshold value is about 10, then it is divided according to a certain ratio. The commonly used data set division ratio is 6:2:2 or 7:2:1, otherwise the data set is divided according to the rule of number of training sets> number of test sets> number of validation sets;
[0053] After the number of training sets, test sets, and validation sets are divided, the order of the three is randomly shuffled, and then the data and its corresponding categories are written into different CSV files and stored in the same folder for call during model training.
[0054] In this embodiment, after filtering the data, it is necessary to construct the data set required for training the model according to the rule of class balance, that is, if a large class of data has many sub-class branches, it is also necessary to evenly select the sub-classes when selecting the data set. As mentioned above, "Eastern equine encephalitis virus" can be divided into 3 categories according to the level of pathogenicity. In order to make the information learned by the model as complete and fair as possible, when extracting the data of "Eastern equine encephalitis virus", the same number of genes with different pathogenicities should be extracted respectively.
[0055] According to the types of hosts infected by alphaviruses, alphaviruses can be divided into 9 categories, and the number of sequences extracted from each category to form the data set should be similar. In the alphavirus data set, the category with the largest amount of data has 726 gene data, while the category with the least amount of data has only 18 gene data. Therefore, the same number of biological gene sequences are randomly extracted from other categories according to the number of biological gene sequences contained in the category with the least number of biological gene sequences, that is, 18 are extracted from each category. If the difference between the number of biological gene sequences in the category with the least number of biological gene sequences and the number of biological gene sequences in other categories is too large and the number of gene sequences in this category is less than 10, then another category with a suitable number of sequences is selected as the benchmark, randomly extract data from the categories with more data than this category, expand the data set for the categories with less data than this category, and divide the data set.
[0056] After the three data sets are divided, the data sets are randomly shuffled, and then the data and their corresponding categories are written into different CSV files, named x_train, y_train, x_validation, y_validation, x_test and y_test respectively, and stored in the same folder for calling during model training.
[0057] In step S4, (a) Genes with a length less than the top 5% in the length ranking in the local database are replicated and filled to the required length (since the lengths of biological gene sequences are all different, genes with too small lengths need to be replicated and filled to a certain length): Randomly select a certain base on a gene with a length less than the top 5% in the length ranking in the local database as the starting position of the self-replicating fragment. The gene sequence segment from the starting position to the last base of this gene sequence is the gene sequence segment used for self-replicating filling, and then this gene segment is filled to the end of the original gene sequence; repeat the above operation until the length of the gene reaches the required length;
[0058] (b) Expand the dataset for classes with insufficient training set data: copy a part of an existing biological gene sequence and treat it as an independent biological gene sequence that can represent this class, so as to achieve the effect of balancing the dataset; for the training set used to train the model, the different number of biological gene sequences in each category will cause the trained model to have a preference for the class with more data in the training set, that is, it is easier to classify the data into this class, resulting in unfairness. Therefore, it is necessary to expand the dataset for classes with insufficient training set data.
[0059] (c) Split the full-length gene: Use the sliding window sampling method to sample a gene fragment at a certain interval. This gene fragment is used as the input data for model training, also known as a subsequence of the gene sequence. When the interval length is small enough, sufficient gene fragments can be sampled. Figure 2 As shown in the figure, the sliding window method is used to extract data, and data with a length of 5 bases is intercepted every 3 bases as an independent gene sequence data, also known as a subsequence. Under the premise of ensuring sufficient data training model and not introducing data not from the original gene sequence, the scale of data input to the model can be reduced, and the potential of the deep learning model can be fully explored to obtain a model with higher fitting degree.
[0060] The length of the input data of the deep learning model is limited. Inputting a complete gene into the model will result in too many parameters, so the entire gene needs to be split into several small gene fragments. Different gene fragment lengths will also affect model training. In order to ensure that the gene fragments obtained by segmentation are all derived from the corresponding original genes in the data set, and to obtain as many gene fragments as possible, it is necessary to segment the full-length gene.
[0061] In the step S5, S01: in converting the biological gene sequence into a digital code expression, one-hot coding, Skip-Gram, CBOW or Elmo model is used to pre-train the biological gene, and the vector representation of each base of the biological gene is output, which is used as the input data of the method; pre-training using a pre-trained model can better translate the original base sequence into a computer-recognizable digital sequence, and using different high-dimensional vectors to represent different bases is better than the traditional one-hot code representation, so that the performance of the trained model can be improved.
[0062] like Figure 1 As shown, a pre-trained model is used for pre-training to build a "dictionary" of the connection between bases (pairs) and vectors. The vectors of the dictionary are used instead of the one-hot code, which can better translate the original base sequence into a computer-recognizable digital sequence. Using different high-dimensional vectors to represent different bases is better than the traditional one-hot code representation, so that the performance of the trained model can be improved.
[0063] S02: The reverse complementary network is used because the DNA of organisms has a regular double helix structure in space, where two single strands are bound together through base complementary pairing, that is, bases A and T, C and G can be complementary paired and bound together, so that the two single strands finally combine into a double helix structure. To solve this problem, the present invention adopts sequence reverse complementary processing. During model training, the DNA strand and its complementary strand are input into the model for training simultaneously. Two independent branch network structures are used in parallel to process two different data respectively. The weight parameters are shared between each pair of identical network layers of the two networks. Before outputting the data from the last layer, the data of the two strands are merged to output the final prediction result;
[0064] S03: When training the model, according to different subsequence lengths, the interval length for extracting subsequences by the sliding window, and the type of deep learning network used for training the model, the parameters are flexibly adjusted and different models are trained; during the training process, the model parameters with the best performance are saved for calling the best model during model testing.
[0065] Using the dataset of alphavirus input by the present invention to train the model, the length of the subsequence is set to 200, the interval for extracting subsequences by the sliding window is set to 1, a pre-trained model is used to vectorize base pairs, and a reverse complementary model is adopted. A 9-classification task is performed on the test data in the test set, and the classification accuracy rate is as high as 94.38%, which is 5.68% higher than the result obtained by testing using the method without the present invention. When the length of the subsequence is reduced to 120, the 9-classification test accuracy rate is as high as 97.32%. Adjusting the length of the subsequence also affects the prediction accuracy rate of the model. Therefore, the parameters can be flexibly adjusted to find a set of parameters that are most suitable for the dataset used in training the model.
[0066] Through the description of the above embodiments, those skilled in the art can clearly understand that the facilities of the present invention can be implemented by means of software plus a necessary general hardware platform. The embodiments of the present invention can be implemented using existing processors, or by dedicated processors used for this purpose or other purposes in a suitable system, or by a hardwired system. The embodiments of the present invention also include non-transitory computer-readable storage media, which include machine-readable media for carrying or having machine-executable instructions or data structures stored thereon; such machine-readable media can be any available medium accessible by a general or special-purpose computer or other machine having a processor. For example, such machine-readable media can include RAM, ROM, EPROM, EEPROM, CD-ROM or other optical disk memories, magnetic disk memories or other magnetic storage devices, or any other medium that can be used to carry or store the required program code in the form of machine-executable instructions or data structures and can be accessed by a general or special-purpose computer or other machine with a processor. When information is transmitted or provided to a machine through a network or other communication connection (hardwired, wireless or a combination of hardwired and wireless), this connection is also regarded as a machine-readable medium.
[0067] Based on the description and drawings of the present invention, those skilled in the art can easily manufacture or use a method for biological sequence processing and model training of the present invention, and can produce the positive effects recorded in the present invention.
[0068] The above are only the preferred embodiments of the present invention, and do not impose any formal limitations on the present invention. Any simple modifications and equivalent changes made to the above embodiments based on the technical essence of the present invention all fall within the protection scope of the present invention.
Claims
1. A method for biological sequence processing and model training, characterized in that: It includes the following steps: S1. Obtain the data of biological gene sequences and integrate the data; S2. Preprocess the data, traverse the read biological gene sequences, and filter out the biological gene sequences that meet the requirements; S3. Construct the dataset required for training the model, and fine-tune the dataset according to the number of each category of data in the dataset to ensure that the scales of various types of data in the dataset are roughly equal; S4. Perform processing on the quantity balance of various types of data in the dataset and the length balance of gene data to obtain the training set; S5. Use the training set to train the model with a reverse complementary network; In step S4, (a) Duplicate and fill the genes with lengths less than the top 5% in the length ranking in the local database to the required length: Randomly select a certain base on a gene with a length less than the top 5% in the length ranking in the local database as the starting position of the self-replicating fragment. The gene sequence segment from the starting position to the last base of this gene sequence is the gene sequence segment used for self-replicating filling. Then fill this gene segment to the end of the original gene sequence; repeat the above operation until the length of the gene reaches the required length; (b) Expand the dataset for the classes with insufficient training set data: Duplicate a part of an existing biological gene sequence and regard it as an independent biological gene sequence that can represent this class, so as to achieve the effect of balancing the dataset; (c) Split the genes with complete lengths: Use the sliding window sampling method to sample a gene segment at a certain interval length. This gene segment is used as the input data during model training and is also called the subsequence of the gene sequence. When the interval length is small enough, sufficient gene segments can be sampled; In step S5, S01: In the expression of converting biological gene sequences into digital encoding, use one-hot encoding, Skip-Gram, CBOW or Elmo model to pre-train the biological genes and output the vector representation of each base of the biological genes, which is used as the input data of this method; S02: Adopt sequence reverse complementary processing. Input the DNA strand and its complementary strand into the model for training while training the model. Use two independent branch network structures in parallel to process two different data respectively. The weight parameters are shared between each pair of the same network layers of the two networks. Before outputting the data from the last layer, merge the data of the two strands to output the final prediction result; S03: When training the model, flexibly adjust the parameters and train different models according to different subsequence lengths, the interval length for extracting subsequences by the sliding window, and the types of deep learning networks used for training the model; During the training process, save the model parameters with the best performance for calling the best model during model testing.
2. The method for processing biological sequences and training models according to claim 1, wherein: In the step S1, obtain the GenBank ID of the biological gene sequence; method for obtaining the GenBank ID of the biological gene sequence: set a file containing the GenBank IDs of a number of biological gene sequences, query and download the biological information corresponding to this index in the public biological gene database according to the GenBank IDs of the biological gene sequences listed in the file, and store it in a FASTA file, or directly obtain a FASTA file containing biological gene sequence information.
3. A method for biological sequence processing and model training according to claim 2, characterized in that: Read the FASTA file, organize the information contained in the FASTA file into a table respectively according to the specified format, and arrange the biological information of the same gene in the same column of the table to obtain a local database containing all the required data. Remove duplicates from the data in the local database and save it as a CSV file.
4. A method for biological sequence processing and model training according to claim 3, characterized in that: In the step S2, data preprocessing method: call the CSV file, traverse each piece of data in the CSV file, analyze the data, and replace the unconventional bases contained in the data with base N; when N appears continuously more than 20 times or non-continuously more than 20 times in a single gene sequence but its quantity occupies 5% of all bases in the whole gene sequence, remove the data of this gene sequence.
5. A method for biological sequence processing and model training according to claim 1, characterized in that: In the step S3, construct the dataset required for training the model according to the rule of class balance: First, determine the categories to be predicted for the prediction task of the training model, and count the number of biological gene sequences contained in each category according to these categories; the number of sequences used for each category to form the dataset is roughly equal. Therefore, randomly extract the same number of biological gene sequences from other categories according to the number of biological gene sequences contained in the category with the least number of biological gene sequences; when the difference between the number of biological gene sequences in the category with the least number of biological gene sequences and the number of biological gene sequences in other categories is greater than the number of sequences in the category with the least number of biological gene sequences itself, select another category with an appropriate number of sequences as the benchmark and extract data from other categories; when dividing the dataset, when the data volume is greater than a certain value, divide it according to a certain ratio, otherwise divide the dataset according to the rule that the number of training sets > the number of test sets > the number of validation sets; After the number of training sets, test sets, and validation sets are divided, shuffle the three randomly, and then write the data and its corresponding categories into different CSV files respectively, and store them in the same folder for the model to call during training.
6. An electronic device, characterized in that: Including: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-5.
7. A computer-readable medium having a computer program stored thereon, characterized in that: When the program is executed by the processor, it implements the method according to any one of claims 1-5.
Citation Information
Patent Citations
Gene sequence identification method and system, and computer readable storage medium
CN110070914A
Cancer type prediction system and method based on tissue and organ differentiation hierarchical relationship
CN110706749A