Deep learning Gap filling method for plant genome

Through the deep learning method DL-GapFilling, the prediction of gap sequences is solved by using the DFillingNet model and the BSCEA algorithm, and the problem of insufficient efficiency and accuracy of existing genome assembly tools in gap filling is solved, and the gap filling rate and accuracy rate are significantly improved.

CN120089202APending Publication Date: 2025-06-03NORTHEAST FORESTRY UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510045344.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Existing genome assembly tools have problems with insufficient efficiency and accuracy in filling gaps, especially when dealing with complex repetitive and mutated regions.

Method used

A DL-GapFilling method based on deep learning is proposed. The gap sequence is predicted through the DFillingNet model and the BeamStarContraction-Expansion Algorithm (BSCEA) algorithm, and genome assembly is combined with traditional tools to improve the gap fill rate and accuracy rate.

Benefits of technology

It significantly improves the efficiency and accuracy of traditional software in gap filling, increases the length of contig and scaffold, can span complex repetitive and mutated regions, and improves the integrity of genome assembly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005238236520000081
    Figure BDA0005238236520000081
  • Figure HDA0005238236530000011
    Figure HDA0005238236530000011
  • Figure HDA0005238236530000021
    Figure HDA0005238236530000021
Patent Text Reader

Abstract

The present invention relates to a gap filling method based on a DL-GapFilling model. Comprising the following steps of: 1, pre-assembling an original genome by using a traditional tool to generate a sacffard file to obtain an assembling result 1; 2, performing deep learning on the data set by using a DFillingNet model network, predicting a gap sequence by using a BeamStar Control-Extension Algorithm (BSCEA) algorithm, and then performing genome assembly on the predicted gene sequence and an original genome together, so as to obtain an assembly result 2; and 3, using Quast, Exonerate and the like to compare the assembly results 1 and 2 to a reference genome, and optimizing the prediction sequence generated in the step 2 by constructing a PredictionFilter mechanism for screening and optimizing the prediction sequence. And finally, carrying out genome packaging on the screened prediction sequence and an original genome. According to the method, a new prediction sequence is provided for traditional software, a PredictionFilter mechanism aiming at the gene sequence is provided, and the gap filling rate and accuracy of the traditional software are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field:

[0001] The present invention relates to the field of bioinformatics gene assembly, and particularly to a deep learning Gap filling method for plant genomes. Background Art:

[0002] The advent and popularization of the second-generation and third-generation sequencing technologies have made sequencing more convenient and faster, thus enabling the determination of reference genomes of more and more organisms. Next-generation sequencing technology has higher throughput than the first-generation sequencing and can generate a large number of short read-length data simultaneously. It is applicable to high-throughput sequencing and genome assembly due to its high accuracy. However, due to insufficient sequencing depth, repetitive sequence structures, and sequencing technology limitations, there are usually undetermined sequence regions between two contigs or scaffolds in genome assembly, which are blank regions or gaps that are not sequenced or cannot be spliced. Filling gaps, as an important part of gene assembly, usually requires additional sequencing data or computational prediction methods to improve the integrity of the assembly. Currently, the main methods [1] for filling gaps include using multiple software for assembly, using the reference genomes of closely related species, amplifying the gap ends using polymerase chain reaction, and using improved assembly methods based on DBG, etc.

[0003] In the field of genome assembly, many tools have been developed to meet the needs of different sequencing technologies. For second-generation sequencing (short read-length) data, tools such as Velvet, ABySS, ABySS2, SPAdes, IDBA-UD, ALLPATHS, ALLPATHS-LG, SOAPdenovo2, and SKESA, etc., focus on processing high-throughput short read-length data and can effectively handle the problems of gaps and repetitive regions that may be encountered during the assembly of short read-lengths. These tools utilize the high coverage of short read-length data to provide high-quality assembly results. In addition, to further improve the quality of the assembly, specialized gap filling tools such as HaploMerger2, GapFiller, TIGRA, Pilon, and FGAP, etc., have been designed to fill the gaps during the assembly process and utilize the paired-end information and local consistency of short read-lengths to optimize the continuity of the genome.

[0004] With the emergence of third-generation sequencing technologies (such as PacBio SMRT and Oxford Nanopore), long-read assembly tools such as Canu, HiCanu, Flye, RFfiller, Wtdbg2, and Shasta have emerged. For the gaps generated during their assembly, there are tools such as Sealer, TGS-GapCloser, Racon, PBJelly, and LR Gapcloser. These tools are good at processing long-read data and can span complex repetitive regions, thus significantly improving the integrity and continuity of the genome. To better combine the advantages of short-read and long-read data, hybrid assembly tools such as HybridSPAde, OPERA-LG, and MaSuRCA have emerged. These tools integrate different types of sequencing data, making the most of the high accuracy of short reads and the coverage advantage of long reads, providing a more reliable basis for genome assembly. With the continuous development of these tools, the efficiency and quality of genome assembly have been significantly improved, providing strong support for genome research. Most of the above tools are basically developed based on algorithms, and a small number combine deep learning, such as Gappredict and DLGapcloser. However, most gene assembly tools have not well explored the broad application prospects of deep learning in gene assembly, nor have they combined the physical properties of gene sequences themselves. Summary of the Invention:

[0005] To overcome the deficiencies and drawbacks of existing methods, the present invention proposes a DL-GapFilling method. The predicted sequences generated by it will be used together with the original sequences as the original materials for gene assembly, helping traditional filling software fill more gaps. The increase in data sources greatly improves the overlap degree of reads splicing, which can increase the lengths of subsequent contigs and scaffolds, bringing the possibility of spanning repetitive regions and variant regions, and thus significantly improving the gap filling rate and accuracy of traditional software. We tested our method on five gene sequences of Micromonas pusilla, Utricularia gibba, Eutrema salsugineum, Eutrema salsugineum, Thalictrum thalictroides, and Oryza longistaminata, and compared it with Sealer, GapPredict, and DLGapCloser.

[0006] The DL-GapFilling filling method based on deep learning includes the following steps:

[0007] Step 1: Use traditional tools to pre-assemble the original genome, generate a scaffold file, and obtain Assembly Result 1;

[0008] Step 2: Use the DFillingNet model network to perform deep learning on the dataset, and use the BeamStarContraction-Expansion Algorithm (BSCEA) to predict the gap sequence. Then, assemble the predicted gene sequence and the original genome together to obtain Assembly Result 2;

[0009] Step 3: Use Quast, Exonerate, etc. to align Assembly Results 1 and 2 to the reference genome. By constructing a PredictionFilter mechanism optimized for screening the predicted sequences, optimize the predicted sequences generated in Step 2. Finally, assemble the screened predicted sequences and the original genome.

[0010] Step 1 specifically includes the following steps:

[0011] Step 1.1: Use traditional assembly software to pre-assemble the short-read sequencing data to form a scaffold file;

[0012] Step 1.2: Obtain the gap information from the scaffold file. After processing with awk / sed, sort through SAMTools and widen the gap sequence with BEDtools, etc. Use tools such as the HomoloGene online database, OrthoMCL, or OrthoFinder to determine the homologous genomes of the genes to be tested;

[0013] Step 1.3:

[0014] Find the read data in the original genome and the homologous genomes that map to the two flanking segments. Screen and trim the matching read information to form the input data required for deep learning. Use Biobloommicategorizer to classify the read lengths, and biobloomimaker to create an index. Thus, construct a mapping file, and then expand the sequence information on both sides of the gap, namely Flank. Classify the Gap flank data into two parts, namely the Fixed file that can be filled by traditional tools and the Unfixed that cannot be filled by traditional tools.

[0015] Thus, the training dataset for the DFillingNet model network is constructed.

[0016] Step 2 specifically includes the following steps:

[0017] Step 2.1:;

[0018] Step 2.2: Send the numerical data into the DFillingNet network architecture. Here, the BeamStarContraction-Expansion Algorithm (BSCEA) is constructed to search for the subsequent expansion nodes of the predicted sequence. First, an OpenList is constructed mainly for storing the currently to-be-expanded nodes. Each node contains the cost of the current path, the heuristic value, and the current base sequence. The ClosedList is used to store the nodes that have been visited to avoid repeated expansion. The nextLevelOpenList is used to store the newly expanded nodes. During initialization, the seed sequence is placed in the OpenList with the initial cost and heuristic value as the first to-be-expanded node for the beam search. When the OpenList is not empty and the entire sequence prediction has not been completed, the algorithm enters a loop. In each iteration, the algorithm selects several nodes with the minimum cost from the OpenList and expands these nodes. Throughout the search process, the beam width (beamWidth) is a dynamically changing parameter that determines the number of candidate nodes retained in each layer of the search. For each node selected from the OpenList, the model makes a prediction based on the sequence of the current node, that is, calculates the next possible base based on this sequence. The output of the model is a probability distribution indicating the occurrence probability of each possible base at this position. According to the prediction results of the model, the algorithm attempts to expand new bases one by one and calculates the new cost newG and the new heuristic value newH for each newly expanded node.

[0019] Step 2.3: After several rounds of the expansion phase, a large number of expanded nodes will be obtained. At this time, these nodes will be trimmed and optimized, that is, enter the contraction phase. Here, attention needs to be paid to the Limit size operation on the nodes. The Limit size operation is affected by the beam width (beamWidth) and needs to be dynamically adjusted continuously. Then, the nodes constrained by the beam are sent into the nextLevelOpenList, and the nextLevelOpenList is used to store the expanded nodes. Repeat the above operations until the storage of the OpenList reaches the set value, and the predicted sequence is obtained. Assemble the predicted gene sequence and the original genome together to obtain the assembly result 2.

[0020] Step 3 specifically includes the following steps:

[0021] Step 3.1: Align the assembly file 1 and assembly file 2 obtained in Steps 1 and 2 to the reference genome using the alignment software Quast respectively to obtain the assembly effect 1 and assembly effect 2.

[0022] Step 3.2: Use Exonerate to customize the assembly evaluation index;

[0023] Step 3.3: Since the quality of the predicted sequences generated in Step 2 is uneven, a PredictionFilter mechanism for screening and optimizing the predicted sequences is constructed. Pay attention to indicators such as the alignment rate of the new evaluation indicators, and eliminate the predicted sequences corresponding to the assembly effect 2 that is inferior to the assembly effect 1. In this way, the new predicted sequences after screening can be obtained. At this time, the optimized predicted sequences are assembled together with the original data.

[0024] Advantages of the present invention: The present invention proposes a brand-new method, namely the deep learning method DL-GapFilling for plants, aiming to assist traditional tools to further fill the gaps in small genomes. 1. A brand-new plant data set for judging the effect of genome gap filling is constructed, providing a good reference for subsequent plant genome gap filling work. The DFillingNet network model is constructed, and a brand-new gene general prediction algorithm - BeamStarContraction-Expansion Algorithm (BSCEA) algorithm is developed. The DFillingNet model greatly improves the feature extraction ability of gene sequences and the context learning ability of sequences on both sides of the gap. The BSCEA algorithm first applies the A* algorithm combined with the BS algorithm to the field of gene sequence prediction. Through an effective beam adjustment strategy, the BSCEA algorithm has great superiority in pruning processing problems and memory occupancy problems. A brand-new gene assembly evaluation standard is proposed, which has better universality for the evaluation standard of gene assembly. A correction mechanism PredictionFilter mechanism for predicted sequences is proposed. Combined with the evaluation method, it can perform secondary error correction screening on the predicted sequences generated by deep learning. By eliminating the predicted sequences that may contaminate the assembly effect, the quality of the extended predicted sequences is further improved, thereby increasing the gap filling rate. Brief Description of the Drawings:

[0025] Figure 1 It is the overall flowchart of a gap filling method based on DL-GapFilling.

[0026] Figure 2 It is the overall framework diagram of the BeamStar Contraction-Expansion Algorithm (BSCEA) algorithm.

[0027] Figure 3 It is the schematic diagram of the PredictionFilter mechanism for screening and optimizing the predicted sequences. Detailed Implementation Manner:

[0028] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0029] Figure 1 is the overall flowchart of the implementation of the present invention, Figure 2 is the structural diagram of the algorithm used, Figure 3 is the schematic diagram of the PredictionFilter mechanism, as Figure 1 Figure 2 Figure 3 shown, the method includes:

[0030] Step 1: Use traditional tools to pre-assemble the original genome to generate a scaffold file and obtain Assembly Result 1;

[0031] Step 2: Use the DFillingNet model network to perform deep learning on the dataset, and use the BeamStarContraction-Expansion Algorithm (BSCEA) to predict the gap sequence. Then, perform gene assembly on the predicted gene sequence and the original genome together to obtain Assembly Result 2;

[0032] Step 3: Use Quast, Exonerate, etc. to align Assembly Results 1 and 2 to the reference genome. By constructing a PredictionFilter mechanism optimized for screening the predicted sequence, optimize the predicted sequence generated in Step 2. Finally, perform gene assembly on the screened predicted sequence and the original genome.

[0033] Step 1 specifically includes the following steps:

[0034] Step 1.1: Use traditional assembly software to pre-assemble the short-read sequencing data to form a scaffold file;

[0035] Step 1.2: Obtain the gap information from the scaffold file. After processing by awk / sed, through sorting by SAMTools and widening the gap sequence by BEDtools, etc., use tools such as the HomoloGene online database or OrthoMCL or OrthoFinder to determine the homologous genome of the gene to be tested;

[0036] Step 1.3:

[0037] Find the read data mapped to the fragments on both sides of the original genome and the homologous genome, filter and trim the matched read information to form the input data required for deep learning. Use Biobloommicategorizer to classify the read length, biobloomimaker to create an index, and then build a mapping file, and then expand the sequence information on both sides of the gap, that is, Flank. Classify the Gap flank data into two parts, namely Fixed files that can be filled by traditional tools and Unfixed files that cannot be filled by traditional tools.

[0038] So far, the training dataset of the DFillingNet model network has been constructed.

[0039] Step 2 specifically includes the following steps:

[0040] Step 2.1: Use one-hot encoding to convert the gene sequence into a digital format that can be processed by computers;

[0041] Step 2.2: Feed the numerical data into the DFillingNet network architecture. Here we construct the BeamStarContraction-Expansion Algorithm (BSCEA) algorithm to search for subsequent expansion nodes of the predicted sequence. First, we construct the OpenList, which is mainly used to store the current nodes to be expanded. Each node contains the cost of the current path, the heuristic value, and the current base sequence. The ClosedList is used to store the nodes that have been visited to avoid repeated expansion. The nextLevelOpenList is used to store the new nodes after expansion. During initialization, the seed sequence is put into the OpenList with the initial cost and heuristic value as the first node to be expanded in the beam search. When the OpenList is not empty and the prediction of the entire sequence has not been completed, the algorithm will enter a loop. In each iteration, the algorithm selects several nodes with the lowest cost from the OpenList and expands these nodes. During the entire search process, the beam width (beamWidth) is a dynamically changing parameter that determines the number of candidate nodes retained in each level of search. For each node selected from the OpenList, the model will predict based on the sequence of the current node, that is, calculate the next possible base based on the sequence. The output of the model is a probability distribution, which indicates the probability of each possible base at that position. According to the prediction results of the model, the algorithm tries to expand new bases one by one and calculates the new cost newG and new heuristic value newH for each newly expanded node.

[0042] Step 2.3: After several rounds of the Expansion phase, we will obtain a large number of expanded nodes. At this time, these nodes will be trimmed and optimized, that is, enter the Contraction phase. Here, it should be noted that the Limit size operation is performed on the nodes. The Limit size operation is affected by the beamWidth and needs to be dynamically adjusted continuously. Then, the nodes that have passed the beam constraint are sent to the nextLevelOpenList, and the nextLevelOpenList is used to store the expanded nodes. Repeat the above operations until the storage of the OpenList reaches the set value, and the predicted sequence is obtained. The predicted gene sequence and the original genome are assembled together to obtain the assembly result 2.

[0043] Step 3 specifically includes the following steps:

[0044] Step 3.1: Align the assembly file 1 and the assembly file 2 obtained in Steps 1 and 2 to the reference genome using the alignment software Quast respectively to obtain the assembly effect 1 and the assembly effect 2.

[0045] Step 3.2: Use Exonerate to customize the assembly evaluation index;

[0046] Step 3.3: Since the quality of the predicted sequences generated in Step 2 is uneven, we have constructed a PredictionFilter mechanism for screening and optimizing the predicted sequences. Pay attention to indicators such as the alignment rate of the new evaluation index, and eliminate the predicted sequences corresponding to the case where the assembly effect 2 is inferior to the assembly effect 1. In this way, the new predicted sequences after screening can be obtained. At this time, the optimized predicted sequences and the original data are assembled together.

[0047] Table 1 shows the comparison of the method of the present invention with other advanced methods in experiments, and the results fully demonstrate the superiority of the method of the present invention in the gap filling task.

[0048]

[0049] It should be understood that the parts not elaborated in detail in this specification belong to the prior art.

[0050] As described above in conjunction with the accompanying drawings, this is only the specific implementation manner and process of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art should understand that this is only an example, and various changes and substitutions can be made to this implementation manner without departing from the essence of the present invention. The scope of the present invention is only defined by the appended claims.

[0051] The embodiments described with reference to the accompanying drawings are exemplary only for the purpose of explaining the present invention and should not be construed as limiting the present invention. The specific scope of the embodiments of the present invention is not limited by this. On the contrary, all embodiments of the present invention include all changes and modifications that fall within the spirit and scope of the appended claims.

Claims

1. A deep learning gap filling method DL-GapFilling for plant genomes, comprising the following steps: Step 1: Use traditional tools to pre-assemble the original genome, generate a sacffold file, and obtain assembly result 1; Step 2: Use the DFillingNet model network to perform deep learning on the data set, and use the BeamStar Contraction-Expansion Algorithm (BSCEA) algorithm to predict the gap sequence, and then assemble the predicted gene sequence together with the original genome to obtain assembly result 2; Step 3: Use Quast, Exonerate, etc. to align assembly results 1 and 2 to the reference genome, and optimize the predicted sequence generated in step 2 by building a PredictionFilter mechanism for predictive sequence screening and optimization. Finally, the screened predicted sequence is assembled with the original genome. Step 1 specifically includes the following steps: Step 1.1: Use traditional assembly software to pre-assemble short-read sequencing data to form a scaffold file; Step 1.2: Get the gap information from the sacffold file, process it with awk / sed, sort it with SAMTools and widen the gap sequence with BEDtools, and use the HomoloGene online database or OrthoMCL or OrthoFinder to determine the homologous gene set of the gene to be tested; Step 1.3: Find the read data mapped to the fragments on both sides in the original genome and homologous genome, filter and trim the matched read information to form the input data required for deep learning. Use Biobloommicategorizer to classify the read length and biobloomimaker to create an index, thereby building a mapping file, and then expand the sequence information on both sides of the gap, namely Flank. Classify the Gap flank data into two parts, namely Fixed files that can be filled by traditional tools and Unfixed files that cannot be filled by traditional tools. So far, the training data set of the DFillingNet model network has been constructed. Step 2 specifically includes the following steps: Step 2.1: Use one-hot encoding on the gene sequence to convert the L-length gene sequence into a 4×L digital matrix; Step 2.2: Feed the numerical data into the DFillingNet network architecture. Here, the BeamStar Contraction-Expansion Algorithm (BSCEA) algorithm is constructed to search for subsequent expansion nodes of the predicted sequence. First, the OpenList is constructed to store the current nodes to be expanded. Each node contains the cost of the current path, the heuristic value and the current base sequence. The ClosedList is used to store the nodes that have been visited to avoid repeated expansion. The nextLevelOpenList is used to store the new nodes after expansion. During initialization, the seed sequence is put into the OpenList with the initial cost and heuristic value as the first node to be expanded in the beam search. When the OpenList is not empty and the prediction of the entire sequence has not been completed, the algorithm will enter a loop. In each iteration, the algorithm selects several nodes with the lowest cost from the OpenList and expands these nodes. During the entire search process, the beam width (beamWidth) is a dynamically changing parameter that determines the number of candidate nodes retained in each level of search. For each node selected from the OpenList, the model will predict based on the sequence of the current node, that is, calculate the next possible base based on the sequence. The output of the model is a probability distribution, which indicates the probability of each possible base at that position. According to the prediction results of the model, the algorithm tries to expand new bases one by one and calculates the new cost newG and new heuristic value newH for each newly expanded node. Step 2.3: After several rounds of expansion phase, a large number of expansion nodes will be obtained. At this time, these nodes will be pruned and optimized, that is, entering the contraction phase. Here, attention should be paid to the Limit size operation of the node. The Limit size operation is affected by the beam width and needs to be adjusted dynamically. After that, the nodes that have passed the beam constraint are sent to the nextLevelOpenList, which is used to store the expanded nodes. Repeat the above operation until the storage of OpenList reaches the set value, and the predicted sequence is obtained. The predicted gene sequence is assembled together with the original genome to obtain the assembly result 2. Step 3 specifically includes the following steps: Step 3.1: Use the alignment software Quast to align the assembly files 1 and 2 obtained in steps 1 and 2 to the reference genome respectively to obtain assembly results 1 and 2. Step 3.2: Use Exonerate to customize the assembly evaluation index; Step 3.3: Since the quality of the prediction sequences generated in step 2 varies, we built a PredictionFilter mechanism for predictive sequence screening and optimization. Focusing on the alignment rate and other indicators of the new evaluation index, the prediction sequences corresponding to the assembly effect 2 are not as good as the assembly effect 1 are eliminated. In this way, a new prediction sequence can be obtained after screening, and then the optimized prediction sequence is assembled together with the original data.