Sequencing gene assembling method and device, electronic equipment and storage medium

Through error correction and joint assembly of Nanopore and PacbioHiFi data, combined with chromosome localization and filling of Hi-C data, the assembly problem of high-repetition regions in T2T sequencing technology is solved, the accuracy and continuity of genome splicing are improved, and a more complete genome map is provided.

CN120020963APending Publication Date: 2025-05-20BGI TECH SOLUTIONS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311554684.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

The existing T2T sequencing technology has the problem that the sequencing read length is short and cannot cross high repeat regions, the limitations of assembly algorithms lead to low splicing continuity, and the limitations of assembly strategy lead to the existence of vacant regions of the splicing genome, and the inability to correctly assemble high repeat regions such as telomeres and centromeres.

Method used

By correcting Nanopore data and assembling with PacbioHiFi data, combining Hi-C data for chromosome localization and assembly interruption region filling, preset indicators are used to evaluate the assembly effect to ensure accurate supplementation of telomere sequences.

Benefits of technology

It improves the continuity and length of the spliced genome, provides a more accurate genome map, and lays the foundation for subsequent research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020963A_ABST
    Figure CN120020963A_ABST
Patent Text Reader

Abstract

The invention discloses a sequencing gene assembling method and device, electronic equipment and a storage medium, and relates to the technical field of genome assemblation.The main technical scheme includes the steps that Nanopore data is subjected to error correction, assembling is conducted according to the Nanopore data subjected to error correction, and an error correction set and a first assembling set are obtained; performing joint assembly according to PacbioHiFi data and the error correction set to obtain a second assembly set; performing chromosome positioning on the second assembly set according to Hi-C data, and determining each assembly interruption area to obtain a third assembly set; and filling the assembly interruption area in the third assembly set according to the first assembly set and the error correction set to obtain an assembled chromosome genome. By combining various sequencing data, splicing the sequencing data and filling the assembly interruption region in the sequencing data, the continuity and length of the spliced genome are improved, the assembly effect is improved, and a more accurate genome map is provided for subsequent research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of gene assembly technology, and particularly to a method and device for assembling sequenced genes, an electronic device, and a storage medium. Background Art

[0002] With the progress of sequencing technology and the improvement of research requirements, it is currently possible to assemble the telomere-to-telomere T2T genome. Through the T2T gene, we can better understand the structure and function of the genome, and contribute to research in aspects such as species evolution, precise agricultural breeding and genetic improvement, and disease variant regions. The market prospect is very large.

[0003] The current T2T sequencing technology has the following disadvantages. First, the sequencing read length is short and cannot cross high-repeat regions. Second, the limitations of the assembly algorithm result in low continuity of splicing. Third, the limitations of the assembly strategy ultimately lead to gaps in the assembled genome, resulting in incorrect results for highly repetitive regions such as telomeres and centromeres in the assembled genome. Summary of the Invention

[0004] The present disclosure provides a method, device, electronic device, and storage medium for assembling sequenced genes. Its main purpose is to solve the problem that correct results cannot be assembled for highly repetitive regions such as telomeres and centromeres of the assembled genome.

[0005] According to the first aspect of the present disclosure, there is provided a method for assembling sequenced genes, which includes:

[0006] Correct the Nanopore data and perform assembly based on the corrected Nanopore data to obtain a corrected set and a first assembled set;

[0007] Perform joint assembly based on the PacbioHiFi data and the corrected set to obtain a second assembled set;

[0008] Perform chromosome localization on the second assembled set according to the Hi-C data to obtain an assembly interruption region, and obtain a third assembled set;

[0009] Fill the assembly interruption region in the third assembled set according to the first assembled set and the corrected set to obtain an assembled chromosomal genome.

[0010] Optionally, before correcting and assembling the Nanopore data to obtain a corrected set and a first assembled set, the method further includes:

[0011] Filter the Nanopore data, the Nanopore data, and the Hi-C data respectively according to preset filtering conditions, where different data correspond to different filtering conditions; the Nanopore data, the Nanopore data, and the Hi-C data are different sequencing data of the same research object.

[0012] Optionally, the error correction of the Nanopore data and the assembly according to the error-corrected Nanopore data to obtain an error correction set and a first assembly set include:

[0013] Perform local comparison on each data to be error-corrected in the Nanopore data to determine repeated sequences, and use the base with the highest repetition rate at the same position in the overlapping sequences as the base at that position to obtain an error correction set;

[0014] Perform genome assembly according to the error correction set to obtain a first assembly set.

[0015] Optionally, the chromosomal localization of the second assembly set according to the Hi-C data to obtain an assembly interruption region, and the steps to obtain a third assembly set further include:

[0016] Compare the Hi-C data with the second assembly set to obtain interaction information;

[0017] Classify, sort, and orient the overlapping fragments according to the interaction information to obtain the third assembly set;

[0018] Mark the assembly interruption regions in the third assembly set with preset bases and count the assembly interruption regions.

[0019] Optionally, the steps to fill the assembly interruption regions in the third assembly set according to the first assembly set and the error correction set to obtain an assembled chromosomal genome include:

[0020] Compare the first assembly set and the error correction set with the third assembly set to determine the gene sequence data corresponding to the assembly interruption regions in the third assembly set in the first assembly set and the error correction set;

[0021] Fill the assembly interruption regions according to the gene sequence data corresponding to the assembly interruption regions to obtain an assembled chromosomal genome.

[0022] Optionally, after filling the assembly interruption regions in the third assembly set according to the first assembly set and the error correction set to obtain an assembled chromosomal genome, the method further includes:

[0023] Determine whether there are telomere repeat units within the preset length regions at both ends of the assembled chromosomal genome;

[0024] If not, determine the corresponding telomere sequences in the first assembly set and the second assembly set, and supplement the determined telomere sequences to the corresponding positions of the assembled chromosomal genome.

[0025] Optionally, after determining the corresponding telomere sequences in the first assembly set and the second assembly set and supplementing the determined telomere sequences to the corresponding positions of the assembled chromosomal genome, the method further includes:

[0026] Evaluate the assembly effect of the assembled chromosomal genome according to preset indicators; wherein, the assembly effect includes at least one of assembly continuity, assembly integrity, and assembly accuracy; the preset indicators include at least one of target gene sequences, genome assembly size, number of assembly interruption regions, and chromosome degree.

[0027] According to a second aspect of the present disclosure, there is provided an assembly device for sequencing genes, including:

[0028] A first assembly unit for correcting Nanopore data and assembling according to the corrected Nanopore data to obtain a corrected set and a first assembly set;

[0029] A joint assembly unit for performing joint assembly according to PacbioHiFi data and the corrected set to obtain a second assembly set;

[0030] A second assembly unit for performing chromosome localization on the second assembly set according to Hi-C data to obtain assembly interruption regions and obtain a third assembly set;

[0031] A filling unit for filling the assembly interruption regions in the third assembly set according to the first assembly set and the corrected set to obtain an assembled chromosomal genome.

[0032] Optionally, the device further includes:

[0033] A filtering unit for filtering the Nanopore data, the Nanopore data, and the Hi-C data according to preset filtering conditions respectively before the first assembly unit corrects and assembles the Nanopore data to obtain a corrected set and a first assembly set, wherein different data correspond to different filtering conditions; the Nanopore data, the Nanopore data, and the Hi-C data are different sequencing data of the same research object.

[0034] Optionally, the first assembly unit is further configured to:

[0035] Perform local comparison on each piece of data to be error-corrected in the Nanopore data respectively, determine the repetitive sequences, and use the base with the highest repetition rate at the same position in the overlapping sequences as the base at that position to obtain an error-correction set;

[0036] Perform genome assembly according to the error-correction set to obtain a first assembly set.

[0037] Optionally, the second assembly unit is further configured to:

[0038] Compare the Hi-C data with the second assembly set to obtain interaction information;

[0039] Classify, sort, and orient the overlapping fragments according to the interaction information to obtain the third assembly set;

[0040] Mark the assembly interruption regions in the third assembly set with preset bases and count the assembly interruption regions.

[0041] Optionally, the filling unit is further configured to:

[0042] Compare the first assembly set and the error-correction set with the third assembly set to determine the gene sequence data corresponding to the assembly interruption regions in the third assembly set in the first assembly set and the error-correction set;

[0043] Fill the assembly interruption regions according to the gene sequence data corresponding to the assembly interruption regions to obtain an assembled chromosomal genome.

[0044] Optionally, the device further includes:

[0045] A determination unit, configured to determine whether there are telomere repeat units within a preset length region at both ends of the assembled chromosomal genome after the filling unit fills the assembly interruption regions in the third assembly set according to the first assembly set and the error-correction set to obtain the assembled chromosomal genome;

[0046] A supplement unit, configured to determine the corresponding telomere sequences in the first assembly set and the second assembly set when there are no telomere repeat units within a preset length region at both ends of the assembled chromosomal genome, and supplement the determined telomere sequences to the corresponding positions of the assembled chromosomal genome.

[0047] Optionally, the device further includes:

[0048] An evaluation unit, configured to evaluate the assembly effect of the assembled chromosomal genome according to a preset index after the supplementation unit determines corresponding telomere sequences in the first assembly set and the second assembly set and supplements the determined telomere sequences to corresponding positions of the assembled chromosomal genome; wherein, the assembly effect includes at least one of assembly continuity, assembly integrity, and assembly accuracy; the preset index includes at least one of a target gene sequence, a genome assembly size, a number of assembly interruption regions, and a chromosome degree.

[0049] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0050] At least one processor; and

[0051] A memory communicatively connected to the at least one processor; wherein,

[0052] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the foregoing first aspect.

[0053] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the foregoing first aspect.

[0054] According to a fifth aspect of the present disclosure, there is provided a computer program product, including a computer program which, when executed by a processor, implements the method described in the foregoing first aspect.

[0055] The main technical solutions of the assembly method, device, electronic device, and storage medium for sequencing genes provided by the present disclosure include: correcting Nanopore data and performing assembly according to the corrected Nanopore data to obtain a correction set and a first assembly set; performing joint assembly on PacbioHiFi data and the correction set to obtain a second assembly set; performing chromosome localization on the second assembly set according to Hi-C data to determine each assembly interruption region and obtain a third assembly set; filling the assembly interruption regions in the third assembly set according to the first assembly set and the correction set to obtain an assembled chromosomal genome. Compared with the related art, the embodiments of the present application improve the continuity and length of the spliced genome and the assembly effect by combining multiple sequencing data, splicing the sequencing data, and filling the assembly interruption regions in the sequencing data, providing a more accurate genome map for subsequent research.

[0056] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The drawings are used to better understand the present solution and do not constitute a limitation to the present disclosure. Among them:

[0058] Figure 1 is a schematic flowchart of a method for assembling a sequencing gene provided by an embodiment of the present disclosure;

[0059] Figure 2 is a schematic flowchart of another method for assembling a sequencing gene provided by an embodiment of the present disclosure;

[0060] Figure 3 is a schematic flowchart of another method for assembling a sequencing gene provided by an embodiment of the present disclosure;

[0061] Figure 4 is a schematic structural diagram of an apparatus for assembling a sequencing gene provided by an embodiment of the present disclosure;

[0062] Figure 5 is a schematic structural diagram of another apparatus for assembling a sequencing gene provided by an embodiment of the present disclosure;

[0063] Figure 6 is a schematic block diagram of an exemplary electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0065] The following describes a method, an apparatus, an electronic device, and a storage medium for assembling a sequencing gene according to embodiments of the present disclosure with reference to the drawings.

[0066] Figure 1 is a schematic flowchart of a method for assembling a sequencing gene provided by an embodiment of the present disclosure.

[0067] As Figure 1 shown, the method includes the following steps:

[0068] Step 101: Correct the Nanopore data and perform assembly based on the corrected Nanopore data to obtain a correction set and a first assembly set.

[0069] Assembly means splicing short read sequences or long read sequences through alignment and stitching them together into a complete and ordered sequence according to the alignment relationship.

[0070] In an implementable manner of the embodiments of the present application, the Nanopore data is the whole-genome long fragment DNA sequence obtained by sequencing with the Oxford Nanopore Technologies (ONT) platform. The Nanopore data can be used for genome sequencing, transcriptome sequencing, and epigenetic research.

[0071] In an implementable manner of the embodiments of the present application, possible sequencing errors are identified by aligning the Nanopore data with a reference sequence. Some common alignment tools can be used, such as BLAST, Bowtie, BWA, etc. Specifically, the embodiments of the present application do not limit this.

[0072] After correcting the Nanopore data, the corrected data can be used for assembly. Assembly is the process of stitching fragmented sequencing reads into longer continuous sequences to restore the original DNA sequence.

[0073] There are various assembly methods, and common ones include: Overlap graph assembly: stitching sequencing reads into longer continuous sequences according to the overlap relationship between them; De Bruijn graph assembly: splitting sequencing reads into fixed-length k-mers, constructing a De Bruijn graph, and then stitching the sequencing reads according to the paths in the graph; Hybrid assembly: combining different assembly methods and taking advantage of their advantages for assembly. For example, some longer sequences can be obtained first using overlap graph assembly, and then De Bruijn graph assembly can be used to fill the gaps. Specifically, the embodiments of the present application do not limit the assembly method.

[0074] Step 102: Perform joint assembly based on the Pacbio HiFi data and the correction set to obtain a second assembly set.

[0075] Integrate the original Nanopore data and the correction set into a data set. This can be achieved by merging the sequencing reads of the two into one file. For the specific implementation method, please refer to any implementation method in the prior art. Specifically, the embodiments of the present application will not elaborate on this one by one here.

[0076] In an implementable manner of the embodiment of the present application, sequencing errors in the overlapping graph can be corrected according to the information in the error correction set. This can be achieved by comparing the sequences in the error correction set with the sequencing reads in the overlapping graph. If there is a mismatch between the sequence in the error correction set and a certain sequencing read at a certain position, then the base at the mismatch position can be corrected, and according to the corrected overlapping graph, one or more assembly paths are selected. According to the selected assembly paths, the corresponding sequencing reads are spliced into longer continuous sequences; an assembly path refers to one or more paths from the start node to the end node, representing the continuous sequence obtained by assembly. The principle for selecting the assembly path can be coverage, mismatch rate, overlap length, etc.

[0077] The PacBio platform is a long-read sequencing technology developed by Pacific Bioscience. Currently, the commercially available sequencers are PacBio RSII and Sequel, and Sequel is mainly used for sequencing animal and plant genomes. The sequencing principle of the PacBio platform is single-molecule real-time sequencing. There are many circular nano-pores, namely ZMWs (Zero-Mode Waveguides), in a chip, which is a reaction tube (SMRT Cell: Single-Molecule Real-Time Reaction Tube). The outer diameter is more than 100 nanometers, which is smaller than the detection laser wavelength (hundreds of nanometers). After the laser hits from the bottom, it cannot penetrate the pores into the upper solution area, and the energy is limited to a small range, just enough to cover the part to be detected, so that the signal only comes from this small reaction area, and the excessive free nucleotide monomers outside the pores still remain in the dark, minimizing the background. There is a polymerase bound to the template DNA at the bottom of a single ZMW. This DNA polymerase is one of the keys to achieving ultra-long read lengths. The read length is mainly related to the activity maintenance of the enzyme. Mainly, the laser will cause damage to it. When the sequencing reaction reagent is added, different bases are added after Watson pairing, and 4-color fluorescence labels 4 bases, which will emit different lights. According to the wavelength and peak value of the light, the type of the entering base can be judged. There are 150,000 ZMWs in a PacBio RSII SMRT Cell (there are 1 million ZMWs in a Sequel instrument's SMRT Cell). There is a single-molecule DNA strand in each pore synthesizing at high speed, like twinkling stars. As a result of the original detection data, each synthesized base is displayed as a pulse peak. At a speed of more than 100 bases per minute, with a high-resolution optical detection system, real-time detection can be carried out.

[0078] Step 103: Perform chromosome localization on the second assembly set according to the Hi-C data to obtain the assembly interruption region, and obtain the third assembly set.

[0079] Chromosome localization: Using the Hi-C contact matrix, the sequences in the second assembly set can be localized to chromosomes; the Hi-C contact matrix reflects the interaction frequencies between different chromosomal regions; this can be achieved by aligning the sequences in the second assembly set with the contact strengths in the Hi-C contact matrix. According to the levels of contact strengths, the chromosomal regions where the sequences in the second assembly set are located can be determined; according to the Hi-C data, the assembly interruption regions can be identified. The assembly interruption regions refer to the regions where the assembly results are incomplete due to sequencing errors, repetitive sequences, etc. during the assembly process. By analyzing the contact strengths in the Hi-C contact matrix, the assembly interruption regions can be identified, and then the third assembly set can be generated.

[0080] Step 104, fill the assembly interruption regions in the third assembly set according to the first assembly set and the error correction set to obtain the assembled chromosomal genome.

[0081] Based on information such as the third assembly set and the Hi-C data, identify the assembly interruption regions, and use the sequence information in the first assembly set and the error correction set to fill the assembly interruption regions; when performing this step, the sequences in the first assembly set and the error correction set can be aligned with the assembly interruption regions in the third assembly set for splicing to fill in the missing sequences.

[0082] In an implementable manner of the embodiment of the present application, during the process of filling the assembly interruption regions, the information in the error correction set can also be used to correct possible sequencing errors. Error correction can be performed by aligning the sequences in the error correction set with the assembly interruption regions in the third assembly set, thereby improving the accuracy of the filled sequences; according to the filled assembly results, the assembled chromosomal genome can be generated.

[0083] The main technical solutions of the sequencing gene assembly method provided by the present disclosure include: performing error correction on Nanopore data and assembling according to the corrected Nanopore data to obtain an error correction set and a first assembly set; performing joint assembly according to PacbioHiFi data and the error correction set to obtain a second assembly set; performing chromosome localization on the second assembly set according to Hi-C data to determine each assembly interruption region and obtain a third assembly set; filling the assembly interruption regions in the third assembly set according to the first assembly set and the error correction set to obtain the assembled chromosomal genome. Compared with the related technologies, in the embodiment of the present application, by combining multiple sequencing data, splicing the sequencing data, and filling the assembly interruption regions in the sequencing data, the continuity and length of the spliced genome are improved, the assembly effect is enhanced, and a more accurate genome map is provided for subsequent research.

[0084] In an implementable manner of the embodiments of the present application, in order to reduce the computational amount and improve the quality of the assembly result, before performing step 101 to correct and assemble the Nanopore data to obtain a corrected set and a first assembly set, it is necessary to clean and filter each data to be assembled according to the corresponding rules, which can be carried out according to the following steps:

[0085] Filter the Nanopore data, the Nanopore data, and the Hi-C data respectively according to preset filtering conditions, where different data correspond to different filtering conditions; the Nanopore data, the Nanopore data, and the Hi-C data are different sequencing data of the same research object.

[0086] In an implementable manner of the embodiments of the present application. When performing step 101 to correct the Nanopore data and assemble according to the corrected Nanopore data to obtain a corrected set and a first assembly set, the following steps can be specifically referred to for assembly:

[0087] Perform local comparison on each data to be corrected in the Nanopore data to determine the repetitive sequences, and use the base with the highest repetition rate at the same position in the overlapping sequences as the base at this position to obtain a corrected set;

[0088] Perform genome assembly according to the corrected set to obtain a first assembly set.

[0089] For the HiFi library, sequence it with the Pacific Biosciences (PacBio) HiFi platform to obtain the whole-genome long-fragment DNA sequence, which is called the PacBio off-machine data here. Use open-source software to perform self-correction and filtering on the PacBio off-machine data, filter out reads with a sequencing accuracy lower than 99% or reads with a read length less than 500 bp, and finally obtain high-quality PacBio HiFi data for subsequent assembly of Contig.

[0090] For the Nanopore Ultra-long library, sequence it through the Oxford Nanopore Technologies (ONT) platform to obtain the whole-genome long-fragment DNA sequence, which is called the ONT off-machine data here. Use a self-written Perl program to filter the ONT off-machine data, filter out reads with a sequencing quality value less than 7 or reads with a read length less than 100 Kb, and obtain ONT Ultra-long data for subsequent assembly of Contig and Scaffold gap filling.

[0091] The Hi-C library is sequenced through platforms such as DNB-seq to obtain DNA sequences containing the interaction relationships between genomic DNA fragments, which are herein referred to as Hi-C raw data. The open-source SoapNuke software is used to process the Hi-C raw data, and reads that meet any of the following conditions are filtered out: A. The proportion of bases with a quality value less than 20 exceeds 50% of the total number of bases in the entire read; B. Reads containing adapters; C. Reads with an N ratio greater than 1% in each read. The processed data is herein referred to as Hi-C data, which is used to cluster, arrange, and orient the Contig / Scaffold sequences to obtain Scaffolds close to the chromosome length.

[0092] The DNA small fragment library is sequenced through the DNB-seq platform to obtain short read sequences of genomic DNA, which are herein referred to as NGS raw data, etc. The open-source SoapNuke software is used to process the Hi-C raw data, and reads that meet any of the following conditions are filtered out: A. The proportion of bases with a quality value less than 20 exceeds 50% of the total number of bases in the entire read; B. Reads containing adapters; C. Reads with an N ratio greater than 1% in each read. The processed data is herein referred to as NGS data, which is used to evaluate the quality of genome assembly.

[0093] Please refer to Figure 2 , Figure 2 For the flow schematic diagram of a method for assembling sequencing genes provided by an embodiment of the present disclosure to perform chromosome localization, obtain the assembly interruption region, and obtain the third assembly set, it further includes:

[0094] Step 201, compare the Hi-C data with the second assembly set to obtain interaction information.

[0095] Align the Hi-C data with the second assembly set. Various alignment algorithms can be used for alignment, such as BLAST, Bowtie, BWA, etc.; the purpose of alignment is to match the reads (fragments) in the Hi-C data with the sequences in the second assembly set to find their corresponding relationships; according to the alignment results, the interaction relationships between the reads in the Hi-C data and the sequences in the second assembly set can be determined. The interaction relationship refers to the physical contact between two genomic regions, which can be determined by the matching positions of the reads in the Hi-C data and the sequences in the second assembly set. It should be noted that this description method is only an exemplary illustration and is not a specific limitation on the specific alignment operation. The embodiments of the present application do not limit this.

[0096] Step 202, classify, sort, and orient the overlapping fragments according to the interaction information to obtain the third assembly set.

[0097] According to the interaction information, the overlapping fragments are classified into different categories; this can be classified according to factors such as their positions in the genome, the frequency of interaction, the strength of interaction, etc.; for example, fragments with similar interaction patterns can be grouped into one category; specifically, the embodiments of the present application do not limit this.

[0098] For each category of overlapping fragments, they can be sorted according to their positions in the genome for subsequent assembly processes, which can be achieved by comparing their starting positions or central positions.

[0099] For each category of overlapping fragments, they can be oriented according to their relative directions. According to the interaction information, the relative directions between the fragments can be determined, that is, whether they are head-to-head or head-to-tail connections.

[0100] Step 203, mark the assembly interruption regions in the third assembly set with preset bases and count the assembly interruption regions.

[0101] Mark the assembly interruption regions: According to the preset base sequence, the assembly interruption regions in the third assembly set can be marked. This can be achieved by aligning the preset base sequence with the assembly interruption regions. If the assembly interruption region matches the preset base sequence, it can be marked as an assembly interruption region.

[0102] Count the assembly interruption regions: For the marked assembly interruption regions, statistical analysis can be performed. Statistical metrics such as the number, length, and distribution of the assembly interruption regions can be calculated. This can help understand the situation of assembly interruption, such as the frequency and position preference of assembly interruption.

[0103] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of an assembly method for sequencing genes provided by an embodiment of the present disclosure, including:

[0104] Step 301, compare the first assembly set and the error correction set with the third assembly set to determine the gene sequence data corresponding to the assembly interruption regions in the third assembly set in the first assembly set and the error correction set.

[0105] Align the third assembly set and the first assembly set: Use a suitable alignment tool (such as BLAST, Bowtie, etc.) to align the sequences in the third assembly set with the sequences in the first assembly set. Through alignment, the corresponding positions of the sequences in the third assembly set in the first assembly set can be found. Similarly, align the sequences in the third assembly set with the sequences in the error correction set. Through alignment, the corresponding positions of the sequences in the third assembly set in the error correction set can be found.

[0106] According to the above comparison results, the corresponding positions of the assembly interruption regions in the third assembly set in the first assembly set and the error correction set can be determined, and the gene sequence data of the assembly interruption regions in the first assembly set and the error correction set can be found, which helps to understand the characteristics, functions of the assembly interruption regions and their relationships with gene sequences.

[0107] Step 302, fill the assembly interruption region according to the gene sequence data corresponding to the assembly interruption region to obtain an assembled chromosomal genome.

[0108] According to the previous comparison results, the gene sequence data corresponding to the assembly interruption region can be obtained from the first assembly set and the error correction set, and the gene sequence data corresponding to the assembly interruption region is compared with the corresponding region in the third assembly set. According to the comparison results, the assembly interruption region can be filled; the filling method can be determined according to the comparison results. For example, methods such as local alignment and sequence splicing can be used. Specifically, the embodiments of the present application do not limit this.

[0109] Repeat the filling step until all the assembly interruption regions are filled; by filling the assembly interruption regions, an assembled chromosomal genome can be obtained. This genome can be used for further genome analysis, annotation and research.

[0110] In an implementable manner of the embodiments of the present application, after obtaining the assembled chromosomal genome in step 103, it is also necessary to detect both ends of the chromosomal genome to determine whether there are repeat units and supplement the repeat units; the specific steps include:

[0111] Determine whether there are telomere repeat units within a preset length region at both ends of the assembled chromosomal genome;

[0112] If not, determine the corresponding telomere sequences in the first assembly set and the second assembly set, and supplement the determined telomere sequences to the corresponding positions of the assembled chromosomal genome.

[0113] According to the characteristics of telomere repeat units, the repeat units of animals are mainly TTAGGG, and the repeat units of plants are mainly TTTAGGG. Open-source software such as Tidk is used for detection and statistics to detect the frequency of telomere repeat units appearing in a region of about 10 kb at both ends of each chromosome. If not, it is judged as telomere deletion. Finally, publicly available software such as Winnowmap is used to search for telomere sequences in ONT and Pacbio sequencing data and fill them into the corresponding telomere deletion regions to obtain a T2T genome assembled from telomere to telomere.

[0114] Evaluate the assembly effect of the assembled chromosomal genome according to preset indicators; wherein, the assembly effect includes at least one of assembly continuity, assembly integrity, and assembly accuracy; the preset indicators include at least one of target gene sequences, genome assembly size, number of assembly interruption regions, and chromosome degree.

[0115] Assembly continuity evaluation: Through a self-written Perl program, statistics are made on indicators such as Contig N50, Contig N90, Scaffold N50, Scaffold N90, genome assembly size, number of GAPs, and length of each chromosome of the T2T genome sequence.

[0116] Assembly integrity evaluation: Benchmarking Universal Single-Copy Orthologs (BUSCO) is a software for evaluating the quality of genome assembly. It evaluates the integrity and redundancy of the assembled genome based on single-copy orthologous genes (orthologs) between species. First, a single-copy homologous gene set is constructed according to the corresponding database, and the splicing result is compared with this gene set. According to the alignment ratio, integrity, and accuracy, the accuracy and integrity of the splicing result are evaluated.

[0117] Assembly accuracy evaluation: Through the open-source software Merqury, NGS data or HiFi data is aligned to the third assembly set to evaluate the accuracy of the assembly result.

[0118] It should be noted that the embodiments of the present disclosure may include multiple steps. For the convenience of description, these steps are numbered, but these labels are not intended to limit the execution time slots and execution orders between the steps; these steps can be implemented in any order, and the embodiments of the present disclosure do not make any limitations in this regard.

[0119] Corresponding to the above-mentioned assembly method of sequencing genes, the present invention also provides an assembly device for sequencing genes. Since the device embodiments of the present invention correspond to the above-mentioned method embodiments, details not disclosed in the device embodiments can be referred to the above-mentioned method embodiments, and will not be elaborated in the present invention.

[0120] Figure 4 This is a schematic structural diagram of an assembly device for sequencing genes provided by an embodiment of the present disclosure, as Figure 4 shown, including:

[0121] The first assembly unit 41 is used to correct the Nanopore data and perform assembly according to the corrected Nanopore data to obtain a corrected set and a first assembly set;

[0122] A joint assembly unit 42 for jointly assembling according to PacbioHiFi data and the error correction set to obtain a second assembly set;

[0123] A second assembly unit 43 for performing chromosome localization on the second assembly set according to Hi-C data to obtain assembly interrupted regions and obtain a third assembly set;

[0124] A filling unit 44 for filling the assembly interrupted regions in the third assembly set according to the first assembly set and the error correction set to obtain an assembled chromosomal genome.

[0125] The assembly device for sequencing genes provided by the present disclosure mainly includes the following technical solutions: correcting Nanopore data and assembling according to the corrected Nanopore data to obtain an error correction set and a first assembly set; jointly assembling according to PacbioHiFi data and the error correction set to obtain a second assembly set; performing chromosome localization on the second assembly set according to Hi-C data to determine each assembly interrupted region and obtain a third assembly set;

[0126] Filling the assembly interrupted regions in the third assembly set according to the first assembly set and the error correction set to obtain an assembled chromosomal genome. Compared with the related art, the embodiment of the present application combines multiple sequencing data, splices the sequencing data, and fills the assembly interrupted regions in the sequencing data, improving the continuity and length of the spliced genome and the assembly effect, and providing a more accurate genomic map for subsequent research.

[0127] Further, in a possible implementation manner of this embodiment, as Figure 5 shown, the device further includes:

[0128] A filtering unit 45 for filtering the Nanopore data, the Nanopore data and the Hi-C data respectively according to preset filtering conditions before the first assembly unit 41 corrects and assembles the Nanopore data to obtain an error correction set and a first assembly set, where different data correspond to different filtering conditions; the Nanopore data, the Nanopore data and the Hi-C data are different sequencing data of the same research object.

[0129] Further, in a possible implementation manner of this embodiment, the first assembly unit 41 is further configured to:

[0130] Perform local comparison on each piece of data to be error-corrected in the Nanopore data to determine repetitive sequences, and use the base with the highest repetition rate at the same position in the overlapping sequences as the base at that position to obtain an error correction set;

[0131] Perform genome assembly according to the error correction set to obtain a first assembly set.

[0132] Further, in a possible implementation manner of this embodiment, the second assembly unit 43 is further configured to:

[0133] Compare the Hi-C data with the second assembly set to obtain interaction information;

[0134] Classify, sort, and orient the overlapping fragments according to the interaction information to obtain the third assembly set;

[0135] Mark the assembly interruption regions in the third assembly set with preset bases and count the assembly interruption regions.

[0136] Further, in a possible implementation manner of this embodiment, the filling unit 44 is further configured to:

[0137] Compare the first assembly set and the error correction set with the third assembly set to determine the gene sequence data corresponding to the assembly interruption regions in the third assembly set in the first assembly set and the error correction set;

[0138] Fill the assembly interruption regions according to the gene sequence data corresponding to the assembly interruption regions to obtain an assembled chromosomal genome.

[0139] Further, in a possible implementation manner of this embodiment, as Figure 5 shown, the device further includes:

[0140] A determination unit 45, configured to determine whether there are telomere repeat units in a preset length region at both ends of the assembled chromosomal genome after the filling unit 44 fills the assembly interruption regions in the third assembly set according to the first assembly set and the error correction set to obtain the assembled chromosomal genome;

[0141] A supplement unit 46, configured to determine the corresponding telomere sequences in the first assembly set and the second assembly set and supplement the determined telomere sequences to the corresponding positions of the assembled chromosomal genome when there are no telomere repeat units in the preset length regions at both ends of the assembled chromosomal genome.

[0142] Further, in a possible implementation manner of this embodiment, as Figure 5 shown, the device further includes:

[0143] An evaluation unit 47, configured to evaluate the assembly effect of the assembled chromosomal genome according to a preset index after the supplementary unit 46 determines corresponding telomere sequences in the first assembly set and the second assembly set and supplements the determined telomere sequences to corresponding positions of the assembled chromosomal genome; wherein, the assembly effect includes at least one of assembly continuity, assembly integrity, and assembly accuracy; the preset index includes at least one of a target gene sequence, a genome assembly size, a number of assembly interruption regions, and a chromosome degree.

[0144] It should be noted that the foregoing explanation of the method embodiments also applies to the device of this embodiment, with the same principle, and will not be limited in this embodiment.

[0145] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0146] Figure 6 FIG. shows a schematic block diagram of an exemplary electronic device 500 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0147] As Figure 6 shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 502 or a computer program loaded from a storage unit 508 into a RAM (Random Access Memory) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An I / O (Input / Output) interface 505 is also connected to the bus 504.

[0148] Multiple components in device 500 are connected to I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disc, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0149] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the method for assembling a sequenced gene. For example, in some embodiments, the method for assembling a sequenced gene can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute the aforementioned method for assembling a sequenced gene in any other suitable manner (e.g., by means of firmware).

[0150] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SoCs (System On Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0151] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0152] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0153] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0154] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a LAN (Local Area Network), a WAN (Wide Area Network), the Internet, and a blockchain network.

[0155] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS" for short). The server may also be a server of a distributed system or a server combined with a blockchain.

[0156] It should be noted that artificial intelligence is a discipline that studies to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and there are both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0157] The various digital numbers such as the first and the second involved in this disclosure are only for the convenience of description and are not used to limit the scope of the embodiments of this disclosure, nor do they represent the order of precedence.

[0158] At least one in this disclosure may also be described as one or more. The plurality may be two, three, four, or more, and this disclosure does not make any restrictions. In the embodiments of this disclosure, for a technical feature, the technical features in this technical feature are distinguished by "the first", "the second", "the third", "A", "B", "C", and "D", etc. There is no order of precedence or size order among the technical features described by "the first", "the second", "the third", "A", "B", "C", and "D".

[0159] It should be understood that various forms of processes shown above can be used, reordering, adding, or deleting steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0160] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for assembling sequenced genes, characterized in that: include: Correcting the Nanopore data and assembling the Nanopore data after the error correction to obtain an error correction set and a first assembly set; Combined assembly is performed according to PacbioHiFi data and the error correction set to obtain a second assembly set; Perform chromosome localization on the second assembly set according to Hi-C data, determine each assembly interruption region, and obtain a third assembly set; The assembly interruption region in the third assembly set is filled according to the first assembly set and the error correction set to obtain an assembled chromosome genome.

2. The method according to claim 1, characterized in that Before error correction and assembly of the Nanopore data to obtain an error correction set and a first assembly set, the method further comprises: The Nanopore data, the Nanopore data and the Hi-C data are filtered respectively according to preset filtering conditions, wherein different data correspond to different filtering conditions; the Nanopore data, the Nanopore data and the Hi-C data are different sequencing data of the same research object.

3. The method according to claim 1, characterized in that The step of correcting the Nanopore data and assembling the Nanopore data after the error correction to obtain the error correction set and the first assembly set comprises: Performing local comparison on each piece of data to be corrected in the Nanopore data, respectively, to determine the repeated sequence, and using the base with the highest repetition rate at the same position in the overlapping sequence as the base at the position, to obtain an error correction set; The genome is assembled according to the error correction set to obtain a first assembly set.

4. The method according to claim 1, characterized in that: The performing chromosome localization on the second assembly set according to the Hi-C data to obtain the assembly interruption region to obtain the third assembly set further comprises: comparing the Hi-C data with the second assembly to obtain interaction information; Classifying, sorting, and orienting the overlapping fragments according to the interaction information to obtain the third assembly set; The assembly interruption region in the third assembly set is marked with a preset base, and statistics are performed on the assembly interruption region.

5. The method according to any one of claims 1 to 4, characterized in that The filling of the assembly interruption region in the third assembly set according to the first assembly set and the error correction set to obtain an assembled chromosome genome comprises: Determine the gene sequence data corresponding to the assembly interruption region in the third assembly set in the first assembly set and the error correction set by comparing the first assembly set and the error correction set with the third assembly set; According to the gene sequence data corresponding to the assembly interruption region, the assembly interruption region is filled to obtain an assembled chromosome genome.

6. The method according to claim 1, characterized in that After filling the assembly interruption region in the third assembly set according to the first assembly set and the error correction set to obtain an assembled chromosome genome, the method further includes: Determining whether telomere repeat units exist in the regions of preset length at both ends of the assembled chromosome genome; If not present, the corresponding telomere sequence is determined in the first assembly set and the second assembly set, and the determined telomere sequence is added to the corresponding position of the assembled chromosome genome.

7. The method according to claim 6, characterized in that After determining the corresponding telomere sequence in the first assembly set and the second assembly set, and adding the determined telomere sequence to the corresponding position of the assembled chromosome genome, the method further includes: The assembly effect of the assembled chromosome genome is evaluated according to preset indicators; wherein the assembly effect includes at least one of assembly continuity, assembly completeness and assembly accuracy; the preset indicators include at least one of target gene sequence, genome assembly size, number of assembly interruption regions and chromosome degree.

8. An assembly device for sequencing genes, characterized in that: include: A first assembly unit, used for correcting the Nanopore data and assembling according to the error-corrected Nanopore data to obtain an error-corrected set and a first assembly set; A joint assembly unit, used for performing joint assembly according to the PacbioHiFi data and the error correction set to obtain a second assembly set; A second assembly unit is used to perform chromosome localization on the second assembly set according to Hi-C data, obtain an assembly interruption region, and obtain a third assembly set; A filling unit is used to fill the assembly interruption region in the third assembly set according to the first assembly set and the error correction set to obtain an assembled chromosome genome.

9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.

11. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 7.