A gene assisted assembly device, chromosome level genome and application
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-26
- Publication Date
- 2026-08-11
AI Technical Summary
但是纠错单元可能会改变原始草图序列,纠错反而存在失误的情况
[0128](1)采用本发明的装置,能够将Hi-C辅助组装的挂载率提高至95%以上。
Smart Images

Figure CN116864008B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a gene-assisted assembly device, the assembled chromosome-level genome, and its applications. Background Technology
[0002] Hi-C technology (High-throughput / resolution Chromosome Conformation Capture) originates from Chromosome Conformation Capture-3C technology. It uses the entire cell nucleus as the research object, employing high-throughput sequencing technology combined with bioinformatics methods to study the interactions of chromosomal DNA across the entire genome. Hi-C-assisted assembly captures these interactions, and based on the principle that intrachromosomal interactions are significantly more frequent than interchromosomal interactions, and that interaction frequencies on the same chromosome decrease with increasing interaction distance, scaffolds or contigs are clustered into groups. Further, contigs / scaffolds within these groups are sorted and oriented to achieve genome mounting approaching the chromosome level.
[0003] Currently, the general steps of a Hi-C-assisted assembly device are as follows: 1. Data statistics and filtering unit; 2. Error correction unit for correcting errors in the reference genome dataset (error correction unit for reference genome drafts based on second-generation sequencing or third-generation sequencing); 3. Hi-C-assisted assembly unit; 4. Post-assembly evaluation unit. However, the error correction unit may alter the original draft sequence, and errors may occur during the correction process itself. Furthermore, the mounting rate of existing Hi-C-assisted assembly devices needs improvement, requiring more mature workflows to support it. Summary of the Invention
[0004] To address the aforementioned problems, this invention provides a gene-assisted assembly device, which can increase the loading rate of Hi-C assisted assembly to over 95%.
[0005] Specifically, the present invention relates to the following apparatus.
[0006] 1. A gene-assisted assembly device, comprising: a Hi-C library construction and sequencing unit, an alignment and selection unit, a preliminary assembly unit, and a processing and screening unit, wherein,
[0007] The Hi-C library construction and sequencing unit is used to construct and sequence a first dataset from a DNA sample using Hi-C library construction and sequencing.
[0008] The comparison and selection unit is used to compare and select the first dataset with the reference genome dataset to obtain the second dataset;
[0009] The preliminary assembly unit is used to perform preliminary assembly of the second dataset to obtain a preliminary assembled dataset;
[0010] The processing and filtering unit is used to process and filter the preliminary assembled dataset.
[0011] 2. According to the apparatus described above, the reference genome dataset is selected from a reference genome draft based on second-generation sequencing or a reference genome draft based on third-generation sequencing; and / or, the dataset includes multiple reads.
[0012] 3. According to the above-described apparatus, the comparison and selection unit comprises a first component, a second component, and a third component, wherein,
[0013] The first component is used to perform a first alignment of each data in the first dataset with a reference genome dataset to obtain a first subset that can be aligned to the reference genome dataset and a second subset that cannot be aligned to the reference genome dataset;
[0014] The second component is used to search for the link sites of the library after each data in the second subset is digested with enzymes, to break the data in the second subset at the link sites of the digested library found, and to perform a second alignment with the reference genome dataset to obtain a third subset that can be aligned to the reference genome dataset.
[0015] The third component is used to merge and select the first subset and the third subset to obtain the second data.
[0016] 4. According to the above-described apparatus, the selection is used to select data whose two ends are both aligned to a unique position in the reference genome dataset.
[0017] 5. According to the apparatus described above, wherein the preliminary assembly unit is based on LACHESIS software, and / or the preliminary assembly dataset includes preliminary assembled Contig / Scaffolds.
[0018] 6. The apparatus described above, wherein the processing and filtering unit is based on LACHESIS software;
[0019] Preferably, the processing and screening unit includes a fourth component and a fifth component, wherein,
[0020] The fourth component is used to allocate the preliminary assembly data to the chromosome group;
[0021] The fifth component is used to sort, orient, and filter the preliminary assembly data assigned to each chromosome group;
[0022] Preferably, the allocation is based on a clustering method;
[0023] Preferably, the fifth component includes a sorting component, a direction component, and a filtering component;
[0024] Preferably, the sorting component constructs an acyclic spanning tree based on the interaction relationships of contigs / scaffolds within a chromosome group, and selects the root tree with the highest confidence from it;
[0025] Preferably, the orientation component is used to traverse the possible directions of Contig / Scaffold through a weighted directed acyclic graph. For the order determination of each Contig / Scaffold, a scoring function is constructed based on the difference between positive and negative interaction relationships and used as the basis for confidence determination, and the orientation with the highest confidence is selected.
[0026] Preferably, the filtering parameters set by the filtering component are selected from one or more of the following: number of clusters, proportion of Contig / Scaffold used during orientation, proportion of Contig / Scaffold in the root tree, and ratio of the shortest chromosome length to the longest chromosome length.
[0027] 7. According to the above-described apparatus, the processing and filtering unit further includes a re-filtering component, used to, when the optimal preliminary assembly data selected cannot be output as assembly data, change the conditions for preliminary assembly again to obtain multiple preliminary assembly data in the preliminary assembly data obtained by preliminary assembly of the second data, and reprocess and filter the preliminary assembly data to select the optimal preliminary assembly data as assembly data.
[0028] 8. According to the above apparatus, wherein the conditions are selected from one or more of the following: the minimum number of restriction enzyme sites in the Contig / Scaffold sequence in cluster analysis, the maximum Link depth in the Contig / Scaffold sequence in cluster analysis, the ratio of the number of interactions between the Contig / Scaffold sequence and the target cluster to the number of interactions between other clusters in cluster back-interpolation, and the minimum number of restriction enzyme sites in the Contig / Scaffold sequence within the root tree in ordination analysis.
[0029] 9. The chromosome-level genome assembled using the gene-assisted assembly device described above.
[0030] 10. The above-mentioned gene-assisted assembly device and the chromosome-level genome obtained by the above-mentioned assembly are used in chromosome interactions, whole genome resequencing, chromosome evolution and differentiation, and epigenomics. Detailed Implementation
[0031] To enable those skilled in the art to better understand the present invention, the technical solutions involved in the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the specific embodiments described are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0033] To facilitate understanding of this plan, some technical terms appearing in this document are explained.
[0034] Terminology Explanation:
[0035] Gene interaction: Gene interaction refers to the phenomenon where non-allelic genes influence the expression of the same trait through mutual interaction. Broadly speaking, gene interaction can be divided into intragene interaction and intergene interaction. Intragene interaction refers to the dominant-recessive effect between alleles. Intergene interaction refers to the interaction between non-allelic genes at different loci, manifested as complementarity, repression, epistasis, etc.
[0036] Positive interaction: When the directions of the interacting regions are consistent, it is a positive interaction.
[0037] Reverse interaction: When the directions of the interacting regions are opposite, it is a reverse interaction.
[0038] Hi-C technology: a chromatin conformation capture technology based on 3C technology and utilizing high-throughput sequencing. Hi-C technology is a novel approach that combines chromatin conformation capture with high-throughput sequencing. This technology cross-links and enriches DNA fragments with large linear distances and similar spatial structures, followed by pair-end sequencing. Analysis of the sequencing data reveals the interactions between different DNA segments in chromatin, thereby deriving the three-dimensional spatial structure of the genome and possible regulatory relationships between genes.
[0039] With the continuous development of NGS, obtaining a draft genome of a species through sequencing is now relatively easy. However, obtaining a complete genome from this draft is not simply a matter of increasing sequencing throughput. To determine the chromosomes corresponding to each contig / scaffold in the draft and their respective sequences on the genome, the classic solution is to construct long-fragment mate pair libraries to determine the connection order of contigs / scaffolds, allowing them to be linked together and extended to the chromosome assembly level. Due to numerous limitations of NGS, such as unavoidable issues like GC content, sequencing read length, and mapping accuracy, assembling genomes with high repetitive sequences and high heterozygosity using NGS sequencing data to the chromosome level is extremely difficult, especially for the genome assembly of large plants and animals. Third-generation sequencing, with its long read advantage, occupies an important position in the field of genome assembly and has overcome many genome assembly challenges for various species; however, its high cost also restricts its widespread application.
[0040] Chromosome mounting rate: The proportion of contigs / scaffolds assembled into chromosomes to the total number of contigs / scaffolds in the sketch.
[0041] Reference genome draft: A genome sequence map with over 90% sequenced and sequencing accuracy at 1%. The reference genome draft contains gaps. The completed genome map performs gap closure on the entire genome to obtain the completed genome map. The completed map is assembled from the draft map using techniques such as gap filling to create sequences that meet evaluation parameters such as N50.
[0042] Raw Reads: The raw image data obtained from sequencing is converted into sequence data through base calling, which is called Raw Data or Raw Reads. The results are stored in fastq file format. The fastq file is the most raw file obtained by the user, which stores the sequence of the reads and the sequencing quality of the reads.
[0043] Clean Reads: In this invention, the high-quality data obtained after filtering the original offline data is called Clean Reads.
[0044] Heatmap: An image showing the interaction between different regions, with the horizontal and vertical axes representing the region names.
[0045] Contig / Scaffold: The genome at the level of Contig and / or Scaffold.
[0046] The software can determine the first subset that can be aligned to the reference genome dataset and the second subset that cannot: Using the software's default parameters and standards, and employing a penalty scoring method, the software can provide both the mapped first subset and the unmapped second subset. For example, using Bowtie2 software, Read1 and Read2 are aligned one-end to the reference genome, respectively, yielding BAM format alignment result files. Those skilled in the art can then obtain the first subset that can be aligned to the reference genome dataset and the second subset that cannot.
[0047] Linkage sites of the library after enzyme digestion: In the construction of Hi-C libraries, the typical procedure includes formaldehyde cross-linking, restriction endonuclease digestion, biotin labeling, DNA ligation, streptavidin enrichment, biotin fishing, final repair, 3' A addition, adapter addition, PCR amplification, and sequencing. In this invention, the linking sites of the library after enzyme digestion are the DNA ligation sites after biotin labeling and before streptavidin enrichment.
[0048] The first aspect of this invention provides a gene-assisted assembly apparatus, comprising: a Hi-C library construction and sequencing unit, an alignment and selection unit, a preliminary assembly unit, and a processing and screening unit, wherein...
[0049] The Hi-C library construction and sequencing unit is used to construct and sequence a first dataset from a DNA sample using Hi-C library construction and sequencing.
[0050] The comparison and selection unit is used to compare and select the first dataset with the reference genome dataset to obtain the second dataset;
[0051] The preliminary assembly unit is used to perform preliminary assembly of the second dataset to obtain a preliminary assembled dataset;
[0052] The processing and filtering unit is used to process and filter the preliminary assembled dataset;
[0053] In this invention, preferably, the gene-assisted assembly device does not include an error correction unit and does not require error correction of the reference genome dataset.
[0054] The gene-assisted assembly apparatus of the present invention does not include a bin-cutting unit, and it is not necessary to cut the contig / scaffold level genome into bin-level genomes.
[0055] In this invention, the sample can be a DNA sample.
[0056] According to the apparatus of the present invention, preferably, the data includes Read.
[0057] According to the apparatus of the present invention, preferably, the reference genome dataset is selected from a reference genome draft based on second-generation sequencing or a reference genome draft based on third-generation sequencing.
[0058] According to the apparatus of the present invention, preferably, the dataset includes multiple reads. More preferably, it includes multiple pairs of reads. A pair of reads includes Read1 and Read2, which are two sequences with opposite directions in the insert sizes of two complementary strands obtained from paired-end sequencing data. Before alignment to the single-stranded reference genome, one of the reads is escaped and then the alignment is performed.
[0059] According to the apparatus of the present invention, preferably, the comparison and selection unit comprises a first component, a second component, and a third component, wherein,
[0060] The first component is used to perform a first alignment of each data in the first dataset with a reference genome dataset to obtain a first subset that can be aligned to the reference genome dataset and a second subset that cannot be aligned to the reference genome dataset;
[0061] The second component is used to search for the link sites of the library after each data in the second subset is digested with enzymes, to break the data in the second subset at the link sites of the digested library found, and to perform a second alignment with the reference genome dataset to obtain a third subset that can be aligned to the reference genome dataset.
[0062] The third component is used to merge and select the first subset and the third subset to obtain the second data.
[0063] According to the apparatus of the present invention, preferably, the selection is used to select data whose two ends are both aligned to a unique position in the reference genome dataset.
[0064] According to the apparatus of the present invention, preferably, the first comparison is a single-end comparison or a double-end comparison; more preferably, it is a single-end comparison. In the present invention, the result file of the single-end comparison can be stored separately, making file processing more convenient.
[0065] According to the apparatus of the present invention, preferably, the second comparison is a single-end comparison or a double-end comparison; more preferably, it is a single-end comparison. In the present invention, the result file of the single-end comparison can be stored separately, making file processing more convenient.
[0066] According to the apparatus of the present invention, preferably, the preliminary assembly unit is based on LACHESIS software.
[0067] According to the apparatus of the present invention, preferably, the preliminary assembled dataset includes preliminary assembled Contigs / Scaffolds. Specifically, the preliminary assembly unit performs preliminary assembly of the second dataset based on the fact that the interaction frequency on the same chromosome within the cell nucleus is higher than the interaction frequency on different chromosomes, to obtain preliminary assembled Contigs and / or Scaffolds.
[0068] According to the apparatus of the present invention, preferably, the processing and screening unit is based on LACHESIS software.
[0069] According to the apparatus of the present invention, preferably, the processing and screening unit comprises a fourth component and a fifth component, wherein,
[0070] The fourth component is used to allocate the preliminary assembly data to the chromosome group;
[0071] The fifth component is used to sort, orient, and filter the preliminary assembly data assigned to each chromosome group.
[0072] According to the apparatus of the present invention, preferably, the allocation is based on a clustering method.
[0073] According to the apparatus of the present invention, preferably, the number of chromosome groups is 1-100, more preferably 2-70. The number of chromosome groups can be set according to prior values; for example, the number of chromosome groups in fruit flies can be 8, and the number of chromosome groups in some aquatic plants can be 60, etc. The number of chromosome groups can be consistent with the number measured in the species laboratory.
[0074] According to the apparatus of the present invention, preferably, the fifth component includes a sorting component, a orientation component, and a filtering component.
[0075] According to the apparatus of the present invention, preferably, the sorting component sorts the preliminary assembly data assigned to each chromosome group based on the strength of the interaction relationships of the preliminary assembly data; more preferably, the sorting component constructs an acyclic spanning tree based on the interaction relationships of the contigs / scaffolds within a chromosome group, and selects the root tree with the highest confidence from it. Even more preferably, the sorting component constructs an acyclic spanning tree based on the interaction relationships of the contigs / scaffolds within a chromosome group, selects the root tree with the highest confidence from it, and inserts the remaining fragmented contigs / scaffolds into appropriate positions at the root according to their association relationships, ultimately constructing an accurate contig / scaffold order for the internal groups of a chromosome. Some extremely short contigs / scaffolds are discarded. This operation can be performed using LACHESIS software.
[0076] According to the apparatus of the present invention, preferably, the orientation component is used to traverse the possible directions of the Contig / Scaffold through a weighted directed acyclic graph, and for the order determination of each Contig / Scaffold, construct a scoring function based on the difference between positive and negative interaction relationships and use it as the basis for confidence determination, and select the orientation with the highest confidence.
[0077] According to the apparatus of the present invention, preferably, the screening conditions set by the screening component are selected from one or more of the following: number of clusters, proportion of Contig / Scaffold used during orientation, proportion of Contig / Scaffold in the root tree, and ratio of the shortest chromosome length to the longest chromosome length.
[0078] According to the apparatus of the present invention, preferably, the processing and filtering unit further includes a re-filtering component, which is used to change the conditions for preliminary assembly again to obtain multiple preliminary assembly data when the optimal preliminary assembly data obtained by preliminary assembly of the second data cannot be output as assembly data, and to reprocess and filter the preliminary assembly data to select the optimal preliminary assembly data as assembly data.
[0079] According to the apparatus of the present invention, preferably, the conditions are selected from one or more of the following: the minimum number of restriction enzyme sites in the Contig / Scaffold sequence in cluster analysis, the maximum link depth in the Contig / Scaffold sequence in cluster analysis, the ratio of the number of interactions between the Contig / Scaffold sequence and the target cluster to the number of interactions between other clusters in cluster back-interpolation, and the minimum number of restriction enzyme sites in the Contig / Scaffold sequence within the root tree in ordination analysis.
[0080] According to the apparatus of the present invention, preferably, the Hi-C library construction and sequencing unit includes a cross-linking and cell lysis component, a chromatin digestion component, a Hi-C sample preparation component, a library construction component, and a sequencing component, wherein,
[0081] The cross-linking and cell lysis components are used to cross-link and lyse DNA samples to obtain lysed cells;
[0082] The chromatin digestion component is used to digest the chromatin of the lysed cells using restriction endonucleases;
[0083] The Hi-C sample preparation component is used to perform biotin labeling, blunt end ligation, and DNA purification and extraction on the digested sample to prepare Hi-C samples.
[0084] The library construction component is used to construct libraries from Hi-C samples to obtain library products;
[0085] The sequencing component is used to sequence the library products to obtain the raw sequencing dataset.
[0086] According to the apparatus of the present invention, preferably, the Hi-C library construction and sequencing unit further includes a sample quality control component for performing Hi-C sample quality control on the purified and extracted DNA to prepare Hi-C samples, and then constructing a library from the quality-controlled Hi-C samples to obtain library products. More preferably, the prepared Hi-C samples are subjected to gel electrophoresis, and a suitable Hi-C sample is selected as the quality-controlled Hi-C sample based on the substrate fragment length selected by the gel electrophoresis results.
[0087] According to the apparatus of the present invention, preferably, the Hi-C library construction and sequencing unit further includes a filtering component for the raw sequencing dataset. More preferably, the filtering includes one or more of the following methods:
[0088] (1) Remove data with bases greater than 5bp that are contaminated with connectors from the original offline dataset;
[0089] (2) Remove data from the original dataset where more than 50% of the bases have a quality value Q ≤ 19; and
[0090] (3) Remove data from the original data that contains more than 5% N bases.
[0091] According to a specific embodiment of the present invention, a gene-assisted assembly device includes, but is not limited to, a Hi-C library construction and sequencing unit, an alignment and selection unit, a preliminary assembly unit, and a processing and screening unit. The Hi-C library construction and sequencing unit is used to construct and sequence a first dataset from a DNA sample using Hi-C libraries. The alignment and selection unit is used to align and select the first dataset with a reference genome dataset to obtain a second dataset. The preliminary assembly unit is used to perform preliminary assembly on the second dataset to obtain a preliminary assembled dataset. The processing and screening unit is used to process and screen the preliminary assembled dataset. The gene-assisted assembly device does not include an error correction unit and does not require error correction of the reference genome dataset. The steps for gene-assisted assembly using this device are as follows:
[0092] The first step is to construct a Hi-C library and sequence the DNA sample to obtain initial data. Specifically,
[0093] First, the experimental DNA samples are cross-linked. After cross-linking, cell lysis is performed, and samples are taken for extraction and quality testing. Once the quality is deemed satisfactory, the samples proceed to the "Hi-C fragment" preparation process.
[0094] In the preparation process of the "Hi-C fragment", restriction endonucleases, such as HindIII / MboI (the "cleavage site" and "cleavage fragment" mentioned later refer to this enzyme), are used to digest chromatin and samples are taken to detect the cleavage effect.
[0095] The enzyme-digested samples were biotin-labeled, blunt-end ligated, and DNA purified and extracted to prepare Hi-C samples, which were then sampled for DNA quality testing.
[0096] After passing the test, the sample proceeds to the standard library construction process, which includes the following steps: the biotin-labeled ends of the Hi-C sample fragments that have passed the quality test are removed, the fragments are sonicated and repaired, then base A is added, the biotin-containing fragments are extracted, and sequencing adapters are added to form adapter-added products; then PCR conditions are screened and amplified to obtain the library products.
[0097] An ideal library consists of DNA fragments from two different restriction enzymes joined together, with the ligation site forming a new restriction enzyme cleavage site (e.g., if HindIII is used in the chromatin digestion step described above, the new cleavage site formed after blunt-end ligation is recognized and cleaved by Nhel). The library amplification products are sampled for "Hi-C fragment ligation site quality control testing," and the entire library preparation is complete once the test is passed.
[0098] In obtaining a quality-controlled Hi-C library, before amplification and sequencing, it is necessary to select substrate fragments of appropriate length (50-1000 bp) using electrophoresis gel mapping. By aligning reads to the draft, the theoretical length distribution of the Hi-C library insert fragments is evaluated to determine whether it conforms to the predetermined length range. This is a typical and intuitive quality control method for verifying library quality in the information analysis stage. Generally, an insert fragment length distribution conforming to 50-1000 bp indicates consistency with the substrate selection range in the experimental stage; furthermore, the more paired-end reads greater than 100 bp in the insert fragment length distribution, the better the library.
[0099] After the constructed library passed quality control, it was sequenced using the Illumina platform with the PE150 sequencing strategy.
[0100] Sequencing results include some raw reads, which contain low-quality sequences, adapter-contaminated sequences, sequences with more than 5% N bases, and clean reads. A higher proportion of clean reads indicates better data quality. To ensure the quality of subsequent information analysis, the raw sequences are filtered to obtain high-quality clean reads before further analysis. The data processing steps are as follows:
[0101] (1) Remove adapter-contaminated reads, such as reads with more than 5 bp of adapter-contaminated bases. For paired-end sequencing, if one end is contaminated by the adapter, remove the reads from both ends.
[0102] (2) Remove low-quality reads, such as reads with a quality value of Q≤19 accounting for more than 50% of the total bases. For paired-end sequencing, if one end is a low-quality read, both ends of the reads will be removed.
[0103] (3) Remove reads containing more than 5% N bases. For example, for paired-end sequencing, if one end contains more than 5% N bases, the reads at both ends will be removed.
[0104] The high-quality Clean Reads obtained after the above data processing steps are what this invention refers to as the first data.
[0105] The second step involves comparing and selecting from the first dataset and the reference genome dataset to obtain the second dataset. Specifically:
[0106] In this step, the first data, after filtering out the reads that are prone to causing analytical errors, is used as the first reference data of the filtered genome draft as the reference sequence. Bowtie2 software is called for alignment. Specifically, two reads in the first data, such as Read1 and Read2, are aligned one-to-one with the reference genome dataset to obtain a BAM format alignment result file. This yields a first subset that can be aligned with the reference genome dataset and a second subset that cannot. The unaligned reads in the second subset (also known as Unmap Reads) are then examined for the linking sites of the digested library. The portion after the linking sites of the digested library is truncated, and the reads are aligned one-to-one with the reference genome dataset again to obtain a third subset that can be aligned with the reference genome dataset. The first subset and the third subset are merged, and pairs of reads with both ends aligned to unique positions in the reference genome dataset are selected as the second data in this invention. This second data is used for subsequent analysis.
[0107] Based on the experimental principles of Hi-C library construction, different molecular types are generated during the Hi-C library construction process. According to the alignment position and orientation of reads in the genome, they can be specifically classified as self-circles, dangling ends, dumped pairs, and valid pairs. The data used for subsequent analysis are valid paired-end reads (i.e., valid pairs), meaning a pair of reads aligned to two different restriction enzyme fragments, and the actual fragment size matches the theoretical fragment size. In this invention, such valid data is considered the second data set. All other data types are considered invalid paired reads and are excluded in this step.
[0108] The third step is to perform preliminary assembly of the second data to obtain preliminary assembled data.
[0109] In this step, parameter traversal is used to perform assisted assembly. Since the interaction frequency on the same chromosome within the cell nucleus is higher than the interaction frequency between different chromosomes, a clustering algorithm can be used to assign the initially assembled contigs / scaffolds to various chromosome groups. This initial assembly process is performed using LACHESIS software.
[0110] The fourth step involves processing and filtering the preliminary assembled dataset. Specifically, the preliminary assembled data is allocated to chromosome groups; and the preliminary assembled data allocated to each chromosome group is sorted, oriented, and filtered. This step is performed using LACHESIS software.
[0111] In this step, the initially assembled data (i.e., Contig / Scaffold) is clustered. Based on the alignment results and the principle that the interaction frequency on the same chromosome within the cell nucleus is higher than the interaction frequency on different chromosomes, the initially assembled Contig / Scaffold is assigned to each chromosome group using agglomerative hierarchical clustering (a bottom-up hierarchical clustering algorithm) with LACHESIS software.
[0112] Chromosome clusters refer to the clustering results obtained by the above algorithm, with each cluster considered as a chromosome cluster.
[0113] The fifth step is to sort and filter the preliminary assembly data assigned to each chromosome group.
[0114] In this step, the order of Contigs / Scaffolds is determined. The initial assembly data (Contigs / Scaffolds) assigned to a chromosome group in the previous step are used to construct the longest acyclic spanning tree based on interactions. The "trunk" (root tree) with the highest confidence is selected from this tree. The remaining fragmented Contigs / Scaffolds are then inserted into appropriate positions in the root tree according to their associations. This process ultimately constructs the accurate Contig / Scaffold order for the internal groups of a chromosome. Some extremely short Contigs / Scaffolds are discarded by the LACHESIS software.
[0115] The sixth step is to determine the orientation of the sorted and filtered preliminary assembled data.
[0116] In this step, within the ordered Contig / Scaffold group of each chromosome, all possible directions of the Contig / Scaffold are traversed using a weighted directed acyclic graph (WDAG). For the order determination of each Contig / Scaffold, a scoring function is constructed based on the difference between forward and reverse interactions to describe the confidence level when determining a direction. This is then determined using LACHESIS software, thereby achieving Contig / Scaffold orientation.
[0117] The seventh step is to evaluate the preliminary assembly data with the determined direction and select the preliminary assembly data with the best results as the final assembly data.
[0118] In this step, the best assembly results are selected. Based on the assembly results obtained by running LACHESIS software under different parameters, different weights are assigned to score the results according to the consistency of the number of chromosomes in the cluster with the expected result, the genome mounting rate, the trunk value, and the consistency of chromosome length. The assembly results with higher scores are selected as the best. Preferably, the selection parameters are selected from one or more of the following: the number of clusters, the proportion of contigs / scaffolds used during orientation, the proportion of contigs / scaffolds in the root tree, and the ratio of the shortest chromosome length to the longest chromosome length.
[0119] More preferably, if the optimal preliminary assembly data cannot be used as the output assembly data, the conditions for preliminary assembly are changed again to obtain multiple preliminary assembly data in the preliminary assembly of the second data, and the preliminary assembly data is reprocessed and screened to select the optimal preliminary assembly data as the assembly data. More preferably, the conditions are selected from the minimum number of restriction enzyme sites in the Contig / Scaffold sequence in the cluster analysis, the maximum Link depth in the Contig / Scaffold sequence in the cluster analysis, the ratio of the number of interactions between the Contig / Scaffold sequence and the target cluster to the number of interactions between other clusters in the cluster back-interpolation, and the minimum number of restriction enzyme sites in the Contig / Scaffold sequence in the root tree in the sorting analysis.
[0120] The above devices can be used to assemble a genome at the chromosome level.
[0121] In this invention, the assembled chromosome-level genome construction heatmap can be verified, specifically:
[0122] Heatmap verification of assembled chromosomes is a standard for checking the accuracy of assisted assembly. The method for constructing heatmaps is as follows: For chromosomes used in assisted assembly, they are cut into equal-length bins (common bin sizes include 1 MB and 500 kb). The effective alignment reads between each pair of bins are used as the strength signal of the interaction between them to construct a heatmap.
[0123] Based on the two fundamental principles of Hi-C assisted assembly, two basic rules exist in heatmap verification: (i) the interaction strength between two bins within a chromosome is stronger than the interaction strength between two bins between chromosomes; (ii) on the same chromosome, the interaction strength between two bins that are close together is stronger than the interaction strength between two bins that are far apart.
[0124] Ideally, the global heatmap or the internal heatmap of a chromosome shows strong interactions along the diagonal, with the interaction strength gradually weakening outwards from the diagonal. In addition, in some species, due to the presence of AB Compartment, the heatmap may also show a grid-like pattern.
[0125] A second aspect of the present invention provides a chromosome-level genome assembled according to the gene-assisted assembly apparatus described above.
[0126] The third aspect of the present invention provides the above-described gene-assisted assembly apparatus and the application of the assembled chromosome-level genome in chromosome interactions, whole-genome resequencing, chromosome evolution and differentiation, and epigenomics.
[0127] The beneficial effects of this invention are:
[0128] (1) The device of the present invention can increase the mounting rate of Hi-C assisted assembly to over 95%.
[0129] (2) The device of the present invention does not require the use of population data; during assembly, it is not necessary to cut the Contig / Scaffold level genome into bin level genomes, that is, it does not include the unit of cutting the Contig / Scaffold level genome into bin level genomes, thus avoiding the loss of some bin information and the resulting incomplete information; the device of the present invention also does not require an error correction unit, thus avoiding the error correction error caused by changing the original draft sequence due to error correction; moreover, the device of the present invention is both time-saving and labor-saving compared to genetic maps.
[0130] (3) The comparison and selection unit of the device of the present invention can obtain more complete and accurate data, thereby improving the mounting rate of Hi-C assisted assembly.
[0131] (4) The preliminary assembly unit of the device of the present invention processes and filters the preliminary assembly dataset, especially under the specific filtering conditions in the preferred case, which can find the optimal result and ensure the reliability of the assembly result.
[0132] (5) The processing and screening unit of the device of the present invention distributes the initially assembled Contig / Scaffold to each chromosome group through a clustering algorithm. Within each chromosome cluster, the Contig / Scaffold is sorted and the Contig / Scaffold direction is determined. In particular, based on the assembly results under different parameters, screening conditions are set to select the better assembly results and ensure the reliability of the assembly results.
[0133] (6) In a preferred embodiment, the processing and screening unit of the present invention further includes a re-screening component. When the optimal preliminary assembly data selected cannot be output as assembly data, the conditions for preliminary assembly will be changed again in the preliminary assembly data obtained by preliminary assembly of the second data to obtain multiple preliminary assembly data again. Through specific screening condition parameters, the mounting rate and reliability of Hi-C assisted assembly can be significantly improved. Attached Figure Description
[0134] Figure 1 This is a schematic diagram of the gene-assisted assembly device of Embodiment 1 of the present invention.
[0135] Figure 2 This is a frequency distribution histogram of the inserted segment length in Embodiment 1 of the present invention.
[0136] Figure 3 This is a heatmap of the global Hi-C interaction of the genome in Example 1 of the present invention.
[0137] Figure 4 This is a heatmap of the Hi-C interaction of a single chromosome in Example 1 of the present invention.
[0138] The present invention will now be described with reference to specific embodiments. It should be noted that these embodiments are merely illustrative and should not be construed as limiting the present invention.
[0139] Example 1
[0140] A gene-assisted assembly device, the structural schematic diagram of which is shown below. Figure 1As shown, the device comprises: a Hi-C library construction and sequencing unit, an alignment and selection unit, a preliminary assembly unit, and a processing and screening unit. The Hi-C library construction and sequencing unit is used to construct and sequence a first dataset from a DNA sample using Hi-C libraries. The alignment and selection unit is used to align and select the first dataset with a reference genome dataset to obtain a second dataset. The preliminary assembly unit is used to perform preliminary assembly on the second dataset to obtain a preliminary assembled dataset. The processing and screening unit is used to process and screen the preliminary assembled dataset. This gene-assisted assembly device does not include an error correction unit and does not require error correction of the reference genome dataset.
[0141] Specifically:
[0142] 1. The Hi-C library construction and sequencing unit includes a cross-linking and cell lysis component, a chromatin digestion component, a Hi-C sample preparation component, a library construction component, and a sequencing component. The cross-linking and cell lysis component is used to cross-link and lyse the DNA sample to obtain lysed cells. The chromatin digestion component is used to digest the lysed cells using restriction endonucleases. The Hi-C sample preparation component is used to perform biotin labeling, blunt end ligation, and DNA purification and extraction on the digested sample to prepare a Hi-C sample. The library construction component is used to construct a library from the Hi-C sample to obtain a library product. The sequencing component is used to sequence the library product to obtain the raw sequencing dataset.
[0143] The Hi-C library construction and sequencing unit also includes a sample quality control component, used for purifying and extracting DNA to prepare Hi-C samples for Hi-C sample quality control, and then constructing a library from the quality-controlled Hi-C samples to obtain the library product. The prepared Hi-C samples are then subjected to gel electrophoresis, and based on the gel electrophoresis results, a suitable Hi-C sample is selected as the quality-controlled Hi-C sample by choosing the substrate fragment length.
[0144] The Hi-C library construction and sequencing unit also includes a filtering component for the raw sequencing dataset. The filtering method is as follows:
[0145] (1) Remove data with bases greater than 5bp that are contaminated with connectors from the original offline dataset;
[0146] (2) Remove data from the original dataset where more than 50% of the bases have a quality value Q ≤ 19; and
[0147] (3) Remove data from the original data that contains more than 5% N bases.
[0148] 2. The alignment and selection unit comprises a first component, a second component, and a third component. The first component performs a first alignment of each data point in the first dataset with a reference genome dataset, obtaining a first subset that can align to the reference genome dataset and a second subset that cannot. The second component performs a linker site search on each data point in the second subset after enzyme digestion, breaks the data in the second subset at the searched linker sites, and performs a second alignment of the broken data with the reference genome dataset, obtaining a third subset that can align to the reference genome dataset. The third component merges and selects from the first and third subsets to obtain the second data. The selection process identifies data whose both ends align to a unique position in the reference genome dataset; both the first and second alignments are single-end alignments.
[0149] 3. The preliminary assembly unit, based on LACHESIS software, is used to perform preliminary assembly of the second dataset to obtain the preliminary assembled Contig / Scaffold dataset.
[0150] 4. The processing and filtering unit is based on LACHESIS software and includes a fourth and a fifth component. The fourth component uses clustering methods to allocate the initial assembled data to chromosome clusters. The fifth component includes a sorting component, an orientation component, and a filtering component, used to sort, orient, and filter the initial assembled data allocated to each chromosome cluster. The sorting component constructs an acyclic spanning tree based on the interaction relationships of contigs / scaffolds within a chromosome cluster and selects the root tree with the highest confidence. The orientation component traverses the possible directions of contigs / scaffolds through a weighted directed acyclic graph. For the order determination of each contig / scaffold, a scoring function is constructed based on the difference between positive and negative interaction relationships, which serves as the basis for confidence determination, and the orientation with the highest confidence is selected. The filtering parameters set by the filtering component are selected from all of the following: the number of clusters, the proportion of contigs / scaffolds used in orientation, the proportion of contigs / scaffolds in the root tree, and the ratio of the shortest chromosome length to the longest chromosome length. The processing and filtering unit also includes a re-filtering component, which is used to change the conditions for preliminary assembly again to obtain multiple preliminary assembly data when the optimal preliminary assembly data cannot be output as assembly data. This is done by reprocessing and filtering the preliminary assembly data obtained from the preliminary assembly of the second data, and then filtering the preliminary assembly data again to select the optimal preliminary assembly data as the assembly data. The conditions are selected from all of the following in cluster analysis: the minimum number of restriction enzyme sites in the Contig / Scaffold sequence, the maximum Link depth in the Contig / Scaffold sequence, the ratio of the number of interactions between the Contig / Scaffold sequence and the target cluster to the number of interactions with other clusters in cluster back-interpolation, and the minimum number of restriction enzyme sites in the Contig / Scaffold sequence within the root tree in ordination analysis.
[0151] The gene-assisted assembly using this gene-assisted assembly device is as follows:
[0152] 1. Construct and sequence a Hi-C library from the DNA sample to obtain initial data.
[0153] First, the DNA sample is cross-linked. After cross-linking, cells are lysed, and samples are taken for extraction and quality testing. Once the quality is satisfactory, the sample proceeds to the "Hi-C fragment" preparation process.
[0154] In the preparation process of the "Hi-C fragment", the restriction endonuclease HindIII / MboI (hereinafter referred to as "restriction site" and "restriction fragment" both refer to this enzyme) is used to digest chromatin, and samples are taken to detect the enzyme digestion effect.
[0155] The enzyme-digested samples were biotin-labeled, blunt-end ligated, and DNA purified and extracted to prepare Hi-C samples, which were then sampled for DNA quality testing.
[0156] After passing the test, the sample proceeds to the standard library construction process, which consists of the following steps: the biotin-labeled ends of the Hi-C sample fragments that have passed the quality test are removed, the fragments are sonicated and repaired, then base A is added, the biotin-containing fragments are extracted, and sequencing adapters are added to form adapter-added products; then PCR conditions are screened and amplified to obtain the library products.
[0157] Some raw sequences obtained from sequencing may contain sequencing adapter sequences and low-quality sequences. In order to ensure the quality of information analysis data, the raw sequences are filtered and duplicate reads are removed to obtain clean reads (high-quality reads).
[0158] 1) Remove adapter-contaminated reads (reads with more than 5 bp of adapter-contaminated bases; for paired-end sequencing, if one end is contaminated by the adapter, remove reads from both ends).
[0159] 2) Remove low-quality reads (reads with a quality value of Q≤19 account for more than 50% of the total bases; for paired-end sequencing, if one end is a low-quality read, both ends of the reads will be removed).
[0160] 3) Remove reads with an N ratio greater than 5% (for paired-end sequencing, if one end has an N ratio greater than 5%, both ends of the reads will be removed). Since the insert fragments in Hi-C library preparation are chimeric fragments, there are cases where single-end reads cross chimeric sites during sequencing. In order to ensure the utilization of sequencing data, we also truncate the reads to a certain length (100bp). The data analyzed later are all based on the truncated reads.
[0161] The data filtering and statistical results are shown in Table 1:
[0162] Table 1
[0163] Length of Reads after truncation (bp) 100 The original number of reads after shutdown (in pairs) 428,015,735 The number of bases in the original sequence (bp). 128,404,720,500 The number of high-quality reads obtained after filtering (in pairs) 422,405,187 The percentage (%) of the high-quality reads obtained after filtering to the original number of reads downloaded. 98.69 The number of low-quality reads removed (in pairs) 2,122,337 The percentage of low-quality reads removed relative to the original number of reads processed (%) 0.5 Remove the logarithm of Reads containing more than 5% N. 291,287 The percentage of Reads logs that exclude those containing more than 5% N, relative to the original total number of Reads (%). 0.07 The number of Reads after removing connector contamination (pairs) 3,196,924 The percentage of Reads that have had their connector contamination removed, relative to the original number of Reads after disconnection (%) 0.75 The percentage of bases with a sequencing quality value greater than 30 (error rate less than 0.1%) in the raw reads, out of the total number of bases, (%). 90.06 The percentage of bases with a sequencing quality value greater than 30 (error rate less than 0.1%) in the high-quality reads obtained after filtering, (%). 92.35
[0164] 2. Base distribution of sequencing data
[0165] Using the base positions of the filtered Clean Reads sequences as the x-axis and the proportions of ATCG bases and N (unidentified bases) at each position as the y-axis, a base distribution map of the Clean Reads was obtained. In the sequencing, the proportions of A and T bases, and the proportions of G and C bases were equal at each position of the Clean Reads, indicating that there was no sequencing bias.
[0166] 3. Comparison and fragment analysis
[0167] 3.1 Genome alignment analysis
[0168] Bowtie2 is used to align each pair of reads in the first dataset. Each pair of reads contains Read1 and Read2. Read1 and Read2 are then aligned one-end to the reference genome dataset to obtain a BAM format alignment result file. This results in a first subset that can be aligned to the reference genome dataset and a second subset that cannot. For the unaligned reads in the second subset (also known as Unmap Reads), the linking sites of the digested library are located. The portion of the digested library after the linking sites is truncated, and then aligned one-end to the reference genome dataset again to obtain a third subset that can be aligned to the reference genome dataset. The first and third subsets are merged, and the pairs of reads where both ends are aligned to a unique position in the reference genome dataset are selected. This second set is used for subsequent analysis.
[0169] The number of reads and the alignment rate are shown in Table 2:
[0170] Table 2
[0171] The number of high-quality reads obtained after filtering (in pairs) 422,405,187 No genome reads were matched at either end. 9,369,416 The percentage of high-quality reads that did not align to either end of the genome sequence (%) 2.22 The number of reads that align to the genome at only one end is (pairs). 60,894,278 The percentage of reads that only aligned to one end of the genome out of the total number of high-quality reads (%) 14.42 The number of reads aligned to multiple locations in the genome (pairs) 126,238,167 The percentage of high-quality reads that aligned to multiple locations in the genome (%) 29.89 The number of reads aligned to a unique location in the genome (pairs) 225,903,326 The percentage of high-quality reads that match a unique genomic location (%) 53.48
[0172] 3.2 Distribution of Insert Fragment Lengths in Hi-C Library
[0173] By aligning reads to a reference genome, the theoretical length distribution of Hi-C library inserts is evaluated to determine whether they conform to the predetermined length range.
[0174] From the filtered high-quality Reads, 50,000 pairs of Reads were randomly selected. The insertion segment length information for each pair of Reads was statistically analyzed. A frequency distribution histogram of the insertion segment length was plotted with the insertion segment length on the x-axis and frequency on the y-axis, as shown below. Figure 2 As shown.
[0175] Typically, the length distribution of inserted fragments should conform to 50~1000bp, indicating consistency with the substrate selection range in the experimental stage; simultaneously, the more paired-end reads longer than 100bp in the inserted fragment length distribution, the better the library quality. Figure 2 It can be seen that the substrate selection range is consistent with that in the experimental stage, and the library quality is high.
[0176] 4. Assisted assembly
[0177] This was accomplished using LACHESIS software.
[0178] 4.1 Contig / Scaffold Cluster Analysis
[0179] For diploid genomes, the LACHESIS software uses a bottom-up hierarchical clustering algorithm to assign the initially assembled Contig / Scaffold to n chromosome groups, where n is 14 in this project.
[0180] As a result, the initial assembled genome contained a total of 255 contigs / scaffolds with a total length of 408,446,700 bp. Clustering using LACHESIS software identified 59 contigs / scaffolds belonging to 14 chromosome groups, representing 23.14% of the total, with a total length of 392,622,225 bp, accounting for 96.13% of the total genome length.
[0181] The results of Contig / Scaffold clustering are shown in Table 3:
[0182] Table 3
[0183] Total number of contigs / scaffolds in the initial genome assembly (number of contigs / scaffolds) 255 Total length of Contig / Scaffold in the preliminary assembled genome (bp) 408,446,700 The total number of contigs / scaffolds clustered into n chromosome groups. 59 The percentage (%) of the total number of contigs / scaffolds clustered into n chromosome groups to the total number of contigs / scaffolds in the initially assembled genome. 23.14 Total length of contigs / scaffolds clustered into n chromosome groups, (bp) 392,622,225 The percentage (%) of the total length of contigs / scaffolds clustered into n chromosome groups to the total length of contigs / scaffolds in the initially assembled genome. 96.13
[0184] 4.2 Contig / Scaffold Ordination Analysis
[0185] Within each chromosome's cluster, contigs / scaffolds are ordered. Based on the standardized interaction strength between contigs / scaffolds, the longest acyclic spanning tree is constructed for each contig / scaffold within a chromosome's cluster, and the "trunk" (root tree) with the highest confidence is selected. The remaining fragmented contigs / scaffolds are then inserted into appropriate positions at the root according to their association relationships, ultimately constructing the accurate contig / scaffold order for each chromosome's internal cluster.
[0186] As a result, among the 59 contigs / scaffolds in the cluster, 52 were sorted, accounting for 88.14% of the total contigs / scaffolds; the total length of the sorted contigs / scaffolds was 392,306,666 bp, accounting for 99.92% of the total contigs / scaffolds. Of the sorted contigs / scaffolds, 51 were located in the root trunk, accounting for 98.08% of the total sorted contigs / scaffolds; the total length of the root trunk contigs / scaffolds was 391,823,533 bp, accounting for 99.88% of the total length of the sorted contigs / scaffolds.
[0187] The Contig / Scaffold sorting results are shown in Table 4:
[0188] Table 4
[0189] Total number of contigs / scaffolds in the sorting (number of contigs / scaffolds) 52 The percentage of sorted contigs / scaffolds out of the total number of cluster contigs / scaffolds, (%) 88.14 Total length of the sorted contig / scaffold (bp) 392,306,666 Ordering(%): The proportion of the total length of the ordered contigs / scaffolds to the total length of the cluster contigs / scaffolds, (%) 99.92 Total number of contigs / scaffolds located in the trunk (number of contigs / scaffolds) 51 The percentage of total Contigs / Scaffolds located in the trunk out of the total number of sorted Contigs / Scaffolds (%) 98.08 Total length of Contig / Scaffold located in the trunk, (bp) 391,823,533 The percentage of total length of contigs / scaffolds located in the trunk relative to the total length of sorted contigs / scaffolds (%) 99.88
[0190] 4.4 Contig / Scaffold Targeted Analysis and Screening
[0191] Orientation: Within the Contig / Scaffold group of each chromosome with a defined order, all possible directions of the Contig / Scaffold are traversed through a weighted directed acyclic graph (WDAG). For the order determination of each Contig / Scaffold, a scoring function is constructed based on the difference between positive and negative interaction relationships and used as the basis for confidence determination. The orientation with the highest confidence is selected to achieve Contig / Scaffold orientation.
[0192] Screening: (1) Number of clusters (number of chromosomes), that is, the number of chromosomes assembled in the end is greater than or equal to the expected number of chromosomes and the smaller the difference, the higher the score. (2) The proportion of Contig / Scaffold used during orientation: the closer to 100%, the higher the score. (3) The proportion of Contig / Scaffold in the root tree: the closer to 100%, the higher the score. (4) The proportion of the shortest chromosome length to the longest chromosome length, the higher the proportion, the higher the score. Finally, the scores of the above four types are added together, and the highest score is taken as the final screening result.
[0193] When the optimal preliminary assembly data cannot be used as the final assembly data, the conditions for preliminary assembly are changed again to obtain multiple preliminary assembly data in the preliminary assembly data obtained from the second data. The preliminary assembly data is then reprocessed and filtered to select the optimal preliminary assembly data as the final assembly data. The conditions for preliminary assembly are changed again in the following ways: minimum number of restriction enzyme sites in the Contig / Scaffold sequence in cluster analysis; maximum link depth in the Contig / Scaffold sequence in cluster analysis; ratio of the number of interactions between the Contig / Scaffold sequence and the target cluster to the number of interactions with other clusters in cluster back-interpolation; and minimum number of restriction enzyme sites in the Contig / Scaffold sequence within the root tree in the ordination analysis.
[0194] As a result, among the 52 sorted contigs / scaffolds, all 52 were oriented, accounting for 100% of the sorted contigs / scaffolds; the total length of the oriented contigs / scaffolds was 392,306,666 bp, accounting for 100% of the total length of the sorted contigs / scaffolds; among the contigs / scaffolds with determined order in the preliminary assembled genome that belong to the trunk, 51 were oriented, accounting for 100% of the trunk contigs / scaffolds; the total length of the oriented contigs / scaffolds was 391,823,533 bp, accounting for 100% of the total length of the trunk contigs / scaffolds.
[0195] 5. Results
[0196] 5.1 Statistics of indicators for assisted genome assembly
[0197] The initial assembly of the genome contig resulted in the acquisition of Pseudomolecules (chromosome-level genomes). The Pseudomolecules were statistically analyzed, and the results are shown in Table 5.
[0198] Table 5
[0199] The final Pseudomolecule after clustering, sorting, and orientation The number of contigs / scaffolds attached to each Pseudomolecule (number of contigs / scaffolds). Pseudomolecule length (bp) chr1 4 33,853,100 chr2 2 34,507,833 chr3 3 31,512,209 chr4 3 29,500,627 chr5 5 28,859,902 chr6 2 28,791,375 chr7 4 28,155,678 chr8 6 34,387,650 chr9 5 26,934,034 chr10 5 26,790,618 chr11 2 25,880,527 chr12 5 24,795,317 chr13 2 22,190,354 chr14 4 16,151,242 Contig / Scaffold attachment status of all chromosomes 52 392,310,466 Unattached Contig / Scaffolds in the initial genome assembly 203 16,140,034
[0200] 5.2 Heatmap Verification
[0201] Figure 3 A global Hi-C interaction heatmap of the genome is displayed, with the horizontal and vertical axes representing chromosomes. The darker the color in the diagram, the stronger the interaction signal.
[0202] Figure 4A heatmap of Hi-C interactions on a single chromosome is shown.
[0203] As shown in Table 5, the gene-assisted assembly device using Hi-C technology of this invention can obtain a near-chromosome-level genome using third-generation assembled genome drafts and Hi-C sequencing data, achieving an attachment rate of over 96%. Furthermore, through... Figure 3 and Figure 4 The heatmap demonstrates that the chromosomes assembled using the gene-assisted assembly device of the present invention, which utilizes Hi-C technology, are of high quality.
[0204] The basic principles of the present invention have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in the present invention are merely examples and not limitations, and should not be considered as essential features of each embodiment of the present invention. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the present invention to the necessity of employing the aforementioned specific details.
[0205] The block diagrams of devices, apparatuses, devices, and systems involved in this invention are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0206] It should also be noted that in the apparatus, device, and method of the present invention, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions of the present invention.
[0207] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the invention. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the invention. Therefore, the invention is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0208] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the invention to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
Claims
1. A gene-assisted assembly device, the device comprising: Hi-C library construction and sequencing units, alignment and selection units, preliminary assembly units, and processing and screening units, among which, The Hi-C library construction and sequencing unit is used to construct and sequence a first dataset from a DNA sample using Hi-C library construction and sequencing. The comparison and selection unit is used to compare and select the first dataset with the reference genome dataset to obtain the second dataset; The preliminary assembly unit is used to perform preliminary assembly of the second dataset to obtain a preliminary assembled dataset; The processing and filtering unit is used to process and filter the preliminary assembled dataset; The comparison and selection unit includes a first component, a second component, and a third component, wherein, The first component is used to perform a first alignment of each data in the first dataset with a reference genome dataset to obtain a first subset that can be aligned to the reference genome dataset and a second subset that cannot be aligned to the reference genome dataset; The second component is used to search for the link sites of the library after each data in the second subset is digested with enzymes, to break the data in the second subset at the link sites of the digested library found, and to perform a second alignment with the reference genome dataset to obtain a third subset that can be aligned to the reference genome dataset. The third component is used to merge and select the first subset and the third subset to obtain the second data; The selection process is used to select data whose two ends are both aligned to a unique position in the reference genome dataset. The processing and filtering unit includes a fourth component and a fifth component, wherein, The fourth component is used to allocate the preliminary assembly data to the chromosome group; The fifth component is used to sort, orient, and filter the preliminary assembly data assigned to each chromosome group; The fifth component includes a sorting component, a direction component, and a filtering component; The sorting component is based on constructing an acyclic spanning tree from the contigs / scaffolds within a chromosome group according to their interaction relationships, and selecting the root tree with the highest confidence from it; The orientation component is used to traverse the possible directions of Contig / Scaffold through a weighted directed acyclic graph. For the order determination of each Contig / Scaffold, a scoring function is constructed based on the difference between positive and negative interaction relationships and used as the basis for confidence determination, and the orientation with the highest confidence is selected. The filtering parameters set by the filtering component are selected from one or more of the following: number of clusters, proportion of Contig / Scaffold used during orientation, proportion of Contig / Scaffold in the root tree, and ratio of the shortest chromosome length to the longest chromosome length.
2. The apparatus according to claim 1, wherein, The reference genome dataset is selected from a second-generation sequencing-based reference genome sketch or a third-generation sequencing-based reference genome sketch; and / or, the dataset includes multiple reads.
3. The apparatus according to claim 1, wherein, The preliminary assembly unit is based on LACHESIS software, and / or the preliminary assembly dataset includes preliminary assembled Contig / Scaffolds.
4. The apparatus according to claim 1, wherein, The processing and filtering unit is based on LACHESIS software.
5. The apparatus according to claim 1, wherein, The allocation is based on a clustering method.
6. The apparatus according to claim 1, wherein, The processing and filtering unit further includes a re-filtering component, which is used to change the conditions for preliminary assembly again to obtain multiple preliminary assembly data when the optimal preliminary assembly data cannot be output as assembly data, and to reprocess and filter the preliminary assembly data to select the optimal preliminary assembly data as assembly data.
7. The apparatus according to claim 6, wherein, The conditions are selected from one or more of the following: the minimum number of restriction enzyme sites in the Contig / Scaffold sequence in cluster analysis, the maximum Link depth in the Contig / Scaffold sequence in cluster analysis, the ratio of the number of interactions between the Contig / Scaffold sequence and the target cluster to the number of interactions between other clusters in cluster back-interpolation, and the minimum number of restriction enzyme sites in the Contig / Scaffold sequence within the root tree in ordination analysis.
8. The application of the gene-assisted assembly device according to any one of claims 1-7 in bioinformatics analysis, wherein the bioinformatics analysis includes chromosome interaction analysis, whole genome resequencing data analysis, chromosome evolution and differentiation analysis, or epigenomics data analysis.
Citation Information
Patent Citations
Data processing method and device for genomes
CN103810402A
Method for obtaining more accurate chromosome level genome
CN112908415A
Method and device for eliminating redundancy of high-heterozygous diploid sequence assembling result and application of method and device
CN113782101A