Macrobrachium rosenbergii reference genome structure correction and quality assessment method and system
By identifying and correcting false split sites and inversion errors in the Macrobrachium rosenbergii genome using Hi-C interaction signals, and combining global and local LAI and BUSCO determinations, the problem of structural errors in the Macrobrachium rosenbergii genome was solved, achieving more accurate genome structure and repetitive sequence continuity, supporting genetic breeding and functional gene research.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG DANSHUI FISHERY RESEARCH INSTITUTE (ZHEJIANG DANSHUI FISHERY ENVIRONMENTAL MONITORING STATION)
- Filing Date
- 2026-03-06
- Publication Date
- 2026-06-19
AI Technical Summary
Because the genome of the giant freshwater prawn is rich in repetitive sequences and has a high degree of heterozygosity, existing technologies are prone to structural errors such as false splits, mislinkings, and inversions during chromosome-level assembly. Furthermore, the lack of objective sequence-level verification of error correction actions affects the reliability of genetic analysis and breeding applications.
Hi-C interactive signals are used to identify false split sites and inversion errors. Structural errors are corrected through merging, splitting or redirection operations. Global LAI and local LAI are combined with BUSCO and consistency threshold for joint judgment to form a repeatable and quantifiable error correction closed loop.
To obtain a more accurate, more continuous, and more complete reference genome of the giant freshwater prawn, providing stable benchmark data to support genetic breeding and functional gene research.
Smart Images

Figure CN121811963B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of bioinformatics and genome assembly quality control, specifically to a method and system for error correction and quality assessment of the reference genome structure of Macrobrachium rosenbergii. Background Technology
[0002] The giant freshwater prawn (Macrobrachium rosenbergii) is one of the world's most important freshwater economic shrimp species, characterized by rapid growth, large size, strong adaptability, and high market demand. With the advancement of molecular breeding and other research, high-quality reference genomes have become a fundamental data resource supporting related studies. However, the giant freshwater prawn genome is typically characterized by its large size, high heterozygosity, and abundant repetitive sequences (especially a large number of transposon-related repetitions), which increases the risk of errors in de novo assembly processes such as contig splicing, scaffold construction, chromosome mounting, and structural correction. Particularly in regions with dense repetitive sequences, assembly algorithms are more prone to two opposing but equally fatal structural errors: "erroneous breakage" (fragmentation) or "over-connection" (incorrect splicing). This results in inconsistencies between the reference genome and the actual karyotype at the macroscopic structural level, thus affecting the reliability of subsequent genetic analysis and breeding applications.
[0003] Current genome assembly technologies typically obtain primary assembled sequences from second- or third-generation sequencing data, and then use various long-distance information methods (such as Hi-C, optical maps, and genetic maps) to construct scaffolds and mount chromosomes at the chromosome level. Among these, Hi-C technology can provide chromatin spatial proximity relationships across the entire genome, obtaining paired interactive reads through cross-linking, enzyme digestion, end labeling, and ligation, and forming interaction matrices or heatmaps in bioinformatics processing. Based on the principle that "interactions within the same chromosome decrease with linear distance, and the interaction strength within the same chromosome is significantly higher than that between chromosomes," Hi-C is widely used for scaffold clustering, orientation, and sorting, as well as for identifying and correcting large-scale structural errors (such as misalignment, inversions, translocations, and pseudo-splits). Several streamlined methods and systems have been proposed in Chinese patent literature regarding Hi-C-assisted assembly and correction to improve the automation and accuracy of chromosome-level assembly.
[0004] For example, Chinese patent CN109326323B provides a method and apparatus for assembling a genome. In its background section, it explicitly discusses the characteristics and shortcomings of Hi-C assisted assembly software such as LACHESIS, SALSA, and 3D-DNA: LACHESIS may result in chromosome fusion during chromosome grouping and lacks error correction and whole-genome heatmap evaluation functions; SALSA mainly improves scaffold indicators but struggles to achieve chromosome-level assembly; while 3D-DNA has the ability to correct errors before assembly, its parameters are complex, prone to "over-correction," and chromosome fusion may still occur. Therefore, it proposes to improve universality and accuracy by clustering assembly results that do not meet preset conditions by region and then reassembling them. The technical contribution of this approach lies in introducing a closed-loop approach of "not meeting the criteria—regional clustering—reassembly" into the Hi-C assisted assembly process, attempting to reduce the impact of software bias on the results through rule-based processing. However, from the perspective of error correction judgment logic, the connotation of its "preset conditions" is mainly aimed at the assembly of macro indicators or cluster consistency, and it does not specifically establish a quantifiable and comparable sequence-level quality feedback mechanism for the continuity and structural authenticity of repetitive sequence regions.
[0005] For example, Chinese patent CN110020726A discloses a method and system for sorting assembled sequences. Its core idea is to orient and sort sequences based on the cross-linking signal intensity between the segmented sequences, and to include a "verification unit" in the system structure. This unit verifies the initial arrangement using an interaction heatmap. If the preset verification conditions are not met, the results are adjusted until the target arrangement is obtained. This patent emphasizes a streamlined control process of "cross-linking signal—arrangement result—heatmap verification—adjustment," making the Hi-C heatmap no longer just a display tool, but an automated result verification process. However, its verification still primarily relies on the consistency judgment of interaction signals (two-dimensional heatmaps). It lacks independent sequence quality indicators as verification criteria for whether structural operations such as "merging / splitting / redirecting" bring real improvements at the sequence level, especially whether highly repetitive regions have been restored from "erroneous breaks" to "continuous and biologically sound" structures. Therefore, in noisy or structurally complex genomes, situations may still occur where "the heatmap appears reasonable, but there are still repetitive region breaks or misassemblies at the sequence level."
[0006] For example, Chinese patent CN113122642A discloses a method for assembling and annotating the Hu sheep genome based on third-generation PacBio and Hi-C technologies. This method explicitly includes a step chain of "Hi-C-assisted assembly, constructing an interaction map, and visual error correction," specifically indicating that ALLHiC can be used for Hi-C-assisted assembly, Juicer can be used to construct the interaction map, and JuiceBox can be used for visual error correction. The value of this approach lies in integrating third-generation assembly, error correction, Hi-C mounting, and visual correction into an end-to-end process, emphasizing toolchain synergy to improve chromosome-level assembly quality. However, its error correction and evaluation focus still leans towards "process completeness" and "tool output correctness." The determination of its error correction effectiveness usually still relies on indicators such as N50, anchoring ratio, interaction heatmap morphology, and gene integrity. For repetitive sequences (especially LTR-type retrotransposon-related regions), there is a lack of quantitative criteria strongly tied to "structural operations" that can directly reflect merging / splitting behavior, regarding whether structural error correction significantly improves these sequences.
[0007] In summary, existing technologies have established a relatively mature approach for utilizing Hi-C interaction signals to achieve chromosome-level assembly and structural correction: this involves obtaining the interaction matrix through alignment, performing clustering / sorting / orientation, verifying the results using heatmaps, and making necessary manual or semi-automatic corrections, supplemented by gene integrity or macroscopic continuity indicators to evaluate the assembly effect. However, for complex genomes like those of the giant freshwater prawn, which are rich in repetitive sequences and have controversial karyotypes, relying solely on Hi-C heatmap morphology and general indicators still has at least three limitations:
[0008] 1) Insufficient independent evidence for error correction: Hi-C heatmaps are essentially statistical results of spatial interaction signals, and are affected by library construction noise, enzyme digestion bias, alignment mismatches, and multiple alignments caused by repetitive sequences. When making "merge / split / inversion correction" decisions based solely on heatmap morphology or crosslinking signal intensity, false positives or false negatives are prone to occur, especially when two real chromosomes are close in three-dimensional space, increasing the risk of over-connection. Existing patents mostly use interaction heatmaps as the main verification method (such as the interaction heatmap verification unit in CN110020726A), but lack sequence-level "gold standard" feedback for improving the continuity of repetitive sequences.
[0009] 2) Gap in Quality Assessment of Repetitive Sequence Regions: The Macrobrachium rosenbergii genome has a high content of repetitive sequences, and structural error correction is often first and most significantly reflected in whether repetitive regions are incorrectly broken, incorrectly spliced, and whether they are restored to more continuous transposon structures after merging. Although CN109326323B points out that some tools lack the ability to assess and correct whole-genome heatmaps and proposes improvement routes, its "preset conditions / standard judgments" do not specifically establish comparable quantitative standards around the assembly integrity of repetitive sequences.
[0010] 3) Insufficient closed-loop error correction mechanism in karyotype dispute scenarios: Given the controversy surrounding the number of chromosomes in *Macrobrachium rosenbergii* and the fact that existing assemblies often use a specific karyotype as a priori basis for mounting, structural error correction requires more than just adjusting the heatmap to resemble a diagonal line. It necessitates a closed-loop mechanism that can objectively evaluate whether "merging / splitting is more biologically accurate" without relying on a priori chromosome number. Existing end-to-end process patents (such as CN113122642A) primarily integrate "assembly, error correction, Hi-C mounting, and visualization correction," but lack a targeted, automated, and quantitatively comparable evaluation framework for jointly determining whether "error correction leads to more complete repetitive regions without sacrificing the integrity of conserved genes in the genome."
[0011] Therefore, for the specific application scenario of structural error correction of the Macrobrachium rosenbergii reference genome, there is still an urgent need for a method that can organically combine "Hi-C heatmap signal-driven structural error correction" with "quantitative index feedback for the integrity of repetitive sequence assembly": on the one hand, using Hi-C interaction heatmaps to identify topological errors such as false split sites and inversions and perform merging / splitting / redirecting operations; on the other hand, introducing sequence-level indicators that are highly sensitive to the continuity of repetitive sequence assembly and can be directly compared before and after error correction, forming a reproducible, quantifiable, and interpretable closed loop for judging the effectiveness of error correction, thereby obtaining a more reliable Macrobrachium rosenbergii reference genome in the context of karyotype disputes and high repetition. Summary of the Invention
[0012] The technical objective of this invention is to address the problem that the Macrobrachium rosenbergii genome, due to its rich repetitive sequences and high heterozygosity, is prone to structural errors such as false splits, mislinkings, and inversions during chromosome-level assembly. Furthermore, existing technologies primarily rely on subjective interpretation of Hi-C heatmaps and lack objective sequence-level verification of error correction actions. This invention provides a method that, without pre-setting the number of chromosomes, utilizes Hi-C interaction signals to identify and execute merge / split / redirect error correction. It introduces a joint judgment and rollback mechanism based on global LAI and breakpoint local LAI, combined with BUSCO and consistency thresholds, to quantitatively evaluate the effectiveness of error correction, thereby obtaining a Macrobrachium rosenbergii reference genome with more accurate structure, more continuous repetitive regions, and higher reproducibility.
[0013] Firstly, in order to achieve the above-mentioned objectives, the present invention adopts the following technical solution:
[0014] A method for error correction and quality assessment of the reference genome of Macrobrachium rosenbergii, comprising the following steps:
[0015] S1, Obtain the primary assembly sequence to be corrected and Hi-C sequencing data, wherein the Hi-C sequencing data is obtained by sequencing after restriction endonuclease digestion of the library;
[0016] S2, Align the Hi-C reads to the sequence, construct a genome-wide contact matrix and generate an interactive heatmap;
[0017] S3, on the interactive heatmap, structural error identification is performed on each candidate connection endpoint and candidate structural variation segment;
[0018] Topology error: High signal interaction characteristics and a bow-tie shape exist in the region deviating from the main diagonal, which is used to indicate inversion error;
[0019] S4, For the structural errors identified in step S3, generate and perform at least one structural correction operation to obtain the corrected candidate genome sequence;
[0020] S5. Before and after error correction, identify complete LTR retrotransposons and calculate their lengths and the total length of LTR to calculate the global LAI. Calculate the local LAI in preset windows on both sides, with the breakpoint or corrected segment as the center.
[0021] S6: When the global LAI increases and the corresponding local LAI does not decrease, and the BUSCO integrity does not decrease or is not lower than the allowed decrease threshold, and the sequence consistency outside the error correction segment reaches the preset threshold, the final reference genome is output. Otherwise, the error correction is rolled back and the substitution error correction is performed on the same region, and then S3-S6 are repeated.
[0022] Preferably, the Hi-C sequencing data in step S1 is sequencing data obtained by digesting chromatin with restriction endonucleases and constructing a library. The restriction endonucleases include one or more of MboI, HindIII, and DpnII, and the sequencing platform is a second-generation or third-generation high-throughput sequencing platform.
[0023] Preferably, the structural error in step S3 includes at least the following:
[0024] False split site: There is a continuous diagonal interaction signal at the junction of two adjacent ends labeled as independent sequences, and the signal intensity relative to the background noise reaches a preset judgment threshold.
[0025] Topology error: High signal interaction characteristics and a bow-tie shape exist in the region deviating from the main diagonal, which is used to indicate inversion error;
[0026] As a further preferred option, the preset judgment threshold in step S3 includes: the signal strength of the continuous diagonal interaction signal at the end connection is higher than the multiple threshold of the background noise, and the multiple threshold is at least 3 times.
[0027] As a further preferred option, the identification of the false splitting site in step S3 further includes: performing a neighborhood scan on the end regions of the two independent sequences in the interaction heatmap, and when the diagonal interaction signal in the end neighborhood extends continuously and crosses the boundary of the two sequences, the boundary is identified as a candidate connection endpoint.
[0028] Preferably, the structural correction operation in step S4 includes: performing physical merging on sequences corresponding to false split sites, and / or performing sequence redirection on segments corresponding to inversion errors, and / or performing splitting on erroneous connections; the physical merging includes: performing a merging operation on two candidate sequences in the interactive correction interface of the visual interactive heatmap to generate a merged super scaffold; the redirection includes: performing reverse rearrangement on the target segments identified as inversions to restore their correct orientation in the assembled sequence.
[0029] Preferably, the calculation of LAI in step S5 includes: identifying the complete LTR retrotransposon sequence and the total LTR sequence length across the entire genome, and obtaining the LAI based on their ratio; wherein, the local LAI is calculated and summarized separately within a preset window on both sides of the breakpoint or correction segment.
[0030] Preferably, in step S6, when the global LAI increases and the local LAI does not decrease but the BUSCO decreases beyond the allowed decrease threshold, the structural correction operation is determined to be invalid and a rollback process to undo the structural correction operation is performed.
[0031] Preferably, the method further includes a consistency assessment step: the corrected candidate genome is sequence-aligned with the uncorrected genome and / or the publicly available reference genome, and the final reference genome is output only after verifying that the sequence consistency, excluding the breakpoint or corrected segment, reaches a preset threshold.
[0032] Preferably, the method is used to correct the number of chromosomes in the Giant River Prawn genome from a preset 59 to 57, and to confirm that the global LAI is improved and the BUSCO is not reduced when the final reference genome is output.
[0033] Secondly, the present invention also provides a macrobrachium rosenbergii reference genome structure correction and quality assessment system, which is used to implement the method described above; wherein, the system includes at least: a data acquisition module, an interactive map construction module, a structural error identification module, a structural correction module, and a quality assessment and judgment module; the quality assessment and judgment module is configured to jointly determine the global LAI before and after error correction and the local LAI centered on the breakpoint or correction segment of the structural correction operation, and decide whether to accept the structural correction operation or perform rollback and alternative correction based on the BUSCO threshold condition.
[0034] Thirdly, the present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein the processor executes the computer program to implement the method described therein.
[0035] Fourthly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described thereon.
[0036] This invention, by employing the aforementioned technical solution, achieves high-confidence error correction and quantifiable verification of structural errors in the reference genome of *Macrobrachium rosenbergii*. Its technical effects are manifested in the following ways: First, by utilizing the diagonal continuous signal and bowtie-like interaction features of the Hi-C interaction heatmap, topological errors such as false split sites and inversions can be accurately located without pre-setting the number of chromosomes. Structural correction is then completed through merging, splitting, or redirection operations, reducing misattachment and miscorrection caused by prior karyotype assumptions or algorithmic biases from the source. Second, the introduction of a comparative evaluation between the global LAI and the local LAI of breakpoints / corrected regions ensures that each structural correction operation receives objective feedback regarding the continuity of repetitive sequences, and the local LAI... Without reducing LAI, the global LAI improvement directly represents the repair of repetitive regions such as transposons from "fragmentation / misspelling" to "continuity / completeness," significantly reducing the excessive merging or false positive corrections that may be caused by relying solely on Hi-C heatmaps. Furthermore, by combining the non-reduced BUSCO integrity and the sequence consistency threshold outside the error correction segment, a multi-dimensional constraint is achieved on "improving the continuity of repetitive sequences through error correction without sacrificing gene region integrity and overall consistency." When the conditions are not met, automatic rollback and alternative error correction are performed. Ultimately, a more accurate, more continuous repetitive sequence region, more reliable gene spatial integrity, and stronger reproducibility reference genome of Macrobrachium rosenbergii are obtained, providing more stable benchmark data for subsequent genetic breeding and functional gene research. Attached Figure Description
[0037] Figure 1 A flowchart of the method for correcting errors and assessing the quality of the reference genome structure of the giant freshwater prawn provided by this invention.
[0038] Figure 2 This is a schematic diagram of the signal characteristics of pseudo-split sites in the Hi-C interactive heatmap.
[0039] Figure 3 This is a schematic diagram of the signal characteristics of inverted topology errors in the Hi-C interactive heatmap.
[0040] Figure 4 This is a schematic diagram comparing the LAI values of the genome before and after error correction. Detailed Implementation
[0041] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.
[0042] I. Terminology Explanation
[0043] 1. Primary assembled sequences: These are collections of contig or scaffold sequences obtained by de novo assembly using second- or third-generation sequencing data. They may contain structural defects such as breaks, mislinks, and incorrect orientations, but they possess basic continuity that can be used as a reference for Hi-C alignment.
[0044] 2. Hi-C sequencing data: refers to paired read data obtained through chromatin cross-linking, restriction endonuclease digestion (such as MboI, HindIII, DpnII), end repair markers, nearest neighbor linkage, decross-linking, library construction and high-throughput sequencing. The read pairs reflect the three-dimensional spatial proximity relationship of chromatin.
[0045] 3. Contact Matrix / Interaction Heatmap: After aligning Hi-C reads to the primary assembled sequence, a matrix is formed by counting interactions between different genomic regions according to a preset bin size. The visualization of this matrix is the interaction heatmap. High interaction signals typically correspond to the same linearly adjacent regions near the main diagonal.
[0046] 4. False splitting sites: These are the boundary locations where a physically continuous chromosome (or the same very long scaffold) is mistakenly split into two independent sequences during primary assembly. A typical Hi-C characteristic is the presence of continuous diagonal interaction signals across the boundary.
[0047] 5. Topological Error / Inversion: This refers to a segment in the primary assembly whose orientation is opposite to the true orientation (inversion). In Hi-C heatmaps, this often manifests as a "bow-tie" shaped high-interaction structure that deviates from the main diagonal.
[0048] 6. Structural correction operations: including but not limited to:
[0049] Physical merge: Connects two incorrectly split sequences into a single sequence or a super scaffold by their endpoints;
[0050] Cut / Split: Cuts an incorrectly connected sequence into two parts at the breakpoint;
[0051] Redirect (Inversion / Flip): Reverses the inverted section to restore the correct direction.
[0052] 7. LAI (Long Terminal Repeat Assembly Index): In this invention, a quantitative indicator is used to characterize the continuity and integrity of repetitive sequence (especially LTR retrotransposon) assembly. For ease of engineering implementation, this invention adopts a calculation framework of "the ratio of complete LTR length to total LTR length," which can be used for both whole-genome evaluation and local evaluation around breakpoints.
[0053] 8. Global LAI / Local LAI:
[0054] Global LAI: LAI obtained from genome-wide statistics;
[0055] Local LAI: The LAI obtained by taking a certain error correction breakpoint or correction segment as the center and calculating within a preset window range on both sides, is used to provide action-level feedback on "whether a certain error correction action has truly improved the continuity of the repeating sequence".
[0056] 9. BUSCO Integrity: A metric for assessing the spatial integrity of genes using a set of single-copy orthologous genes. BUSCO does not decrease (or does not fall below the allowable decrease threshold) to constrain the error correction process from introducing significant gene region deletions / duplications.
[0057] 10. Consistency assessment: The candidate genome after error correction is compared with the genome before error correction and / or the publicly available reference genome to verify that the sequence consistency outside the error-corrected region reaches a preset threshold, in order to avoid the error correction operation causing unexpected disturbances to non-target regions.
[0058] 11. Rollback / Alternative Correction: When the joint judgment criteria are not met after correction (such as global LAI not increasing, local LAI decreasing, BUSCO significantly decreasing, or consistency not meeting the standard), the correction operation is revoked (rollback) and another correction strategy (alternative correction) is tried on the same area, forming a closed-loop iteration.
[0059] II. System Structure
[0060] This invention can be implemented as a method or as a software / hardware system. The system structure can be abstracted into the following functional modules (which can be deployed on servers, workstations, or in cloud computing environments), and their collaborative relationships can be referred to... Figure 1 The process shown is as follows:
[0061] Data acquisition module: used to receive primary assembled sequences (FASTA / AGP, etc.) and Hi-C sequencing data (FASTQ), and perform basic quality control (read quality, adapter contamination, repetition rate, effective interaction ratio, etc.).
[0062] Interactive graph construction module: used for Hi-C read segment alignment, filtering invalid interactions, constructing contact matrices and generating multi-resolution interactive heatmap files, supporting subsequent automatic identification and manual verification.
[0063] Structural error identification module: Extracts candidate structural errors based on interactive heatmaps, including false split sites and inverted topological errors; this module supports both rule-based threshold identification and multi-scale consistency verification (signals that are consistent across different bin resolutions are included as candidates).
[0064] The structure correction module generates error correction actions (merging / splitting / redirecting) for candidate errors and generates corrected candidate genomes (including new sequences, AGP, breakpoint records, and coordinate mappings) under version management.
[0065] Quality assessment and judgment module: Calculates the global LAI and local LAI before and after error correction, and performs joint judgment based on BUSCO and consistency assessment results; when the conditions are not met, it triggers rollback and alternative error correction, and records the reasons, threshold triggers and next steps.
[0066] Results output module: Outputs the final reference genome sequence, chromosome / super scaffolding list, error correction report (breakpoint list, action type, changes before and after LAI / BUSCO, consistency indicators, etc.), and can generate Figure 4 The LAI comparison visualization results are shown below.
[0067] The above modules can be integrated into a standalone software or executed in series in a scheduling system as a pipeline; the inputs and outputs of each step adopt the common format in the field, ensuring that those skilled in the art can implement it without creative effort.
[0068] III. Specific Technical Route for Implementing the Method of the Invention
[0069] Reference Figure 1 The overall route of this invention is as follows: S1 Data acquisition → S2 Interactive map construction → S3 Structural error identification → S4 Structural correction → S5 LAI (global + local) quality assessment → S6 Joint judgment (LAI + BUSCO + consistency) and rollback iteration → Output the final reference genome.
[0070] (I) Step S1 Data Acquisition
[0071] 1. Sample and DNA preparation
[0072] Select healthy, vigorous adult giant freshwater prawns, and harvest muscle or gill tissues, which are then flash-frozen in liquid nitrogen and... Preservation. Genomic DNA should be extracted using a commercially available high molecular weight DNA extraction kit or a modified CTAB method. For high-quality third-generation sequencing and Hi-C library preparation, it is recommended to control DNA integrity (e.g., main band >50kb on electrophoresis) and purity (…). Approximately 1.8 >2.0) and the concentration meets the requirements for database construction.
[0073] 2) Acquisition of primary assembly sequence
[0074] Long read data can be obtained using third-generation sequencing platforms (such as PacBio or Nanopore), and contig-level assemblies can be generated using conventional assemblers; second-generation data can be used for polishing if necessary. The output is a primary assembly FASTA file and its statistical information (total length, N50, number of fragments, N content, etc.). This invention is not limited to specific assembly software; the key is to output primary assembled sequences that can be aligned with Hi-C.
[0075] 3) Hi-C sequencing data acquisition
[0076] The Hi-C standard workflow is as follows: tissue fixation and cross-linking → chromatin digestion → end repair / labeling → spatial nearest neighbor ligation → decross-linking → DNA purification → library construction and sequencing. Restriction endonucleases can be one or more of MboI, HindIII, and DpnII. The sequencing platform can be a second- or third-generation high-throughput platform, yielding paired-end or paired-read FASTQ files. To ensure the reliability of the subsequent contact matrix, the effective interaction ratio (effective paired alignments, effective logarithms after removing self-loops, duplicates, and invalid ligations of the same fragment) can be calculated during the data acquisition stage, and low-quality samples can be retested or removed.
[0077] 4) Input organization and version management
[0078] Establish the project directory structure as follows: raw_data / , assembly_v0 / , hic_map / , candidate_fixes / , assembly_v1...vn / , reports / , etc. Set the initial assembly to version. Each error correction results in a new version. Record the breakpoint coordinates and reasons for each error correction action (for rollback and auditing purposes).
[0079] (II) Step S2 Interactive Graph Construction
[0080] 1. Hi-C read segment comparison and filtering
[0081] The Hi-C reads are aligned to the primary assembled sequences. Alignment tools are not limited; commonly used procedures in the field can be employed: BWA-MEM, etc., can be used for second-generation data, and minimap2, etc., can be used for third-generation data. After alignment, Hi-C-specific filtering is performed: PCR duplicates, invalid loops, ligation of the same fragment, low-quality alignments, and unreliable pairs resulting from multiple alignments are removed, retaining only high-confidence pairs for matrix construction.
[0082] 2. Contact Matrix and Multi-resolution Heatmap Generation
[0083] Interaction counts are statistically analyzed using preset bin sizes (e.g., 50kb, 100kb, 250kb, 1Mb, etc.) to form a contact matrix. Appropriate normalization is then performed (e.g., row and column balancing or coverage normalization) to reduce the impact of uneven sequencing depth. An interactive heatmap file is output for automatic identification and manual verification.
[0084] 3. Background noise estimation
[0085] To achieve the "intensity ≥ background noise threshold (e.g., 3 times)" determination in step S3, it is necessary to define the calculation method for background noise: several matrix cells outside the neighborhood of the breakpoint and at the same scale can be sampled and counted, and their mean / median can be taken as the background estimate. This estimate will be used for subsequent ratio determination and robustness control.
[0086] (III) Step S3 Structural Error Identification
[0087] Step S3 is the entry point for the closed loop of this invention. Its goal is to output a candidate structural error list, including false split sites and inversion errors, by combining the calculable signal features of the interactive heatmap without presetting the number of chromosomes, and to provide a basis for generating subsequent error correction actions. This implementation supports both fully automatic rule recognition and a semi-automatic mode of automatic candidate selection and manual review and confirmation; neither changes the core innovation of this invention, which uses local LAI action-level feedback as the error correction acceptance condition.
[0088] 1. Candidate Region Generation and Multi-Scale Consistency Strategy
[0089] (1) Candidate endpoint set
[0090] For each scaffold / pseudo-chromosome sequence, extract its terminal regions (e.g., each terminal region). bp range This can be in the Mb range or defined by bin number, forming an endpoint candidate set. The endpoint candidate set is used to identify false split sites where there is a continuous diagonal interaction between the ends of two independent sequences.
[0091] (2) Candidate segment set
[0092] A sliding window scan is performed within each sequence to look for segments that may contain inversions or misconnections; priority is given to regions that show obvious diagonal breaks, abnormal local block structures, or strong interactions deviating from the diagonal in the heatmap. This candidate set is used for inversion and splitting error identification.
[0093] (3) Confirmation of multi-resolution consistency
[0094] Considering that Hi-C signals exhibit different details at different bin scales, this invention preferably uses "at least two resolutions being consistent" as the confirmation criterion for candidate errors. For example, if the same type of feature (continuous diagonal lines or bows) is observed at both 100kb and 250kb resolutions, it is then included in the candidate list to reduce false positives caused by noise.
[0095] 2. Quantitative implementation of pseudo-split site identification (e.g.) Figure 2 (As shown)
[0096] The core characteristic of pseudo-separation is that the connection between the ends of two sequences labeled as independent still exhibits continuous diagonal interaction signals, with a significantly higher intensity than the background. To enable direct implementation by those skilled in the art, an operable quantitative determination method is presented.
[0097] (1) Definition of the neighborhood of the breakpoint
[0098] Let the two candidate sequences be respectively and Its terminal candidate connection position is and At a given bin resolution, take End inward bin, End inward Each bin constitutes a matrix sub-block. It is used to observe cross-boundary interactions.
[0099] (2) Signal strength along the continuous diagonal
[0100] Define a set of "quasi-diagonal" cells across boundaries. (For example from) Calculate the average interaction count of the main diagonal elements from the top left to the bottom right as the continuous diagonal intensity:
[0101] ;
[0102] in, The first in the contact matrix Line number Column interaction count; A set of cross-boundary quasi-diagonal elements; The number of elements in the set.
[0103] (3) Background noise intensity
[0104] Exclude within the same sub-block or its adjacent areas. After several bins nearby, take the remaining unit set. Calculate the background interaction mean:
[0105] ;
[0106] in, This is the set of background sampling units.
[0107] (4) Signal-to-noise ratio determination
[0108] The ratio of the continuous diagonal signal to the background is defined as:
[0109] ;
[0110] in, This is the ratio for determining a false split. When When this happens, it is determined to be a false split candidate; This is a preset threshold. The present invention preferably uses... (That is, the intensity is at least 3 times that of the background).
[0111] (5) Robustness check
[0112] To avoid false signals caused by multiple alignments due to repeated sequences, the following robustness check can be added (without changing the core of the invention): repeated calculations at adjacent different resolutions. All endpoints must meet the threshold; check the alignment quality distribution and effective interaction ratio near the endpoints, and remove obviously abnormal regions; if multiple candidate endpoint combinations exist, select the appropriate endpoints according to the criteria. Sort the sequences from highest to lowest intensity and proceed to S4 to generate error correction. When two independent sequences have a continuous diagonal line across the boundary at the end junction and the intensity is significantly higher than the background, it indicates that the two sequences are more likely to belong to the same continuous chromosome and should be considered as merging candidates.
[0113] 3. Quantization Implementation of Inverted Topology Error Detection
[0114] like Figure 3 As shown, inversions commonly exhibit a "bow tie" structure in Hi-C heatmaps. Essentially, near the inversion boundary, the interacting signals show symmetrical enhancement on both sides of the diagonal. For feasibility, this invention provides a quadrant statistics-based determination method.
[0115] (1) Inverted candidate boundary and submatrix
[0116] Let the candidate inversion segment be , with the boundary and Extract neighborhood submatrices centered on the point, or extract segments from the midpoint of the segments. Cut off at the center submatrix .
[0117] (2) Quadrant interaction intensity statistics
[0118] Will Divide the space into four quadrants based on the center: And calculate the average interaction strength in each quadrant. :
[0119] ;
[0120] in, For the first Quadrant unit set.
[0121] (3) Bow tie scoring
[0122] In a typical inverted case, the crossing quadrants on both sides of the diagonal (e.g.) and This will be significantly enhanced. It can be defined as:
[0123] ;
[0124] in, This is the inverted characteristic ratio. If If the condition holds true at least in two different resolutions, then the segment is considered a candidate for inversion. The preset threshold can be set according to the data quality (e.g.) (or higher).
[0125] (4) with Figure 3 Corresponding morphological verification
[0126] In semi-automatic mode, candidate segments can be selected through a visual interface. Figure 3 The "bow tie" shape shown is used for verification to further reduce false positives. This verification does not change the objectivity of the closed-loop determination in this invention, because the final decision on whether to accept the error correction still needs to be made in steps S5-S6 by the joint determination of local LAI, BUSCO, and consistency.
[0127] (iv) Step S4 Structural Correction
[0128] The goal of step S4 is to convert the candidate structural errors output in step S3 into executable error correction actions and generate a corrected candidate genome version. To ensure repeatability and rollback capability, this implementation emphasizes: clear action granularity, traceable coordinates, and version rollback capability.
[0129] 1. Action generation strategy
[0130] (1) Merging action for pseudo-split sites
[0131] For each pseudo-split candidate Generate merge action .in Indicates the direction of connection (e.g.) tail splice (Head), can be determined by combining Hi-C diagonal continuity scores in different directions; if the direction is uncertain, two candidate actions can be generated to enter step S5-step S6 for evaluation, and the better one is automatically selected by local LAI and consistency constraints.
[0132] (2) Redirection action for inverted candidates
[0133] Candidate segments for inversion Generate redirection action This involves reversing the segment and updating the coordinate mapping. If there is uncertainty regarding the inversion boundary, a small-range sliding motion can be performed near the boundary (e.g., Multiple candidate segments are generated from each bin, and the optimal segment is selected through a closed-loop evaluation process.
[0134] (3) Splitting action for faulty connections
[0135] When a clear diagonal break appears in the heatmap, and the interaction across the break points is lower than the background or exhibits an abnormal block structure, a splitting action is generated. ,in These are the coordinates of the breakpoint. Splitting actions are often used as an "alternative strategy after error correction failure": for example, if a merging action causes a local LAI decrease or inconsistency to be unacceptable, then a rollback is performed and an attempt is made to split the suspected faulty connection.
[0136] 2. Error correction execution and version management
[0137] (1) Execution layer implementation
[0138] Error correction can be performed by editing FASTA / AGP using scaffolding editing tools, exporting results from the Hi-C error correction visualization tool, or using self-developed scripts. The key is that a fixed length can be introduced at the connection point during merging. The choice between gaps or gapless splicing depends on whether there is explicit sequence overlap; when inverting, the segment sequence needs to be reversed and its direction marker in AGP needs to be updated; when splitting, the sequence needs to be split at the breakpoint and a new sequence ID or a new scaffold record needs to be generated.
[0139] (2) Coordinate mapping and audit records
[0140] Each action generates an "Action Log File," which includes: old version ID, new version ID, action type, breakpoint / segment coordinates, involved sequence ID, direction information, expected error type to be corrected, and corresponding S3 score (e.g., ...). or The evaluation results of S5–S6 are also generated. Simultaneously, coordinate mapping (lift-over) information is generated to facilitate the mapping of gene annotation, duplicate annotation, and other results from the old version to the new version.
[0141] (3) Rollback mechanism
[0142] Rollback can be achieved through a versioned file system: retain each version of FASTA / AGP and its index; if step S6 fails, directly restore to the previous successful or baseline version, and perform alternative error correction actions on this basis.
[0143] (V) Step S5LAI Quality Assessment
[0144] Step S5 is one of the key steps that distinguishes this invention from "correction solely based on Hi-C heatmaps". Its core idea is to establish a direct correspondence between the error correction action and the quantitative improvement of the continuity of the repetitive sequence, especially by providing action-level feedback for a single error correction action through "local LAI of the breakpoint / correction segment", thereby suppressing excessive merging and false positive errors.
[0145] 1. LTR Recognition and LAI Calculation Process
[0146] (1) Recognition of intact LTR retrotransposons
[0147] For the previous version With the corrected candidate version Run the LTR identification process separately (a commonly used toolchain in this field can be used, such as calling LTR_FINDER and LTR_HARVEST for candidate identification, and then using LTR_retriever to integrate and filter to obtain a high-confidence complete LTR-RT set). The output includes: a complete LTR-RT list, statistics on the total length of LTR-related sequences, etc.
[0148] (2) Global LAI calculation
[0149] According to the engineering definition adopted in this invention, the "total length of the complete LTR-RT sequence" across the entire genome is defined as... "Total LTR sequence length (including complete and incomplete candidates)" is Then the global LAI is defined as:
[0150] ;
[0151] in, For global LAI; The sum of the complete LTR-RT lengths; This is the sum of the total LTR length. If you need to express it as a percentage, you can multiply it by 100%, but this will not affect the "increase / decrease" comparison logic.
[0152] (3) Global LAI change
[0153] Calculate the changes before and after the bug fix:
[0154] ;
[0155] in, For global LAI before error correction, This is the global LAI after error correction.
[0156] 2. Action level calculation of local LAI
[0157] (1) Local window definition
[0158] For each error correction action Determine the center coordinates of the corresponding breakpoint or correction section. (e.g., merge breakpoints, split breakpoints, inverted segment centers). Preset windows are selected on both sides of the genome linear coordinate system. (For example Desirable (on the order of Mb, or defined by bin number), forming a local interval. This invention is not limited. The specific value, The "preset window range" can be selected by the implementer based on the genome size and data resolution.
[0159] (2) Local LTR statistics and local LAI
[0160] Statistics within a local interval and And calculate:
[0161] ;
[0162] in, For local LAI; The length of the complete LTR-RT within the local window; This represents the total length of the LTR within the local window.
[0163] (3) Local LAI variation
[0164] For the same action Calculations before and after error correction:
[0165] ;
[0166] in, and These are the local LAIs defined in the same window before and after error correction.
[0167] like Figure 4 As shown, the global LAI or multiple local window LAIs before and after error correction can be presented as box plots or distribution plots. In terms of implementation: the global LAI can be used for single-value comparison; the local LAI can be sampled from all error correction action windows to form a set of sample distributions, used to show whether "the error correction action as a whole improves the continuity of the repetitive sequence"; or the local LAI can be drawn separately for key merging breakpoints (such as the two merging points involved in correcting 59 to 57) as core evidence.
[0168] (vi) Step S6: Joint judgment, rollback and alternative error correction
[0169] Step S6 is the "decision layer" of the closed loop of this invention. Its key lies not in simply calculating indicators, but in using the global LAI + local LAI as the acceptance condition for structural error correction, and using BUSCO and consistency constraints to prevent "introducing erroneous splicing in order to improve LAI", ultimately achieving dual verification of "structure-sequence".
[0170] 1. Implementation of the joint decision criterion
[0171] (1) LAI criteria (necessary conditions)
[0172] For each error correction action The generated candidate versions must at least satisfy:
[0173] (Global LAI increases);
[0174] (This action does not reduce the local LAI).
[0175] This requires that error correction actions must at least not cause local degradation in the repeating sequence space, and bring overall improvement, making it more difficult to be fooled by "only good-looking heatmaps but more fragmented repeating regions" or "excessive merging".
[0176] (2) BUSCO constraints (necessary conditions or strong preference conditions)
[0177] Calculate the BUSCO integrity before and after error correction. Let the integrity before error correction be... After correction, it is The allowed drop threshold is (It can be set to 0 to indicate that a decrease is not allowed, or set to allow a small decrease), then the requirements are:
[0178] ;
[0179] in, You can select the proportion or quantity of Complete BUSCOs; The threshold for allowing a decrease.
[0180] (3) Consistency constraints (preferred conditions, used to stabilize non-target regions)
[0181] Sequence alignment is performed outside the error-correction region, with a consistency metric of . (For example, comparing a combination of consistency rate or coverage rate), the threshold is ,Require: ;in, Consistency indicators outside the error correction area; This is a preset threshold. (Under implementation) It can be determined by the "compare consistency rate" The structure can be constructed using methods such as "coverage rate" or "high consistency segment ratio," and technicians can select the appropriate method according to project requirements.
[0182] If the above combined conditions are met, the error correction action is deemed valid, the candidate version is upgraded to a "passed version", and the next candidate error is processed or the final reference genome is output.
[0183] 2. Rollback and alternative error correction strategies when failure occurs
[0184] If any condition is not met, the system executes:
[0185] (1) Rollback: Undo the error correction action and restore the version before the error correction. ;
[0186] (2) Cause identification: Record the conditions that caused the failure to trigger (e.g. BUSCO drop exceeds threshold (etc.), used to guide alternative strategies;
[0187] (3) Alternative error correction: Evaluate the same candidate region by taking different actions or fine-tuning breakpoints / boundaries. Common alternative strategies include: Merging failure (local LAI decrease or inconsistency) → switching to splitting or adjusting connection direction; inaccurate inversion boundaries (local LAI decrease or inconsistency). Unstable) → Sliding boundary bin re-executes redirection; global LAI does not increase after splitting → prioritizes returning to S3 to re-filter higher confidence candidates to avoid excessive fragmentation.
[0188] This strategy transforms error correction from a "single subjective decision" to a "quantifiable multi-round trial-verification-rollback," significantly improving reproducibility and noise resistance.
[0189] V. Application Example: Validation of Macrobrachium rosenbergii Reference Genome Structure Correction and Quality Improvement Based on "Hi-C Structure Evidence, LAI (Global / Local) Feedback, and BUSCO Constraints"
[0190] This application example focuses on constructing a reference genome for the giant freshwater prawn. Following the technical solution described in this invention, without relying on a fixed karyotype prior for error correction, it identifies, merges / redirects, and corrects structural errors in primary chromosome-level assembly. Furthermore, it uses a joint assessment based on LAI (Global and Local Breakpoint Analysis) and BUSCO integrity to demonstrate that this invention can significantly improve the assembly continuity and structural accuracy of repetitive sequence regions while avoiding "over-assembly." The overall workflow corresponds to... Figure 1 The key structural evidence corresponds to Figure 2 , Figure 3 The corresponding quality improvement results Figure 4 .
[0191] (a) Experimental subjects and data sources (corresponding to steps S1–S2)
[0192] (1) Sample acquisition and preservation
[0193] Select healthy, vigorous adult giant freshwater prawns, quickly extract approximately 1g of muscle tissue, flash-freeze in liquid nitrogen, and then place... Preservation. This sampling method minimizes DNA degradation and microbial contamination, meeting the requirements of third-generation sequencing and Hi-C library preparation for high molecular weight DNA.
[0194] (2) Acquisition of primary assembly sequence
[0195] High-quality genomic DNA was extracted using commercially available kits, a portion of which was used for third-generation sequencing (PacBio or Oxford Nanopore) to obtain long read data. This data was then de novo assembled into primary assembled sequences (contig / scaffold level), serving as a reference for subsequent Hi-C alignment and as the basis for chromosome-level mounting. This example does not limit the specific assembly software, but requires the output of primary assembled sequences in FASTA format with version numbers retained (e.g., denoted as...). ).
[0196] (3) Hi-C library construction and sequencing
[0197] A separate muscle tissue sample was used for Hi-C library construction. Library construction followed a standard procedure of restriction endonuclease digestion, end labeling, nearest-neighbor ligation, decrosslinking, and purification. In this example, MboI was used as the restriction endonuclease. It should be noted that those skilled in the art can substitute enzymes such as HindIII or DpnII that recognize 4bp or 6bp palindromic sequences, or use a method that does not depend on specific restriction sites (such as the Arima Genomics kit), all of which will yield Hi-C interaction data that satisfy the requirements of this invention. The Hi-C library was subjected to paired-end sequencing on platforms such as DNBseq or Illumin aNovaSeq, yielding approximately 100–200 Gb of raw data.
[0198] (4) Interaction map construction and preliminary chromosome-level assembly
[0199] The raw Hi-C data was quality controlled, aligned, and filtered for effective interactions using Juicer (v1.6) to construct a contact matrix and generate an interaction heatmap. Using third-generation primary assembly sequences as a reference, an automated scaffold was constructed using the 3D-DNA (v180922) workflow to obtain a preliminary chromosome-level assembly version. Subsequently, Juicebox (v1.11.08) was used to visualize the interaction heatmap, providing a basis for structural error identification and correction.
[0200] Key outputs of this step: A primary chromosome-level assembly version (initially assumed to consist of 59 pseudochromosomes / linkage groups, denoted as the "59-chromosome version"); the corresponding Hi-C interaction heatmap (used to identify structural errors such as pseudo-splits and inversions); and breakpoint candidates and interaction signal evidence for subsequent action-level verification (see...). Figure 2 , Figure 3 ).
[0201] (II) Structural error identification and correction implementation (corresponding to steps S3–S4)
[0202] A systematic review of the interactive heatmap of the "59 versions" revealed at least two typical structural errors: false splitting sites and inverted topological errors. According to this invention, error correction is not aimed at "the prior must be 59 versions," but rather at locating errors using Hi-C structural evidence, performing structural correction, and then objectively judging the error correction action using indicators such as LAI and BUSCO.
[0203] 1. Identification and Merging Error Correction of False Split Sites
[0204] A significant diagonal continuous interaction signal was observed at the junction of the original assembled Chr_X (temporary number) and Chr_Y (temporary number). Figure 2 (Arrow), and the signal intensity is more than 3 times higher than the surrounding background noise. This phenomenon conforms to the identification rule of the present invention for "false split sites": if there is still a continuous diagonal interaction across the boundary at the end connection of two sequences labeled as independent sequences, it indicates that the two sequences are more likely to belong to the same continuous chromosome in three-dimensional space and linear structure, and the original assembly splitting it is a structural error.
[0205] Error correction action: Perform a physical merge operation on Chr_X and Chr_Y in Juicebox to generate a merged super scaffold and update the assembly version (e.g., by...). Updated to Record the coordinates, direction, and operation log of the merge breakpoint. This action corresponds to "performing physical merging of sequences corresponding to pseudo-split sites" in this invention.
[0206] 2. Inverted topology error identification and redirection error correction
[0207] In the Chr_Z region interaction heatmap, strong interaction signals deviating from the main diagonal were identified, exhibiting a typical "bow-tie pattern." Figure 3 (Arrow). This pattern is a classic Hi-C feature of large segment inversion: inversion causes the interaction signals to be symmetrically enhanced on both sides of the diagonal, thus forming a "bow tie" pattern, indicating that the direction of this segment is inconsistent with the true direction.
[0208] Error correction action: Perform an inversion operation on the region to restore it to the correct direction and update the assembled version (e.g., by...). Updated to This records the inverted segment boundary, reverse operation, and version differences. This action corresponds to "redirecting the segment corresponding to the inverted error" in this invention.
[0209] 3. Correction results: The 59 versions have been consolidated into 57 super scaffolding versions.
[0210] By systematically investigating and correcting the two types of errors mentioned above (including but not limited to merging and inversion redirection in the examples), the sequences that were originally split into 59 linkage groups were reintegrated, ultimately resulting in a corrected candidate genome composed of 57 super-scaffolds (referred to as the "57 versions"). It is important to emphasize that this "59→57" change is not a pre-set target, but rather a natural result after the structural errors have been corrected. Further analysis using the LAI and BUSCO criteria of this invention is still necessary to rule out the possibility of over-assembly.
[0211] (III) Verification of the effectiveness of error correction based on LAI (global, local critical breakpoints) and BUSCO (corresponding to S5–S6)
[0212] To verify whether the above structural corrections (especially the "merging" action) constitute "over-assembly," this example introduces LAI as quantitative evidence of the continuity of repetitive sequences and combines it with BUSCO as a constraint on gene spatial integrity to jointly evaluate the versions before and after correction, forming a reproducible chain of evidence.
[0213] 1. Comparison of LAI calculation methods and global methods
[0214] The LTR_retriever (v2.9.0) workflow, combined with tools such as LTR_FINDER and LTR_HARVEST, is used to identify complete LTR retrotransposon sequences across the entire genome and calculate the LAI. For ease of engineering implementation, this example employs the following computational framework:
[0215] ;
[0216] in, The total length of complete LTR retrotransposon sequences identified across the entire genome. The total length of LTR-related sequences (including complete and incomplete LTR candidate sequences).
[0217] Global LAI results: such as Figure 4 As shown, the average LAI before correction (59 versions) was 12.5; after correction (57 versions), the average LAI increased to 14.2. The significant increase in global LAI indicates that the corrected versions have an overall improvement in the continuity and integrity of the repetitive sequence space (especially the LTR-related region), rather than just "looking more like chromosomes" in terms of macroscopic structure.
[0218] 2. Action-level verification of local LAI at key breakpoints (evidence of merged sites)
[0219] This invention emphasizes "action-level objective verification": not only comparing the global LAI, but also focusing on the key sites where the merging operation was performed in Example 2 to verify whether the merging truly repaired the breakage of the repeat sequence.
[0220] For the region surrounding the Chr_X and Chr_Y connection sites (i.e., within a pre-defined window on both sides of the merge breakpoint), the local LAI was calculated and the changes before and after error correction were compared. The results showed that the LAI in this local region increased from less than 10 (Draft level) before merging to over 20 (Reference level). This evidence is directly indicative: if merging is excessive splicing, it often introduces erroneous connections and disrupts the local repetitive sequence structure, making it difficult to achieve the result of "significantly increased intact LTR and significantly increased local LAI" in the breakpoint neighborhood. Conversely, this example observed a significant increase in local LAI, indicating that the merging action repaired the transposon structure that was previously broken due to erroneous splitting, allowing the repetitive sequence fragments to be assembled more completely into the same continuous structure, supporting at the sequence level that "merging is a biologically real error correction."
[0221] 3. BUSCO integrity assessment (error correction without sacrificing genomic spatial integrity)
[0222] To rule out the possibility of sacrificing gene region correctness to improve repetitive region metrics, this example further employed BUSCO (v5.4.3) and selected the arthropoda_odb10 gene set to assess the integrity of single-copy orthologous genes. The results showed that the Complete BUSCOs of the corrected genome were 96.5%, which was not lower than before correction, and no gene loss or duplication anomalies were observed due to incorrect splicing. This result demonstrates that structural error correction not only improves the continuity of repetitive sequence regions (reflected by LAI) but also maintains the integrity and reliability of the gene space.
[0223] Based on the above implementation process and data results, the following conclusions can be drawn, which prove the technical effects of the present invention:
[0224] 1) Effective and interpretable structural error correction: False splits are identified through Hi-C interactive heatmaps ( Figure 2 (Diagonal continuous signal) and inversion ( Figure 3 (Bow Knot Signal), and implement merging and redirection error correction, which can systematically repair the structural errors existing in the initial 59 versions, resulting in 57 more reasonable super scaffolding candidate assemblies.
[0225] 2) LAI provides objective "error-correcting feedback" to the repeating sequence space: after error correction, the global LAI increased from 12.5 to 14.2. Figure 4 This indicates an overall improvement in the continuity of the repeating sequence; more importantly, the local LAI around the merge breakpoint jumped from <10 to >20, proving that the merging action actually repaired the break in repeating sequences such as LTR in the neighborhood of the breakpoint, rather than causing excessive splicing. This "global + local" chain of evidence makes the error correction conclusion quantifiable and verifiable.
[0226] 3) BUSCO constraints ensure that the gene space is not damaged: After error correction, the Complete BUSCOs is 96.5% and does not decrease, indicating that structural error correction does not come at the cost of gene integrity, thus eliminating the risk of gene deletion or duplication caused by incorrect splicing.
[0227] In summary, this application example demonstrates that the present invention can achieve a more accurate, more continuous repetitive sequence region, and eliminate over-assembly reference genome construction effect in the genome of Macrobrachium rosenbergii, which has controversial karyotype and abundant repetitive sequences. This improves the quality of the genome to a higher standard and provides reliable benchmark data for subsequent molecular breeding and functional studies.
[0228] The foregoing description of embodiments of the present invention, through which those skilled in the art are able to implement or use the present invention, will be readily apparent to those skilled in the art. Various modifications to these embodiments will be readily apparent to those skilled in the art. The general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novelty disclosed herein.
Claims
1. A method for error correction and quality assessment of the reference genome structure of *Macrobrachium rosenbergii*, characterized in that, The method includes the following steps: S1, obtaining the primary assembly sequence to be corrected and Hi-C sequencing data, wherein the Hi-C sequencing data is obtained by sequencing after restriction endonuclease digestion of the library; S2, aligning the Hi-C reads to the sequence, constructing a whole-genome contact matrix and generating an interaction heatmap; S3, identifying structural errors in each candidate connection endpoint and candidate structural variation segment on the interaction heatmap; S4, generating and performing at least one structural correction operation on the structural errors identified in step S3 to obtain the corrected candidate genome sequence; S5: Before and after error correction, identify complete LTR retrotransposons and calculate their lengths and the total LTR length to calculate the global LAI. Calculate the local LAI in preset windows on both sides of the breakpoint or corrected region of each error correction. S6: When the global LAI increases and the corresponding local LAI does not decrease, and the BUSCO integrity does not decrease or is not lower than the allowable decrease threshold, and the sequence consistency outside the error-corrected region reaches the preset threshold, output the final reference genome. Otherwise, roll back the error correction and perform substitution error correction on the same region, then repeat S3-S6. In step S5, the local LAI is calculated as follows: For each error correction action, determine the corresponding breakpoint or the coordinates c of the center of the correction segment; take two pre-defined windows W on the genomic linear coordinates to form a local interval. Statistics within a local interval and And calculate: ; in, For local LAI; The length of the complete LTR-RT within the local window; This represents the total length of the LTR within the local window. In step S6, the BUSCO constraints are as follows: Calculate the BUSCO integrity before and after error correction. Let the integrity before error correction be... After correction, it is The allowed drop threshold is Then the following is required: Sequence alignment is performed outside the error-correction region, with a consistency index of 1. The preset threshold is ,Require: If the above combined conditions are met, the error correction action is deemed valid, the candidate version is upgraded to a pass version, and the next candidate error is processed or the final reference genome is output. The alternative error correction involves taking different actions or fine-tuning breakpoints / boundaries for the same candidate region before re-evaluation. Alternative strategies include: if merging fails, splitting or adjusting the connection direction; if the inverted boundary is inaccurate, sliding the boundary. bin will re-execute redirection; if the global LAI does not increase after splitting, it will prioritize returning to S3 to re-filter higher confidence candidates.
2. The method of claim 1, wherein, The Hi-C sequencing data mentioned in step S1 is sequencing data obtained by digesting chromatin with restriction endonucleases and constructing libraries. The restriction endonucleases include one or more of MboI, HindIII, and DpnII, and the sequencing platform is a second-generation or third-generation high-throughput sequencing platform.
3. The method according to claim 1, characterized in that, In step S3, the structural errors include at least: spurious split points: continuous diagonal interaction signals exist at the connection points of adjacent ends of two sequences labeled as independent, and the signal strength relative to the background noise reaches a preset judgment threshold; topological errors: high signal interaction features exist in the region deviating from the main diagonal and are bow-shaped, used to indicate inversion errors; the preset judgment threshold includes: the signal strength of the continuous diagonal interaction signals at the connection points is higher than a multiple threshold of the background noise, and the multiple threshold is at least 3 times; and / or, the identification of spurious split points in step S3 includes: performing a neighborhood scan on the end regions of the two independent sequences in the interaction heatmap, and when the diagonal interaction signals in the end neighborhood extend continuously and cross the boundary of the two sequences, the boundary is determined as a candidate connection endpoint.
4. The method according to claim 1, characterized in that, The structural correction operation in step S4 includes: performing physical merging on sequences corresponding to false split sites, and / or performing sequence redirection on segments corresponding to inversion errors, and / or performing splitting on erroneous connections; the physical merging includes: performing a merging operation on two candidate sequences in the interactive correction interface of the visual interactive heatmap to generate a merged super scaffold; the redirection includes: performing reverse rearrangement on the target segments identified as inversions to restore their correct orientation in the assembled sequence.
5. The method of claim 1, wherein, The calculation of LAI in step S5 includes: identifying the complete LTR retrotransposon sequence and the total LTR sequence length across the entire genome, and obtaining the LAI based on their ratio; wherein, the local LAI is calculated and summarized separately within a preset window on both sides of the breakpoint or correction segment.
6. The method according to claim 1, characterized in that, The method further includes a consistency assessment step: the corrected candidate genome is sequence-aligned with the uncorrected genome and / or the publicly available reference genome, and the final reference genome is output only after verifying that the sequence consistency, excluding the breakpoint or corrected segment, reaches a preset threshold.
7. The method according to claim 1, characterized in that, The method is used to correct the number of chromosomes in the Giant River Prawn genome from a preset 59 to 57, and to confirm that the global LAI is improved and the BUSCO is not reduced when the final reference genome is output.
8. A Macrobrachium rosenbergii reference genome structure correction and quality assessment system, characterized in that, The system is used to implement the method described in any one of claims 1 to 7; wherein the system includes at least: a data acquisition module, an interactive graph construction module, a structural error identification module, a structural correction module, and a quality assessment and judgment module; the quality assessment and judgment module is configured to jointly determine the global LAI before and after error correction and the local LAI centered on the breakpoint or correction section of the structural correction operation, and decide whether to accept the structural correction operation or perform rollback and alternative correction based on the BUSCO threshold condition.
9. An electronic device comprising a processor, a memory, and a computer program stored in the memory and executable by the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
A method and apparatus for assembling a genome
CN109326323B
Method and system for sorting assembly sequences
CN110020726A
Method for assembling and annotating Hu sheep genome based on three-generation PacBio and Hi-C technologies
CN113122642A
A method for analysis and interpretation of crop bioinformatics repeats sequence pattern
AU2021101343A4
Reference-level high-quality genome assembling method
CN120260685A