A Hi-C-based microbial metagenomic sequencing analysis method and system
Through an automated analysis process based on Hi-C technology, the problems of distinguishing MAGs and defining host relationships in metagenomic sequencing are solved, and efficient and convenient metagenomic research and analysis of complex environmental samples are achieved, especially the discovery of new species and the study of host source of antibiotic resistance genes.
Patent Information
- Application Number
- CN202411022278.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2044-07-29
AI Technical Summary
The existing metagenomic sequencing technology is difficult to efficiently and conveniently distinguish and analyze MAGs, especially for single bacterial genome research in complex microbiome samples. The existing methods are difficult to distinguish different species with less obvious genome sequence characteristics, and lack the definition of the host relationship between microorganisms, plasmids and viruses.
Using microbial metagenome sequencing analysis methods and systems based on Hi-C technology, Nextflow and docker are used to build an automated analysis process, and combining software Flye, Pilon, bin3C, checkM, gtdb-tk, prokka, etc., quality control, assembly, comparison, clustering and annotation are carried out, Hi-C interaction signals are identified, connection maps are constructed, and host relationships between MAGs, plasmids and viruses are determined.
It has achieved that non-biological information professionals can conduct metagenomic research efficiently and with high quality, discover new species in complex environmental samples, define the host relationship between microorganisms and plasmids and viruses, and improve the quality and analysis efficiency of MAGs.
Smart Images

Figure CN118782149B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of metagenomic sequencing technology, and in particular to a Hi-C-based microbial metagenomic sequencing analysis method and system. Background Art
[0002] The number of microbial individuals accessible using culture-based techniques is relatively small relative to the Earth's total microbial population. Although the conditions necessary to cultivate a relatively small number of species in the laboratory have been identified, scaling up this process to the majority of the remaining species remains elusive. Metagenomics offers a direct, culture-independent sampling approach for studying uncultivable microorganisms. Recent advances in this field have enabled the systematic resolution of microbial genomes (MAGs) from metagenomes. Most current methods for constructing MAGs rely on genomic sequence characteristics (GC content and base frequencies) and abundance information. These methods struggle to distinguish between species with less distinct genomic sequence features. Abundance information is typically derived from cross-sectional or longitudinal cohort studies, which can pose a barrier due to sequencing costs and computational constraints for data analysis.
[0003] To overcome these technical limitations, Hi-C technology has been applied to metagenomic sequencing. Hi-C (high-throughput chromosome conformation capture) was originally used to investigate the spatial relationships of chromatin DNA across the genome, thereby generating high-resolution maps of interactions between chromatin regulatory elements. In metagenomic studies, Hi-C cross-linking relationships can be used to identify sequences originating from the same cell, i.e., DNA molecules from the same cell (microorganism) interact more strongly than those from different cells (microorganisms). Based on this principle, contig sequences from the same microorganism can be clustered into the same group, and each group is considered a MAG. This technical feature allows for the differentiation of nucleic acids within complex microbiome samples, addressing the challenges of studying single bacterial genomes in complex samples. Compared to other MAG-study techniques, Hi-C-based metagenomics not only yields more high-quality MAGs, enabling the discovery of new species in complex environmental samples, but also allows for the definition of host relationships between bacteria, plasmids, and viruses in environmental samples, enabling the investigation of the host origins of genes such as antibiotic resistance genes in complex microbiome samples.
[0004] Metagenomic analysis based on Hi-C technology involves a series of complex analysis steps and the integration and use of multiple analysis software. It requires a lot of professional knowledge and operational experience, which poses certain difficulties for researchers without basic knowledge. Summary of the Invention
[0005] This invention provides a Hi-C-based microbial metagenomic sequencing and analysis method and a corresponding sequencing and analysis system. This method and system utilizes the Nextflow software and Docker serialization to implement an automated metagenomic Hi-C analysis method (hereinafter referred to as Meta-HIC). This method and system enables researchers without bioinformatics expertise to conduct efficient, high-quality, and convenient metagenomic Hi-C studies. This is achieved through the following technologies.
[0006] In one aspect of the present invention, a Hi-C-based microbial metagenomic sequencing and analysis method is provided, comprising the following steps:
[0007] Quality control was performed on the second-generation metagenomics raw sequencing data, second-generation Hi-C raw sequencing data, and third-generation metagenomics raw sequencing data of microbial DNA samples;
[0008] The quality-controlled third-generation Hi-C sequencing data were assembled, and after assembly, the quality-controlled second-generation metagenomic sequencing data were used for error correction to obtain high-quality genome contigs.
[0009] The quality-controlled second-generation metagenomic sequencing data are aligned with the high-quality genomic contigs, and the unique alignment result is retained. The Hi-C interaction signals between the high-quality genomic contigs are identified based on the unique alignment result, and the intensity of the Hi-C interaction signals is normalized to construct a Hi-C connection map; the high-quality genomic contigs are clustered based on the Hi-C connection map to obtain original MAGs;
[0010] The original MAGs were quality assessed to obtain high-quality MAGs, and the high-quality MAGs were annotated with species classification, structure annotation (gene prediction), function annotation (e.g., Nr, Pfam, Uniprot, KEGG, GO, COG, a total of 6 databases) and abundance analysis;
[0011] Plasmid and virus prediction was performed on the high-quality genomic contigs, and they were associated with the corresponding high-quality MAGs according to the Hi-C connection map to obtain the host relationship between plasmids, viruses and the high-quality MAGs.
[0012] Furthermore, the method for quality control of the second-generation metagenomic raw sequencing data and the second-generation Hi-C raw sequencing data is as follows:
[0013] Remove reads with N base content >5%;
[0014] Furthermore, reads with a quality value ≤ 5 and a base number of 50% were removed;
[0015] Also, reads contaminated by adapters were removed;
[0016] Furthermore, the repetitive sequences caused by PCR amplification were removed;
[0017] The quality control method for the raw sequencing data of the third-generation metagenomics was to filter reads with Q < 7 and length < 1000 bp.
[0018] Furthermore, based on the third-generation metagenomic sequencing data after quality control, the meta-mode of Flye software is used for metagenomic assembly; based on the second-generation metagenomic sequencing data after quality control, several rounds (for example, 5 rounds) of error correction are performed using plion software, and contigs with a length of less than 500bp are filtered to obtain the high-quality genome contigs.
[0019] Furthermore, the second-generation Hi-C sequencing data after quality control was compared with the high-quality genomic contigs using bwa software to obtain a sam file;
[0020] Use the view method in the samtools software to convert the sam file into a bam file;
[0021] The bam file was sorted using the sort method in the software samtools to obtain a sorted bam file.
[0022] Furthermore, combining the high-quality genomic contigs and the sorted bam files, the mkmap method in the bin3C software was used to identify and normalize Hi-C interaction signals to construct a Hi-C contact map.
[0023] Based on the Hi-C connection map, the high-quality genomic contigs were clustered using the cluster method in the software bin3C to obtain the original MAGs.
[0024] Furthermore, the software checkM is used to perform quality assessment on the original MAGs. The quality assessment mainly includes completeness and contamination indicators. MAGs with a completeness of not less than 90% or a contamination of not more than 10% are retained to obtain the high-quality MAGs.
[0025] Furthermore, the software gtdb-tk was used to perform species classification annotation on the high-quality MAGs, i.e., species classification information such as kingdom, phylum, class, order, family, genus, and species of the biological world;
[0026] The high-quality MAGs were structurally annotated using the software prokka to obtain the gene and protein sequences of each high-quality MAG;
[0027] The method for functional annotation of the high-quality MAGs is as follows:
[0028] The protein sequences of the high-quality MAGs were compared with the existing protein databases Uniprot, NR, and EggNOG using the software diamond blastp to obtain functional information; GO and COG annotation results were obtained based on the association between the databases;
[0029] The protein sequences of the high-quality MAGs were compared with the Pfam database using the software hmmscan to predict the domain structure of the genes;
[0030] The protein sequences of the high-quality MAGs were searched for homology with the Kofam database using the software kofam_scan to obtain information on metabolic pathways in which the genes were potentially involved;
[0031] The quality-controlled second-generation Hi-C sequencing data were aligned with the high-quality MAGs using the bwa software, and the abundance of MAGs was obtained using the coverm software based on the coverage depth obtained by the alignment.
[0032] Furthermore, the software plassflow was used to identify the plasmid sequences in the high-quality genomic contigs, the software VirSorter2 was used to identify the viral sequences in the high-quality genomic contigs, the software checkV was used to evaluate the integrity and contamination of the viral sequences, and finally the host relationship between the virus and plasmid and MAGs was determined based on the Hi-C connection map.
[0033] In another aspect of the present invention, a Hi-C-based microbial metagenomic sequencing and analysis system is provided, comprising a quality control module, a metagenomic assembly module, a Hi-C module, a MAGs module, and a plasmid virus prediction module;
[0034] The quality control module is used to perform quality control processing on the second-generation metagenomic raw sequencing data, the second-generation Hi-C raw sequencing data and the third-generation metagenomic raw sequencing data respectively;
[0035] The metagenomic assembly module is used to assemble the third-generation metagenomic sequencing data after quality control, and after assembly, error correction is performed using the second-generation metagenomic sequencing data after quality control to obtain high-quality genome contigs;
[0036] The Hi-C module is used to align the quality-controlled second-generation metagenomic sequencing data with the high-quality genomic contigs, retain unique alignment results, identify Hi-C interaction signals between the high-quality genomic contigs based on the unique alignment results; normalize the intensity of the Hi-C interaction signals, and construct a Hi-C connection map;
[0037] The MAGs module is used to cluster the high-quality genomic contigs based on the Hi-C connection map to obtain original MAGs; perform quality assessment on the original MAGs to obtain high-quality MAGs, and perform species classification annotation, structure annotation, function annotation and abundance analysis on the high-quality MAGs;
[0038] The plasmid virus prediction module is used to predict plasmids and viruses for the high-quality genomic contigs, and associate them with the corresponding high-quality MAGs based on the Hi-C connection map to obtain the host relationship between plasmids, viruses and high-quality MAGs.
[0039] Optionally, the system further includes a sequencing module for sequencing the microbial DNA sample to obtain second-generation metagenomic raw sequencing data, second-generation Hi-C raw sequencing data, and third-generation metagenomic raw sequencing data;
[0040] Compared with the prior art, the present invention is beneficial in that:
[0041] 1. Based on nextflow and docker, this paper builds an integrated Hi-C technology-based automated metagenomic analysis method and system, making the analysis process more portable, repeatable, and portable. High-quality metagenomic assembly results are obtained using the software Flye and pilon. The software bin3C realizes the identification of Hi-C interaction relationships, the construction of connection maps, and the clustering of MAGs. High-quality MAGs are obtained for downstream analysis using the checkM software.
[0042] 2. The present invention uses an automated metagenomic analysis method based on Hi-C technology. It uses gtdb-tk software to analyze the species information of high-quality MAGs and the evolutionary relationship information between existing species. Then, through functional annotation using prokka and seven databases, it elaborates on the functions of MAGs in many aspects, which can realize the discovery and analysis of new species in complex environmental samples.
[0043] 3. The automated metagenomic analysis method based on Hi-C technology in the present invention can directly define the host relationship between microorganisms and plasmids and viruses in the sample through Hi-C interaction relationships, and study the host origin of genes such as antibiotic resistance genes in complex microbiome samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 This is a detailed flowchart of the automated metagenomic analysis method based on Hi-C technology;
[0045] Figure 2 This is the Hi-C interaction signal diagram between normalized contigs; Figure 2 In the figure, the horizontal and vertical axes represent contigs, and the colors in the figure represent the strength of the signal between contigs;
[0046] Figure 3 This is a host relationship diagram of virus and plasmid sequences and high-quality MAGs; in the figure, green circles represent high-quality MAGs, red diamonds represent plasmids, blue triangles represent viruses, and the connecting line indicates the existence of Hi-C interaction signals between the two, and the color depth represents the signal strength. DETAILED DESCRIPTION
[0047] The technical solutions of the present invention are described clearly and completely below. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0048] The microbial metagenomic sequencing and analysis method based on Hi-C provided in this embodiment has the following specific process: Figure 1 shown.
[0049] 1. Sample pretreatment, obtaining DNA required for sequencing, and constructing a library
[0050] The samples were preprocessed to obtain sequencing DNA; second-generation libraries and third-generation libraries were constructed to obtain second-generation metagenomic raw sequencing data, second-generation Hi-C raw sequencing data, and third-generation metagenomic raw sequencing data.
[0051] 2. Sequencing Data Quality Control and Self-Control
[0052] 1. Second-generation metagenomic sequencing data statistics
[0053] Since the original sequencing data may contain low-quality sequences, adapter sequences, etc., in order to ensure the reliability of the information analysis results, the second-generation metagenomic raw sequencing data uses fastp software to remove adapters and low-quality sequences. Specifically:
[0054] Remove reads with N base content >5%;
[0055] Furthermore, reads with a quality value ≤ 5 and a base number of 50% were removed;
[0056] Also, reads contaminated by adapters were removed;
[0057] In addition, repetitive sequences caused by PCR amplification are removed.
[0058] The second-generation metagenomic sequencing data after quality control were obtained. The statistical results of the second-generation metagenomic sequencing data before and after quality control are shown in Table 1 below.
[0059] Table 1 Statistics of the amount of second-generation metagenomic sequencing data before and after quality control
[0060]
[0061] 2. Second-generation Hi-C sequencing data statistics
[0062] Since the original sequencing data may contain low-quality sequences, adapter sequences, etc., in order to ensure the reliability of the information analysis results, the second-generation Hi-C original sequencing data uses fastp software to remove adapters and low-quality sequences. Specifically:
[0063] Remove reads with N base content >5%;
[0064] Furthermore, reads with a quality value ≤ 5 and a base number of 50% were removed;
[0065] Also, reads contaminated by adapters were removed;
[0066] In addition, repetitive sequences caused by PCR amplification are removed.
[0067] The second-generation Hi-C sequencing data after quality control were obtained. The statistical results of the second-generation Hi-C data volume before and after quality control are shown in Table 2 below.
[0068] Table 2 Statistics of second-generation Hi-C data before and after quality control
[0069]
[0070] 3. Statistics of third-generation metagenomic sequencing data
[0071] The raw data format of the third-generation metagenomic sequencing data is the fast5 format, which contains all the raw sequencing signals. The fast5 format data was converted to the fastq format after base calling using the software GUPPY (version: 5.0.16). The third-generation raw sequencing data was then filtered according to the following criteria: reads with Q < 7 and length < 100 bp were filtered. The statistical results of the third-generation metagenomic sequencing data volume before and after quality control are shown in Table 3 below.
[0072] Table 3 Statistics of the amount of data from three-generation metagenomic sequencing before and after quality control
[0073]
[0074] 3. Third-generation metagenomic data assembly and error correction
[0075] Flye's metagenome assembly model, MetaFlye, uses a different approach than setting fixed K-mers. Instead, it calculates local K-mer distributions to form a global K-mer. This algorithm detects repetitive regions in the draft metagenome assembly and efficiently detects highly inconsistent sequence distributions within the assembled genome. Flye's assembled contigs are subjected to five rounds of error correction using plion software based on the second-generation metagenome data to produce high-quality assembled genomes. This approach effectively avoids single-base errors and base insertion and deletion errors introduced by Nanopore data.
[0076] Therefore, this specific embodiment uses the Flye software to assemble the quality-controlled third-generation metagenomic data. After assembly, the software Plion is used to correct errors in the second-generation metagenomic sequencing data after quality control. Finally, contigs with a length of less than 500 bp are filtered to form the final assembled genome, i.e., high-quality genome contigs. The assembly results are shown in Table 4 below.
[0077] Table 4. Assembly results of three-generation metagenomic data
[0078]
[0079] IV. Hi-C Data Analysis
[0080] First, the bwa software (parameter: -5SP) was used to align the quality-controlled second-generation Hi-C sequencing data with high-quality genomic contigs to obtain the original sam file.
[0081] The alignment results (sam files) were then converted to bam files using the view method in samtools. Finally, the bam files were sorted using the sort method in samtools to obtain sorted bam files. The statistical table of the second-generation Hi-C alignment results after quality control is shown in Table 5 below.
[0082] Table 5 Second-generation Hi-C alignment results after quality control
[0083]
[0084] bin3C is a tool for reconstructing MAGs (Metagenome-Assembled Genones) from metagenome-assembled genomes based on Hi-C signals.
[0085] The bin3C software was used to obtain the Hi-C interaction signals between pairs of contigs from the sorted bam files, and the Knight-Ruiz algorithm was used to normalize the Hi-C interaction signals. Then, a Hi-C connection graph was constructed based on the Hi-C interaction signals (the nodes are contigs, and the edge weights are the normalized Hi-C interaction signals between contigs). The normalized Hi-C interaction signal graph between contigs is shown in Figure 2. Figure 2 shown. Figure 2 In the figure, the horizontal and vertical axes represent contigs, and the colors in the figure represent the strength of the signal between contigs.
[0086] Finally, the Infomap algorithm was used to cluster the Hi-C connectivity graph, and each category was regarded as a MAG to obtain the original MAGs.
[0087] V. MAGs Evaluation
[0088] The checkM software was used to evaluate the quality of each MAG, mainly evaluating the genome integrity and contamination indicators. The genome integrity and contamination were determined by measuring the proportion of single-copy genes that were almost universally distributed in the MAG.
[0089] MAGs with genome integrity ≥ 90% and contamination < 10% were retained, i.e., high-quality MAGs, for subsequent analysis. The evaluation results of some MAGs are shown in Table 6 below.
[0090] Table 6 Evaluation results of some MAGs
[0091]
[0092] 6. High-quality MAGs annotation
[0093] 1. Species classification annotation and abundance analysis.
[0094] The software gtdb-tk was used to perform species classification annotation on high-quality MAGs. This specific embodiment mainly performs bacterial genome classification and species classification annotation based on 120 single-copy genomes that are commonly found in bacterial species.
[0095] The quality-controlled second-generation metagenomic sequencing data were then compared with high-quality MAGs using the bwa software, and the abundance of high-quality MAGs was assessed using the coverm software. The species classification annotation results and abundance of some high-quality MAGs are shown in Table 7 below.
[0096] Table 7 Species classification annotation results and abundance analysis results of some high-quality MAGs
[0097]
[0098] 2. Structural Annotation
[0099] Coding gene prediction was performed on high-quality MAG genomes using the Prokka software (Version 1.14.6). Prokka is a collection of gene element prediction tools, incorporating Prodigal to predict coding genes, Aragorn to predict tRNAs, RNAmmer to predict rRNAs, and Infernal to predict miscRNAs. The predicted gene elements are then aggregated to complete preliminary structural annotation. For the MAG012 gene, the prediction results are shown in Table 8 below.
[0100] Table 8 Prediction results of MAG012 gene
[0101]
[0102] 3. Functional annotation
[0103] Based on the BLAST or hmmer comparison method, the protein sequence of the predicted gene was compared with existing protein function databases such as Nr, Pfam, Uniprot, KEGG, GO, EggNOG, etc., with an E value of 1e-5, to obtain the functional annotation results of each gene in each database.
[0104] The method for functional annotation of the high-quality MAGs is as follows:
[0105] The protein sequences encoded by the above genes were compared with the existing protein databases Uniprot, NR, and EggNOG using the software diamond blastp to obtain functional information; GO and COG annotation results were obtained based on the association between the databases;
[0106] The genes were compared with the Pfam database using the software hmmscan to predict the gene domains;
[0107] The gene was searched for homology with the Kofam database using the software kofam_scan to obtain information on metabolic pathways that the gene is potentially involved in. Taking MAG012 as an example, the statistical results of its functional annotation in various databases are shown in Table 9 below.
[0108] Table 9 Statistical results of MAG012 functional annotation in various databases
[0109]
[0110] 7. Interaction Analysis
[0111] The software PlasFlow was used to identify the plasmids in the assembly results, and the results with a confidence level greater than 0.8 were selected as the final plasmid sequences.
[0112] VirSorter2 software was used to identify viral sequences in the assembly results. The integrity and contamination of the viral sequences were then assessed using checkV software. Viral sequences with an integrity of ≥90% and a contamination of <10% were retained. Finally, the host relationships of viral and plasmid sequences with high-quality MAGs were determined using Hi-C connectivity maps. The statistical table of plasmid and virus identification results is shown in Table 10 below. Figure 3 shown. Figure 3 In the figure, green circles represent high-quality MAGs, red diamonds represent plasmids, and blue triangles represent viruses. The connecting line indicates the presence of Hi-C interaction signals between the two, and the color depth indicates the signal strength.
[0113]
[0114] The above specific embodiments describe the implementation of the present invention in detail, but the present invention is not limited to the specific details of the above embodiments. Within the scope of the claims and technical concept of the present invention, various simple modifications and changes can be made to the technical solution of the present invention, and these simple modifications all fall within the scope of protection of the present invention.
Claims
1. A microbial metagenomic sequencing and analysis method based on Hi-C, characterized in that: The following steps are involved: Quality control was performed on the second-generation metagenomics raw sequencing data, second-generation Hi-C raw sequencing data, and third-generation metagenomics raw sequencing data of microbial DNA samples; The quality-controlled third-generation metagenomic sequencing data were assembled, and after assembly, the quality-controlled second-generation metagenomic sequencing data were used for error correction to obtain high-quality genome contigs. Aligning the quality-controlled second-generation Hi-C sequencing data with the high-quality genomic contigs, retaining unique alignment results, identifying Hi-C interaction signals between the high-quality genomic contigs based on the unique alignment results, normalizing the intensities of the Hi-C interaction signals, and constructing a Hi-C connectivity map; Clustering the high-quality genomic contigs based on the Hi-C connection map to obtain original MAGs; Performing quality assessment on the original MAGs to obtain high-quality MAGs, and performing species classification annotation, structure annotation, function annotation and abundance analysis on the high-quality MAGs; Plasmid and virus prediction was performed on the high-quality genomic contigs, and they were associated with the corresponding high-quality MAGs according to the Hi-C connection map to obtain the host relationship between plasmids, viruses and the high-quality MAGs.
2. The Hi-C based microbial metagenomic sequencing and analysis method according to claim 1, characterized in that: The method for quality control of the second-generation metagenomic raw sequencing data and the second-generation Hi-C raw sequencing data is as follows: Remove reads with N base content >5%; Furthermore, reads with a quality value ≤ 5 and a base number of 50% were removed; Also, reads contaminated by adapters were removed; Furthermore, the repetitive sequences caused by PCR amplification were removed; The quality control method for the raw sequencing data of the third-generation metagenomics was to filter reads with Q < 7 and length < 1000 bp.
3. The Hi-C based microbial metagenomic sequencing and analysis method according to claim 1, characterized in that: Based on the third-generation metagenomic sequencing data after quality control, the meta-mode of Flye software was used for metagenomic assembly; based on the second-generation metagenomic sequencing data after quality control, plion software was used to perform 5 rounds of error correction, and contigs with a length of less than 500bp were filtered to obtain the high-quality genome contigs.
4. The Hi-C based microbial metagenomic sequencing and analysis method according to claim 1, characterized in that: Use bwa software to align the quality-controlled second-generation Hi-C sequencing data with the high-quality genomic contigs to obtain a sam file; Use the view method in the samtools software to convert the sam file into a bam file; The bam file was sorted using the sort method in the software samtools to obtain a sorted bam file.
5. The Hi-C based microbial metagenomic sequencing and analysis method according to claim 4, characterized in that: Combining the high-quality genomic contigs and the sorted bam files, the mkmap method in the bin3C software was used to identify and standardize Hi-C interaction signals to construct a Hi-C connection map; Based on the Hi-C connection map, the high-quality genomic contigs were clustered using the cluster method in the software bin3C to obtain the original MAGs.
6. The Hi-C based microbial metagenomic sequencing and analysis method according to claim 1, characterized in that: The quality of the original MAGs was evaluated using the checkM software, and MAGs with an integrity of not less than 90% or a contamination of not more than 10% were retained to obtain the high-quality MAGs.
7. The Hi-C based microbial metagenomic sequencing and analysis method according to claim 1, characterized in that: The high-quality MAGs were annotated with species classification using the software gtdb-tk; The high-quality MAGs were structurally annotated using the software prokka to obtain the gene and protein sequences of each high-quality MAG; The method for functional annotation of the high-quality MAGs is as follows: The protein sequences of the high-quality MAGs were compared with the existing protein databases Uniprot, NR, and EggNOG using the software diamond blastp to obtain functional information; GO and COG annotation results were obtained based on the association between the databases; The protein sequences of the high-quality MAGs were compared with the Pfam database using the software hmmscan to predict the domain structure of the genes; The protein sequences of the high-quality MAGs were searched for homology with the Kofam database using the software kofam_scan to obtain information on metabolic pathways in which the genes were potentially involved; The quality-controlled second-generation Hi-C sequencing data were aligned with the high-quality MAGs using the bwa software, and the abundance of MAGs was obtained using the coverm software based on the coverage depth obtained by the alignment.
8. The Hi-C based microbial metagenomic sequencing and analysis method according to claim 1, characterized in that: The plasmid sequences in the high-quality genomic contigs were identified using the software plassflow, the viral sequences in the high-quality genomic contigs were identified using the software VirSorter2, the integrity and contamination of the viral sequences were evaluated using the software checkV, and finally the host relationships of the viruses and plasmids with MAGs were determined based on the Hi-C connection map.
9. A Hi-C based microbial metagenomic sequencing and analysis system, characterized in that: Including quality control module, metagenomic assembly module, Hi-C module, MAGs module, plasmid virus prediction module; The quality control module is used to perform quality control processing on the second-generation metagenomic raw sequencing data, the second-generation Hi-C raw sequencing data and the third-generation metagenomic raw sequencing data respectively; The metagenomic assembly module is used to assemble the third-generation metagenomic sequencing data after quality control, and after assembly, error correction is performed using the second-generation metagenomic sequencing data after quality control to obtain high-quality genome contigs; The Hi-C module is used to align the quality-controlled second-generation metagenomic sequencing data with the high-quality genomic contigs, retain unique alignment results, identify Hi-C interaction signals between the high-quality genomic contigs based on the unique alignment results; normalize the intensity of the Hi-C interaction signals, and construct a Hi-C connection map; The MAGs module is used to cluster the high-quality genomic contigs based on the Hi-C connection map to obtain original MAGs; perform quality assessment on the original MAGs to obtain high-quality MAGs, and perform species classification annotation, structure annotation, function annotation and abundance analysis on the high-quality MAGs; The plasmid virus prediction module is used to predict plasmids and viruses for the high-quality genomic contigs, and associate them with the corresponding high-quality MAGs based on the Hi-C connection map to obtain the host relationship between plasmids, viruses and high-quality MAGs.
Citation Information
Patent Citations
Metagenome GutHi-C library building method suitable for microbial population and application of metagenome GutHi-C library building method
CN116606910A
Metagenome structure variation automatic analysis method and device
CN116884480A