Pit mud metagenome data automatic analysis method and system

By using automated analysis methods and systems and integrating multi-step data processing workflows, the problem of low efficiency in metagenomic data analysis has been solved, enabling efficient and accurate metagenomic data analysis of pit sludge and lowering the professional threshold.

CN121811973APending Publication Date: 2026-04-07WULIANGYE +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing metagenomic data analysis methods are inefficient, rely on specialized knowledge and manual operation, and are prone to result bias, making it difficult for non-bioinformatics professionals to perform efficient analysis.

Method used

This invention provides an automated method and system for metagenomic data analysis of sediment, which integrates steps such as raw data quality control, genome assembly, binning and optimization, and species annotation. It uses pre-written scripts to automatically complete data analysis, adapts to different sequencing data types, and generates analysis reports.

Benefits of technology

It has achieved standardization and automation of the metagenomic analysis process, reduced reliance on professional knowledge and manual operation, improved data processing efficiency and accuracy, simplified operation steps, and lowered the analysis threshold for non-professionals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811973A_ABST
    Figure CN121811973A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of metagenomics, discloses an automatic analysis method and system for pit mud metagenomic data, and aims at solving the problem that an existing method is poor in efficiency and accuracy, and the scheme mainly comprises the steps that a sequencing data type, a file path and analysis parameters are received; performing quality control on the original offline data; sequence assembly is carried out, and a contigs file is generated; carrying out assembly quality evaluation on the contigs file; carrying out genome binning by using at least two binning tools; integrating output results of the binning tool, and performing optimization based on a preset integrity threshold value and a preset pollution degree threshold value to obtain an optimized binning genome data set; evaluating and optimizing the integrity, the pollution degree and the strain heterogeneity of the binning genome; calculating coverage and relative abundance; performing species classification annotation and function annotation; and integrating the result data of the previous steps to generate an analysis report. According to the method, automatic analysis of metagenome data is realized, and the analysis efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of metagenomics technology, specifically to an automated method and system for analyzing metagenomic data from pit mud. Background Technology

[0002] Metagenomics is a key technology for studying the composition and functional genes of microbial communities in the environment. It allows for direct sequencing and analysis of the genetic material of all microorganisms in a sample without the need for microbial culture, thus enabling the analysis of microbial diversity and potential functions at the species and gene levels. It plays an indispensable role in environmental microbiology research and fermentation process analysis. Currently, metagenomic sequencing technologies are mainly divided into second-generation short-read sequencing, represented by the Illumina platform, and third-generation long-read sequencing, represented by the Nanopore platform. Second-generation sequencing offers high throughput, low cost, and good accuracy, but its short read lengths are not conducive to complete genome assembly. Third-generation sequencing significantly increases read lengths, facilitating the acquisition of more complete genomic information, but its single-base error rate is relatively high. In recent years, hybrid sequencing strategies combining second- and third-generation sequencing technologies have emerged as an effective means of in-depth microbiome research, balancing data quality and assembly integrity.

[0003] Although metagenomic sequencing technology has been widely applied to the study of microorganisms in sediment, subsequent bioinformatics analysis of metagenomic data still faces significant challenges. Current analysis workflows typically involve multiple steps, including raw data quality control, sequence assembly, genome binning, quality assessment, species and functional annotation, and abundance calculation. Each step relies on different specialized software, such as Trimmomatic, SPAdes, MetaBAT2, CheckM2, and GTDB-Tk. Researchers must master the usage, parameter settings, and input / output format requirements of numerous software programs beforehand, and perform tedious manual data processing and conversion at each step. This highly specialized and manual analysis model not only makes data analysis difficult, time-consuming, and inefficient, but also highly susceptible to operational errors that can lead to biased results. This severely restricts the speed of research output and its application, posing a significant obstacle, especially for researchers without a bioinformatics background. Summary of the Invention

[0004] This invention aims to address the problems of poor efficiency and accuracy in existing metagenomic data analysis methods, and proposes an automated method and system for analyzing metagenomic data from pit silt.

[0005] The technical solution adopted by the present invention to solve the above-mentioned technical problems is as follows: In a first aspect, the present invention provides an automated method for analyzing metagenomic data of pit mud, the method comprising: Step 1: Receive the sequencing data type, file path, and analysis parameters input by the user; the sequencing data type includes Illumina data, Nanopore data, or a mixture of both. Step 2: Perform quality control on the raw sequencing data from the gene sequencer according to the sequencing data type to obtain high-quality sequencing data; Step 3: Perform sequence assembly on the high-quality sequencing data to generate contigs files; Step 4: Assess the assembly quality of the contigs files; Step 5: Perform genome binning on the contigs file using at least two binning tools; Step 6: Integrate the output results of at least two binning tools described in Step 5, and optimize them based on preset integrity thresholds and contamination thresholds to obtain an optimized binned genome dataset; Step 7: Evaluate the integrity, contamination level, and strain heterogeneity of the optimized binned genome; Step 8: Calculate the coverage and relative abundance of the genome in each optimized bin; Step 9: Perform species classification and functional annotation on the optimized binned genome; Step 10: Integrate the results data from Steps 4, 7, 8, and 9 to generate an analysis report.

[0006] Furthermore, in step 1, the analysis parameters include parameters for dynamically specifying the analysis termination step, which corresponds to the end of step 2, the end of step 3, the end of step 6, or the end of step 10.

[0007] Further, step 2 includes: When the sequencing data is Illumina data, Trimmomatic is used for filtering, and FastQC is used for quality assessment and visualization. When the sequencing data is Nanopore data, Porechop is used for adapter removal, FiltLong is used for quality filtering, and Nanoplot is used for quality assessment and visualization. When the sequencing data type is mixed data, the quality control process for the Illumina data portion is executed, and the quality control process for the Nanopore data portion is executed.

[0008] Further, step 3 includes: selecting an assembly tool according to the type of sequencing data; For Illumina data, use SPAdes or MEGAHIT for assembly; For Nanopore data, Flye or Canu are selected for assembly; For mixed data, Unicycler or OPERA-MS can be used for assembly.

[0009] Furthermore, in step 5, the at least two bin-separating tools are selected from MetaBAT2, MaxBin2, and CONCOCT; For Illumina data, MetaBAT2, MaxBin2, and CONCOCT are invoked in parallel. For Nanopore data or mixed data, call MetaBAT2 and CONCOCT.

[0010] Furthermore, in step 6, the preset integrity threshold is greater than or equal to 70%, and the contamination threshold is less than or equal to 10%.

[0011] Furthermore, step 9 specifically includes: Use GTDB-Tk for species classification annotation; Using Prokka for gene prediction and functional annotation; Based on the output of Prokka, functional enrichment analysis was performed using EggNOG-mapper.

[0012] Furthermore, step 6 also includes: When the optimized binned genome dataset is empty, the original binning results generated in step 5 are output as the final binned genome file. The suffixes of the binned genome files corresponding to the optimized binned genome or the original binned genome results are standardized.

[0013] Further, step 10 includes: Extract binning identifiers from the binning genome file names corresponding to the optimized binning genome or the original binning results output in step 6, and clean up the normalized suffixes in the binning identifiers; Using the cleaned bin identifier as the association key, merge the result data from steps 4, 7, 8, and 9. The process terminates when any critical data is detected to be missing from the quality assessment data in step 7, the species annotation data in step 9, or the coverage data in step 8.

[0014] In a second aspect, the present invention provides an automated analysis system for metagenomic data of pit silo mud, used to implement the automated analysis method for metagenomic data of pit silo mud as described in the first aspect, the system comprising: The parameter parsing module is used to receive the sequencing data type, file path, and analysis parameters input by the user; the sequencing data type includes Illumina data, Nanopore data, or a mixture of both. The quality control module is used to perform quality control on the raw sequencing data from the gene sequencer according to the sequencing data type, so as to obtain high-quality sequencing data. An assembly execution module is used to assemble the high-quality sequencing data into contigs files. An assembly quality assessment module is used to assess the assembly quality of the contigs files. A multi-tool binning module is used to perform genome binning on the contigs file using at least two binning tools; The binning optimization module is used to integrate the output results of the multi-tool binning module and optimize them based on preset integrity thresholds and contamination thresholds to obtain an optimized binned genome dataset. The binning quality assessment module is used to evaluate the integrity, contamination, and strain heterogeneity of the optimized binning genome; The coverage calculation module is used to calculate the coverage and relative abundance of each optimized bin genome; The annotation module is used to perform species classification annotation and functional annotation on the optimized binned genome; The report generation module is used to integrate the result data output by the assembly quality assessment module, the boxing quality assessment module, the coverage calculation module, and the annotation module to generate an analysis report.

[0015] The beneficial effects of this invention are as follows: The automated analysis method and system for metagenomic data from sediment provided by this invention integrates all analytical steps from raw data quality control, genome assembly, binning and optimization, species annotation, functional annotation, and statistical analysis of genome quality data, which can meet the analytical needs of researchers. Furthermore, it dynamically adapts corresponding tools and parameters for second-generation, third-generation, and mixed sequencing data, and can automatically complete data analysis and generate results reports using pre-written scripts. This achieves standardization and automation of the metagenomic analysis process, significantly reducing the reliance on professional bioinformatics knowledge and manual operation. It realizes fully automated and standardized analysis from raw data to final report, greatly lowering the threshold for data analysis for non-bioinformatics professionals, simplifying the operational steps for bioinformatics analysts, reducing program errors caused by human error, and significantly improving data processing efficiency and accuracy. Attached Figure Description

[0016] Figure 1 A schematic diagram of the workflow structure of the automated analysis method for metagenomic data of pit mud provided in the embodiment; Figure 2 A schematic diagram illustrating the binning composition of the second-generation sequencing data of the pit mud provided in this embodiment; Figure 3 A schematic diagram illustrating the binning composition of the third-generation sequencing data of pit mud provided in this embodiment; Figure 4 A schematic diagram illustrating the binning composition of second-generation and third-generation sequencing data of pit mud provided for this embodiment; Figure 5 A schematic diagram of the structure of the automated analysis system for metagenomic data of pit mud provided in the example. Detailed Implementation

[0017] Currently, analyzing metagenomic sequencing data often requires researchers to spend significant time and effort mastering bioinformatics analysis methods and learning to use specialized software. During the analysis process, manual data processing is necessary for each step to meet the input requirements of different software. These issues make data analysis challenging and result in delays in obtaining results. Therefore, there is an urgent need to develop an automated analysis workflow for sediment metagenomic data of different sequencing types to achieve rapid result acquisition, lower the analysis threshold, reduce manual steps, and lower labor costs.

[0018] Based on this, the technical solution of this invention is proposed. In this invention, the parameter parsing module receives and parses the data type and parameters specified by the user, and automatically selects and executes an appropriate quality control scheme accordingly. Subsequently, the process transfers the quality-controlled data to the assembly step to generate continuous sequences (contigs files) and perform preliminary quality assessment. Then, the system employs a multi-tool joint strategy to bin the sequences and optimizes the binning results through integration and threshold filtering. Afterward, the process automatically performs quality assessment, coverage calculation, and species and functional annotation on the optimized binned genomes. Finally, the system automatically collects the key output data from each of the aforementioned steps and integrates them to generate a complete analysis report. The entire process replaces manual data conversion and manipulation through automatic matching and transfer of input and output formats between steps, achieving automated processing from raw data to analysis reports without requiring manual intervention in data transfer and parameter adjustment between analysis steps.

[0019] The technical solutions in this embodiment will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0020] Figure 1 A flowchart illustrating an automated method for analyzing metagenomic data from pit sludge is shown below. Please refer to [link / reference]. Figure 1 The method includes the following steps: Step 1: Receive the sequencing data type, file path, and analysis parameters input by the user; the sequencing data type includes Illumina data, Nanopore data, or a mixture of both.

[0021] This step provides a unified, configurable entry point for analysis and establishes the initial state and execution path of the entire automation process. It sets specific execution rules and data sources for all subsequent analysis steps by parsing the metadata input by the user.

[0022] Specifically, by receiving the sequencing data type, the system can anticipate whether it will process short-read, long-read, or mixed data, thereby activating entirely different sub-processes such as quality control and assembly. By receiving the file path, the system can automatically locate the original sequencing data file, eliminating the need for users to repeatedly enter it in each tool. By receiving analysis parameters (such as specifying the tools to use, adjusting quality thresholds, and setting analysis termination steps), users can flexibly control the depth, accuracy, and scope of analysis without modifying the underlying script, making the workflow both standardized and adjustable.

[0023] The analysis parameters include a parameter for dynamically specifying the analysis termination step, which corresponds to the end of step 2, step 3, step 6, or step 10. In practical applications, after receiving the user's analysis termination step parameter, the system checks whether the current step matches the user-specified termination step after each corresponding key step is completed. If they match, the process termination routine is triggered, automatically organizing and outputting all results up to the current step, and then exiting normally without executing subsequent steps. For example, termination occurs at the end of step 2 (quality control), at which point the quality-controlled sequence file and quality report will be output; termination occurs at the end of step 3 (sequence assembly), at which point the assembled Contigs file and its quality assessment report will be output; termination occurs at the end of step 6 (binning optimization), at which point the optimized binned genome set will be output; and termination occurs at the end of step 10 (report generation), i.e., the complete workflow is run and the final analysis report is generated.

[0024] Step 2: Perform quality control on the raw sequencing data from the gene sequencer according to the sequencing data type to obtain high-quality sequencing data.

[0025] In this embodiment, when the sequencing data type is Illumina data, Trimmomatic is used for filtering, and FastQC is used for quality assessment and visualization; when the sequencing data type is Nanopore data, Porechop is used for adapter removal, FiltLong is used for quality filtering, and Nanoplot is used for quality assessment and visualization; when the sequencing data type is mixed data, the Illumina data quality control process is executed for the Illumina data portion, and the Nanopore data quality control process is executed for the Nanopore data portion.

[0026] Raw data refers to the initial data files directly output from gene sequencers such as Illumina and Nanopore after sequencing, without any bioinformatics processing. These are typically stored in a specific standard file format for subsequent analysis software to read, such as the FASTQ file format.

[0027] This step is used to perform standardized cleaning and quality verification tailored to the physical characteristics and error patterns of different sequencing technologies. Through a conditional logic, it dynamically maps general quality control instructions to three dedicated processing workflows for second-generation sequencing (Illumina), third-generation sequencing (Nanopore), or their mixed data. This eliminates systematic errors and technically specific interferences introduced during the sequencing process in the raw data, outputting high-confidence sequence data that meets the requirements for subsequent analysis.

[0028] Specifically, errors in Illumina data primarily manifest as misidentification of single bases. Therefore, using Trimmomatic for quality-score-based trimming and filtering is the most efficient approach. The main issues with Nanopore data lie in adapter sequence remnants and low-quality fragments embedded in long reads. Therefore, Porechop is used to remove adapters first, followed by FiltLong filtering of the overall read length based on length and quality. Before and after cleaning, FastQC (for Illumina) or NanoPlot (for Nanopore) is used to generate quality reports, providing visualizations of key indicators such as base quality distribution, GC content, adapter contamination, and read length distribution. This not only verifies the cleaning effect but also provides users with intuitive diagnostic data quality information. For mixed data, this step involves parallel and independent processing, treating the Illumina and Nanopore portions as two separate data streams, each processed through its optimized dedicated workflow. This ensures that each type of data receives the most professional processing, maximizing its value.

[0029] Step 3: Perform sequence assembly on the high-quality sequencing data to generate contigs files.

[0030] In this embodiment, the assembly tool is selected according to the type of sequencing data; wherein: For Illumina data, use SPAdes or MEGAHIT for assembly; For Nanopore data, Flye or Canu are selected for assembly; For mixed data, Unicycler or OPERA-MS can be used for assembly.

[0031] As is understandable, assembly is a crucial step in metagenomic analysis, its role being to reassemble quality-controlled, discrete sequencing reads into longer, continuous DNA fragments. This step, based on the read length characteristics and error patterns of the input data, calls an optimized algorithm to reconstruct short sequence fragments into long, continuous sequences. It leverages the algorithm's advantages to overcome inherent defects in various data types, generating the most complete and accurate continuous sequences possible for downstream analysis.

[0032] In practical applications, Illumina short-read data offers high accuracy but suffers from short fragments, making it prone to losing continuity in repetitive sequence regions. Therefore, SPAdes or MEGAHIT, based on de Bruin diagrams and adept at handling high-coverage short-read data, are chosen. Nanopore long-read data can span repetitive regions but has a higher single-base error rate. Therefore, Flye or Canu, specifically designed for handling long, high-error-rate reads, are selected, employing overlap-layout-consensus or similar strategies for error correction and assembly. Hybrid data leverages the advantages of both approaches, using Unicycler or OPERA-MS to synergistically integrate short-read accuracy and long-read continuity information for hybrid assembly, achieving superior assembly completeness.

[0033] The final output of all assembly processes is standardized into a contig file (usually in FASTA format). This file contains all the continuous sequences assembled from the microbial community and serves as the common basis for all subsequent analyses such as binning and annotation.

[0034] Step 4: Assess the assembly quality of the contigs files.

[0035] This step is used to automate the quality assessment of the contigs files generated during assembly. In practical applications, it uses standard tools such as QUAST to calculate key statistical indicators, including the number of contigs, N50, the longest contig, and the total assembly length, generating structured quality assessment results data. This provides objective and quantitative quality metrics for the assembly results, serving both as a verification of upstream assembly processes and as a decisive basis for determining whether the assembly is worthy of downstream sorting and analysis.

[0036] Step 5: Perform genome binning on the contigs file using at least two binning tools.

[0037] In this embodiment, the at least two binning tools are selected from MetaBAT2, MaxBin2, and CONCOCT; wherein: For Illumina data, MetaBAT2, MaxBin2, and CONCOCT are invoked in parallel. For Nanopore data or mixed data, call MetaBAT2 and CONCOCT.

[0038] Different binning tools (such as MetaBAT2, MaxBin2, and CONCOCT) have different focuses in their core algorithms; for example, they may focus on sequence tetranucleotide frequency, coverage variation, or a combined model of both. Using at least two tools in parallel is equivalent to performing cluster analysis on the same set of data from multiple dimensions. This can compensate for the blind spots or biases that may exist in a single algorithm, thereby improving the detection capability of different species (especially rare species) in complex communities.

[0039] This step utilizes multiple complementary algorithms to cluster the mixed-assembled sequences, and improves the success rate of recovering individual microbial genomes from complex communities through a redundancy strategy of parallel execution.

[0040] In practical applications, for Illumina data, due to its large volume, high coverage, and clear sequence composition signals, all three mainstream tools (MetaBAT2, MaxBin2, and CONCOCT) are used in parallel to fully utilize its data advantages and pursue the highest resolution binning results. For Nanopore or mixed data, the characteristics of long-read or mixed data (such as higher error rates and more complex coverage distributions) may make the default MaxBin2 model unstable. Therefore, a more robust tool combination (MetaBAT2 and CONCOCT) is strategically used on this type of data, maintaining the advantage of algorithm redundancy while ensuring the reliability of the results.

[0041] After performing genome binning in this step, multiple sets of raw binning results are generated, providing a data foundation for integration and screening based on integrity and contamination thresholds.

[0042] Step 6: Integrate the output results of at least two binning tools described in Step 5, and optimize them based on preset integrity and contamination thresholds to obtain an optimized binned genome dataset.

[0043] In this embodiment, the preset integrity threshold is greater than or equal to 70%, and the contamination threshold is less than or equal to 10%.

[0044] This step integrates and filters based on quality thresholds, merging and elevating redundant and inconsistent raw results from multiple binning tools into a unified, high-quality standard genome set.

[0045] In practical applications, the multiple sets of original binning results generated in step 5 are cross-referenced and integrated through voting. A preset genome quality standard (integrity ≥70% and contamination ≤10%) is applied for automated filtering, thus retaining reliable binnings with high integrity and low contamination while eliminating low-quality or contradictory clustering results. This process replaces manual comparison and screening with a standardized workflow, significantly improving the reliability and efficiency of result acquisition.

[0046] In this embodiment, when the optimized binning genome dataset is empty, the original binning results generated in step 5 are used as the final binned genome file output.

[0047] Finally, the file extensions of the binned genome files corresponding to the optimized binned genome or the original binned results were standardized to ".fa".

[0048] Step 7: Evaluate the integrity, contamination, and strain heterogeneity of the optimized binned genome.

[0049] This step is used to perform standardized quality metrics based on the core gene set on the optimized binned genome dataset output from step 6. In practical applications, dedicated tools such as CheckM2 can be used to automatically calculate quality assessment data, including integrity (characterizing the completeness of genome recovery), contamination (characterizing the degree of foreign sequence contamination), and strain heterogeneity (characterizing whether the bin contains a single strain), by comparing each binned genome sequence with a large microbial reference gene database.

[0050] Step 8: Calculate the coverage and relative abundance of the genome in each optimized bin.

[0051] This step is used to back-attach the quality-controlled raw sequencing reads to each optimized bin genome, calculate its coverage depth and community proportion, and thus quantify the abundance information of each candidate genome in the original sample.

[0052] In practical applications, tools such as CoverM can be used to take the high-quality sequencing data obtained in step 2 as input and compare it with each optimized bin genome file generated in step 6. The coverage (i.e., the depth of the average sequencing read length covering each genome base, reflecting the sequencing adequacy of the genome) and relative abundance (i.e., the proportion of reads belonging to the genome to the total read length, reflecting its relative proportion in the microbial community) of each optimized bin genome can be automatically calculated.

[0053] Step 9: Perform species classification annotation and functional annotation on the optimized binned genome.

[0054] In this embodiment, it specifically includes: Use GTDB-Tk for species classification annotation; Using Prokka for gene prediction and functional annotation; Based on the output of Prokka, functional enrichment analysis was performed using EggNOG-mapper.

[0055] This step utilizes a standardized, automated annotation pipeline to provide a layered biological interpretation of optimized binned genomes, from species classification to gene function. In practical application, firstly, GTDB-Tk, based on the latest taxonomic database, is used to assign standardized species classification labels from phylum to species to each binned genome; then, Prokka is used for gene prediction, and basic functional annotations (such as COG, PFAM, etc.) are performed on the predicted genes; finally, the gene prediction results from Prokka are submitted to EggNOG-mapper for more in-depth functional enrichment analysis.

[0056] The above process is fully automated. The system calls the tools in sequence. The launch of EggNOG-mapper is a prerequisite for the successful operation of Prokka in the previous step. Finally, it generates a structured species classification table, gene catalog and multi-level functional annotation results, providing a core data foundation for understanding the species composition and metabolic potential of microbial communities.

[0057] Step 10: Integrate the results data from Steps 4, 7, 8, and 9 to generate an analysis report.

[0058] In this embodiment, it specifically includes: Extract binning identifiers from the binning genome file names corresponding to the optimized binning genome or the original binning results output in step 6, and clean up the normalized suffixes in the binning identifiers; Using the cleaned bin identifier as the association key, merge the result data from steps 4, 7, 8, and 9. The process terminates when any critical data is detected to be missing from the quality assessment data in step 7, the species annotation data in step 9, or the coverage data in step 8.

[0059] This step is used to structurally integrate key results scattered across multiple independent analysis steps through an automated data fusion and verification mechanism, generating a unified and complete final analysis report.

[0060] In practical applications, a unique binning identifier is first extracted and standardized from the binning genome file names and used as the association key. Then, based on this identifier, four key data categories—assembly quality (step 4), binning quality (step 7), coverage and abundance (step 8), and species and functional annotation (step 9)—are automatically associated and merged. This process also incorporates data integrity verification; the workflow automatically terminates when any key data is detected as missing to ensure report consistency. This step uses scripts (such as Python / Pandas) to automatically perform data cleaning, association, merging, and formatted output, ultimately generating a structured comprehensive report (such as TSV / Excel format), achieving automated conversion from multi-source heterogeneous intermediate results to a directly interpretable end product.

[0061] The technical solution in this embodiment will be clearly and completely described below using Illumina data, Nanopore data, and a mixture of the two as examples.

[0062] Example 1: Automated analysis workflow for second-generation sequencing data (Illumina): Step 101: Input the raw data from the instrument into the automated analysis program by specifying the Illumina data type, file path, and analysis parameters; Step 102: Use the software Trimmomatic to perform quality control on the raw data. Delete bases with a quality value lower than 20 before and after the data is processed. Scan a wide sliding window of 4 bases. When the average quality of each base drops below 20, perform shearing. After the above steps are completed, delete reads with less than 50 bases. Step 103: Specify Megahit as the assembly tool to assemble the quality-controlled data and generate contigs files; Step 104: Use the QUAST tool to assess the quality of the assembly results and generate a quality report in the contigs file; Step 105: Call the MetaBAT2, MaxBin2, and CONCOCT tools of the MetaWrap software to perform parallel binning. The results generated by each binning tool are stored in their respective folders. Step 106: Use the bin_refinement function of the metawrapp software to integrate the outputs of multiple binning tools, set the integrity threshold to ≥70% and the contamination threshold to ≤10% to generate an optimized binned genome; Step 107: Use CheckM2 software to evaluate the integrity, contamination, and other indicators of the optimized binned genomes, see Table 1; Table 1. Basic parameters for high-quality binning of second-generation sequencing data (integrity >95%, contamination <5%). Step 108: Use CoverM to calculate the coverage and relative abundance of the genome in each bin. See the visualization results below. Figure 2 , Figure 2 Different colored sector regions represent different binned genomes, and the area of ​​each sector represents the proportion of the relative abundance or coverage of each binned genome in the entire microbial community. Step 109: Use GTDB-Tk to perform species classification annotation on the optimized binned genome; use Prokka to perform gene prediction and functional annotation; use the Prokka output results for functional enrichment analysis using EggNOG-mapper software. Step 110: Integrate the results data, such as quality assessment, species annotation, coverage and relative abundance data, to generate a comprehensive report.

[0063] Example 2: Automated analysis workflow for third-generation sequencing data (Nanopore): Step 201: Input the raw data from the machine into the automated analysis program by specifying the data Nanopore type, file path, and analysis parameters; Step 202: Use porechop software to identify and remove the connector sequence of the original offline data. Use FiltLong to perform quality control on the data after removing the connectors. The minimum length is 200 bp and the minimum average quality value is greater than or equal to 7. Use NanoPlot software to visualize the data after quality control. Step 203: Specify Flye as the assembly tool to assemble the quality-controlled data and generate contigs files; Step 204: Use the QUAST tool to assess the quality of the assembly results and generate a quality report in the contigs file; Step 205: Call the MetaBAT2 and CONCOCT tools of the Metawrapp software to perform parallel binning. The results generated by each binning tool are stored in their respective folders. Step 206: Use the bin_refinement function of the metawrapp software to integrate the outputs of multiple binning tools, set the integrity threshold to ≥70% and the contamination threshold to ≤10% to generate an optimized binned genome; Step 207: Use CheckM2 software to evaluate the integrity, contamination, and other indicators of the optimized binned genome, see Table 2; Table 2. Basic parameters for high-quality binning of third-generation sequencing data (integrity >95%, contamination <5%). Step 208: Use CoverM to calculate the coverage and relative abundance of the genome in each bin. See the visualization results below. Figure 3 , Figure 3 Different colored sector regions represent different binned genomes, and the area of ​​each sector represents the proportion of the relative abundance or coverage of each binned genome in the entire microbial community. Step 209: Use GTDB-Tk to perform species classification annotation on the optimized binned genome; use Prokka to perform gene prediction and functional annotation; use EggNOG-mapper software to perform functional enrichment analysis based on the Prokka output results. Step 210: Integrate the results data, such as quality assessment, species annotation, coverage and relative abundance data, to generate a comprehensive report.

[0064] Example 3: Automated analysis workflow for combining second-generation sequencing data with third-generation sequencing data: Step 301: Input the raw data from the machine into the automated analysis program by specifying that the data is of mixed type, the file path, and the analysis parameters.

[0065] Step 302: Use the software Trimmomatic to perform quality control on the raw second-generation sequencing data, deleting bases with a quality value below 20 before and after sequencing. Scan a 4-base wide sliding window, and cut the data when the average quality of each base drops below 20. After completing the above steps, delete reads with less than 50 bases. Use the Porechop software to identify and remove adapter sequences from the raw third-generation sequencing data. Use FiltLong to perform quality control on the data after adapter removal, with a minimum length of 200 bp and a minimum average quality value greater than or equal to 7. Use the NanoPlot software to visualize the quality of the quality-controlled data. Step 303: Specify unicycler as the assembly tool to assemble the quality-controlled data and generate contigs files; Step 304: Use the QUAST tool to assess the quality of the assembly results and generate a quality report in the contigs file; Step 305: Call the MetaBAT2 and CONCOCT tools of the Metawrapp software to perform parallel binning. The results generated by each binning tool are stored in their respective folders. Step 306: Use the bin_refinement function of the metawrapp software to integrate the outputs of multiple binning tools, set the integrity threshold to ≥70% and the contamination threshold to ≤10% to generate an optimized binned genome; Step 307: Use CheckM2 software to evaluate the integrity, contamination, and other indicators of the optimized binned genome, see Table 3; Table 3. Basic parameters for high-quality binning of second-generation and third-generation sequencing data (integrity >90%, contamination <5%). Step 308: Use CoverM to calculate the coverage and relative abundance of the genome in each bin. See the visualization results below. Figure 4 , Figure 4 Different colored sector regions represent different binned genomes, and the area of ​​each sector represents the proportion of the relative abundance or coverage of each binned genome in the entire microbial community. Step 309: Use GTDB-Tk to perform species classification annotation on the optimized binned genome; use Prokka to perform gene prediction and functional annotation; use EggNOG-mapper software to perform functional enrichment analysis on the Prokka output results. Step 310: Integrate the results data, such as quality assessment, species annotation, coverage and relative abundance data, to generate a comprehensive report.

[0066] In summary, the automated analysis method for metagenomic data from sewer mud provided in this embodiment integrates all core steps from raw data quality control, sequence assembly, multi-tool binning optimization to species functional annotation and report generation by constructing a fully automated and intelligently adaptable analysis pipeline. Furthermore, it dynamically selects the optimal tools and parameters based on the characteristics of second-generation, third-generation, and mixed sequencing data. This significantly reduces the professional threshold and manual operation costs of metagenomic data analysis while ensuring accurate and reliable analysis results, and substantially improves data processing efficiency and result reproducibility. It provides an efficient and convenient integrated solution for research and applications in related fields.

[0067] Based on the above technical solution, this embodiment also proposes an automated analysis system for pit mud metagenomic data, used to implement the automated analysis method for pit mud metagenomic data described in the embodiment. Please refer to [link to relevant documentation]. Figure 5 The system includes: The parameter parsing module is used to receive the sequencing data type, file path, and analysis parameters input by the user; the sequencing data type includes Illumina data, Nanopore data, or a mixture of both. The quality control module is used to perform quality control on the raw sequencing data from the gene sequencer according to the sequencing data type, so as to obtain high-quality sequencing data. An assembly execution module is used to assemble the high-quality sequencing data into contigs files. An assembly quality assessment module is used to assess the assembly quality of the contigs files. A multi-tool binning module is used to perform genome binning on the contigs file using at least two binning tools; The binning optimization module is used to integrate the output results of the multi-tool binning module and optimize them based on preset integrity thresholds and contamination thresholds to obtain an optimized binned genome dataset. The binning quality assessment module is used to evaluate the integrity, contamination, and strain heterogeneity of the optimized binning genome; The coverage calculation module is used to calculate the coverage and relative abundance of each optimized bin genome; The annotation module is used to perform species classification annotation and functional annotation on the optimized binned genome; The report generation module is used to integrate the result data output by the assembly quality assessment module, the boxing quality assessment module, the coverage calculation module, and the annotation module to generate an analysis report.

[0068] It is understood that since the automated analysis system for metagenomic data of pit mud described in this embodiment is a system for implementing the automated analysis method for metagenomic data of pit mud described in the embodiment, the system disclosed in the embodiment is relatively simple to describe because it corresponds to the method disclosed in the embodiment. For relevant parts, please refer to the description of the method, and it will not be repeated here.

Claims

1. An automated method for analyzing metagenomic data of pit mud, characterized in that, The method includes: Step 1: Receive the sequencing data type, file path, and analysis parameters input by the user; the sequencing data type includes Illumina data, Nanopore data, or a mixture of both. Step 2: Perform quality control on the raw sequencing data from the gene sequencer according to the sequencing data type to obtain high-quality sequencing data; Step 3: Perform sequence assembly on the high-quality sequencing data to generate contigs files; Step 4: Assess the assembly quality of the contigs files; Step 5: Perform genome binning on the contigs file using at least two binning tools; Step 6: Integrate the output results of at least two binning tools described in Step 5, and optimize them based on preset integrity thresholds and contamination thresholds to obtain an optimized binned genome dataset; Step 7: Evaluate the integrity, contamination level, and strain heterogeneity of the optimized binned genome; Step 8: Calculate the coverage and relative abundance of the genome in each optimized bin; Step 9: Perform species classification and functional annotation on the optimized binned genome; Step 10: Integrate the results data from Steps 4, 7, 8, and 9 to generate an analysis report.

2. The automated analysis method for metagenomic data of pit mud according to claim 1, characterized in that, In step 1, the analysis parameters include parameters for dynamically specifying the analysis termination step, which corresponds to the end of step 2, the end of step 3, the end of step 6, or the end of step 10.

3. The automated analysis method for metagenomic data of pit mud according to claim 1, characterized in that, Step 2 includes: When the sequencing data is Illumina data, Trimmomatic is used for filtering, and FastQC is used for quality assessment and visualization. When the sequencing data is Nanopore data, Porechop is used for adapter removal, FiltLong is used for quality filtering, and Nanoplot is used for quality assessment and visualization. When the sequencing data type is mixed data, the quality control process for the Illumina data portion is executed, and the quality control process for the Nanopore data portion is executed.

4. The automated analysis method for metagenomic data of pit mud according to claim 1, characterized in that, Step 3 includes: selecting an assembly tool according to the sequencing data type; For Illumina data, use SPAdes or MEGAHIT for assembly; For Nanopore data, Flye or Canu are selected for assembly; For mixed data, Unicycler or OPERA-MS can be used for assembly.

5. The automated analysis method for metagenomic data of pit mud according to claim 1, characterized in that, In step 5, the at least two binning tools are selected from MetaBAT2, MaxBin2, and CONCOCT; For Illumina data, MetaBAT2, MaxBin2, and CONCOCT are invoked in parallel. For Nanopore data or mixed data, call MetaBAT2 and CONCOCT.

6. The automated analysis method for metagenomic data of pit mud according to claim 1, characterized in that, In step 6, the preset integrity threshold is greater than or equal to 70%, and the contamination threshold is less than or equal to 10%.

7. The automated analysis method for metagenomic data of pit mud according to claim 1, characterized in that, Step 9 specifically includes: Use GTDB-Tk for species classification annotation; Using Prokka for gene prediction and functional annotation; Based on the output of Prokka, functional enrichment analysis was performed using EggNOG-mapper.

8. The automated analysis method for metagenomic data of pit mud according to claim 1, characterized in that, Step 6 also includes: When the optimized binned genome dataset is empty, the original binning results generated in step 5 are output as the final binned genome file. The suffixes of the binned genome files corresponding to the optimized binned genome or the original binned genome results are standardized.

9. The automated analysis method for metagenomic data of pit mud according to claim 8, characterized in that, Step 10 includes: Extract binning identifiers from the binning genome file names corresponding to the optimized binning genome or the original binning results output in step 6, and clean up the normalized suffixes in the binning identifiers; Using the cleaned bin identifier as the association key, merge the result data from steps 4, 7, 8, and 9. The process terminates when any critical data is detected to be missing from the quality assessment data in step 7, the species annotation data in step 9, or the coverage data in step 8.

10. An automated analysis system for metagenomic data of pit mud, characterized in that, The system for implementing the automated analysis method of pit mud metagenomic data as described in any one of claims 1 to 9 comprises: The parameter parsing module is used to receive the sequencing data type, file path, and analysis parameters input by the user; the sequencing data type includes Illumina data, Nanopore data, or a mixture of both. The quality control module is used to perform quality control on the raw sequencing data from the gene sequencer according to the sequencing data type, so as to obtain high-quality sequencing data. An assembly execution module is used to assemble the high-quality sequencing data and generate contigs files. An assembly quality assessment module is used to assess the assembly quality of the contigs files. A multi-tool binning module is used to perform genome binning on the contigs file using at least two binning tools; The binning optimization module is used to integrate the output results of the multi-tool binning module and optimize them based on preset integrity thresholds and contamination thresholds to obtain an optimized binned genome dataset. The binning quality assessment module is used to evaluate the integrity, contamination, and strain heterogeneity of the optimized binning genome; The coverage calculation module is used to calculate the coverage and relative abundance of each optimized bin genome; The annotation module is used to perform species classification annotation and functional annotation on the optimized binned genome; The report generation module is used to integrate the result data output by the assembly quality assessment module, the boxing quality assessment module, the coverage calculation module, and the annotation module to generate an analysis report.