A method and system for analyzing biogeochemical element cycles in metagenomic datasets

By preprocessing and modeling the metagenomic dataset, the problems of long analysis time and high error rate in existing technologies were solved, and fast and flexible biogeochemical element cycle analysis was achieved, improving analysis efficiency and accuracy.

CN119580852BActive Publication Date: 2025-09-26ENVIRONMENT & PLANT PROTECTION INST CHINESE ACADEMY OF TROPICAL AGRI SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411653970.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-19
Publication Date
2025-09-26
Estimated Expiration
2044-11-19

AI Technical Summary

Technical Problem

Existing technologies for analyzing biogeochemical element cycles in metagenomic datasets have problems such as long analysis time, high error rate, complex software dependency, high cost, and high installation requirements. It is also difficult to effectively distinguish element cycle pathways and related microorganisms.

Method used

A biogeochemical element cycle analysis method for a metagenomic dataset is provided, which includes obtaining original experimental data, preprocessing annotation files and sample grouping files, constructing a chemical element cycle analysis model, performing classification and statistical analysis, and visualizing and storing the results.

Benefits of technology

It enables fast, comprehensive and flexible acquisition of microbial-related information, reduces analysis complexity and cost, and improves analysis efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119580852B_ABST
    Figure CN119580852B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for analyzing biogeochemical element cycles in a metagenomic dataset, comprising: obtaining raw experimental data, obtaining annotation files, sample grouping files, and non-redundant genome profile files based on the raw experimental data; preprocessing the annotation files, sample grouping files, and non-redundant genome profile files to obtain a sample dataset; constructing a chemical element cycle analysis model, obtaining a chemical element profile file based on the sample dataset and the chemical element cycle analysis model; classifying and arranging the chemical element profile files according to corresponding experimental processing conditions, performing statistical analysis on each classification and then visualizing it, and storing the statistical analysis results and the visual processing results. When users need to analyze large metagenomic datasets, the system enables users to comprehensively, quickly, and flexibly obtain intuitive information about microorganisms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of bioinformatics, and in particular relates to a method and system for analyzing biogeochemical element cycles of a metagenomic dataset. Background Art

[0002] Detecting genes and associated microorganisms related to biogeochemical cycles is a key step in studying soil element cycling. Analyzing metagenomic datasets for soil element cycling often consumes considerable time, and the cumbersome process can lead to increased analytical errors. Furthermore, while previously developed gene chip methods can simultaneously and quantitatively detect 72 genes involved in the carbon, nitrogen, phosphorus, and sulfur cycles, they are inadequate in distinguishing element cycling pathways from associated microorganisms. Furthermore, users still need to manually construct complex element cycling pathway maps and identify microorganisms with biogeochemical cycle-related genes from tens of thousands of NR (Non-Redundant Protein Database) species annotations. Furthermore, most software used for metagenomic sequencing is customized for Linux systems, requiring high computing performance and facing issues such as high cost and complex software installation dependencies. Summary of the Invention

[0003] To solve the above technical problems, the present invention proposes a method and system for analyzing the biogeochemical element cycles of a metagenomic dataset to solve the problems existing in the above-mentioned prior art.

[0004] To achieve the above objectives, the present invention provides a method for analyzing biogeochemical element cycles of a metagenomic dataset, comprising:

[0005] Obtaining original experimental data, and obtaining an annotation file, a sample grouping file, and a non-redundant genome profile file based on the original experimental data; preprocessing the annotation file, the sample grouping file, and the non-redundant genome profile file to obtain a sample data set;

[0006] Constructing a chemical element cycle analysis model, and obtaining a profile file of the chemical element based on the sample data set and the chemical element cycle analysis model;

[0007] The profile files of chemical elements are classified and sorted according to the corresponding experimental processing conditions, and each classification is statistically analyzed and then visualized, and the statistical analysis results and visualization results are stored.

[0008] Optionally, the process of obtaining the annotation file includes:

[0009] The original experimental data are assembled, gene prediction is performed on the assembled long sequence fragments, and the amino acid sequence is translated; the amino acid sequence is compared with the KEGG database and the NR database to obtain the corresponding KO entries, and the amino acid sequence is functionally annotated and classified based on the KO entries to obtain an annotation file.

[0010] Optionally, the annotation file includes: KO entry profile file, mapping between non-redundant genome and KO and KEGG PATHWAY, and BLAST results between non-redundant genome and NR.

[0011] Optionally, the process of preprocessing the annotation file, the sample grouping file, and the non-redundant genome profile file to obtain a sample data set includes:

[0012] The description column in the KO entry profile file was deleted and the first column was set as the row name; the first two columns in the mapping between the non-redundant genome and KO and KEGG PATHWAY were extracted and duplicate rows were deleted; the domain classification in the BLAST results of the non-redundant genome and NR was deleted; the first column in the non-redundant genome profile file was set as the row name to obtain the sample data set.

[0013] Optionally, the chemical element cycle analysis model includes four cycle analysis branches, corresponding to the four chemical elements of carbon, nitrogen, phosphorus and sulfur.

[0014] Optionally, the process of obtaining a profile file of a chemical element based on the sample data set and the chemical element cycle analysis model includes:

[0015] arranging KO entries related to corresponding chemical element cycles in the sample data set, extracting abundance information based on the KO entries, and merging the abundance information according to the chemical element cycle process to obtain a cycle abundance data set;

[0016] After performing a normality test on the circulating abundance data set, a significance analysis is performed and the analysis results are stored in a folder named Heatmap; wherein the significance analysis method is selected based on the number of species corresponding to the experimental conditions in the original experimental data and the normality test results;

[0017] Check the carbon cycle gene IDs of predefined categories to determine the existing categories; create carbon cycle pathway folders based on the existing categories, count the relevant microorganisms at different taxonomic levels and save them in the corresponding carbon cycle pathway folders;

[0018] β-diversity analysis was performed on the circulating abundance dataset to evaluate the similarities and differences of the microbial communities. The analysis results were visualized and stored in a folder named the same as the visualization method.

[0019] Optionally, the statistical analysis process for each category includes:

[0020] A normality test was performed on each classified data set. If it conformed to the normal distribution, a parametric test was selected for significance analysis. If it did not conform to the normal distribution, a non-parametric test was selected for significance analysis. Based on the results of the normality test and significance analysis, the differential changes in chemical element cycle genes were calculated. Envelope multivariate variance analysis, similarity analysis and multi-reaction envelope procedure were used to analyze the interactions between multiple chemical element cycle processes.

[0021] Optionally, element-related sub-files are constructed under the root folder according to the number of elements. Each sub-file has the same structure and stores statistical analysis results and visualization processing results.

[0022] The present invention also provides a biogeochemical element cycle analysis system for a metagenomic dataset, comprising:

[0023] The data preparation module is used to obtain the original experimental data and obtain the annotation file sample grouping file and non-redundant genome profile file based on the original experimental data;

[0024] A data processing module is used to pre-process the annotation file, the sample grouping file and the non-redundant genome profile file to obtain a sample data set;

[0025] a cycle analysis module, configured to construct a chemical element cycle analysis model and obtain a profile file of the chemical element based on the sample data set and the chemical element cycle analysis model;

[0026] The statistical analysis module is used to classify and organize the profile files of chemical elements according to the corresponding experimental processing conditions, and perform statistical analysis on each classification and then perform visualization processing;

[0027] The storage module is used to store statistical analysis results and visualization processing results according to preset rules.

[0028] The present invention also provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the program is executed by a processor, corresponding steps involved in a method for analyzing the biogeochemical element cycle of a metagenomic dataset are implemented.

[0029] Compared with the prior art, the present invention has the following advantages and technical effects:

[0030] The present invention provides a method for analyzing biogeochemical element cycles in metagenomic datasets. The method first obtains raw experimental data and, based on the raw experimental data, obtains annotation files, sample grouping files, and non-redundant genome profile files. The annotation files, sample grouping files, and non-redundant genome profile files are then preprocessed to obtain a sample dataset. A chemical element cycle analysis model is then constructed, and a chemical element profile file is obtained based on the sample dataset and the chemical element cycle analysis model. Finally, the chemical element profile files are classified and organized according to the corresponding experimental treatment conditions. Each classification is statistically analyzed and then visualized, and the statistical analysis and visualization results are stored. This method enables users to comprehensively, quickly, and flexibly obtain intuitive microbial information when analyzing large metagenomic datasets. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0032] Figure 1 is a flow chart of a method according to an embodiment of the present invention;

[0033] Figure 2 2 is a system structure diagram of an embodiment of the present invention. DETAILED DESCRIPTION

[0034] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0035] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0036] Example 1

[0037] like Figure 1 As shown, this embodiment provides a method for analyzing biogeochemical element cycles of a metagenomic dataset, including:

[0038] Obtaining original experimental data, and obtaining an annotation file, a sample grouping file, and a non-redundant genome profile file based on the original experimental data; preprocessing the annotation file, the sample grouping file, and the non-redundant genome profile file to obtain a sample data set;

[0039] In some embodiments, the process of obtaining the annotation file includes:

[0040] The raw sequencing data were assembled, open reading frames (ORFs) were predicted for the assembled contig sequences (long sequence fragments), and the ORFs were translated into amino acid sequences according to 11 sets of genetic codons (NCBI Genetic Code 11). The amino acid sequences were compared with the KEGG database and the NR database to obtain corresponding KO entries. The amino acid sequences were functionally annotated and classified based on the KO entries to obtain annotation files.

[0041] Furthermore, metagenomic assembly software such as metaSPAdes and MEGAHIT are used to assemble raw reads, which refer to the original experimental data obtained during the sequencing process. These data are usually short fragments of DNA sequences, also known as reads. The length of each read is usually between 50 and 500 base pairs (bp). MetaGeneMark, Prodigal, or Glimmer-MG are then used to predict genes on the assembled long sequence fragments. Finally, the predicted amino acid sequences are compared with the KEGG and NR databases using tools such as KofamScan and DIAMOND to obtain KO numbers to understand the functional categories, thereby achieving functional and taxonomic annotations to obtain annotation files.

[0042] For example, users provide metagenomic dataset files annotated using the KEGG and NR databases, as well as sample data files under different treatments. For example, the translated protein sequence is submitted to KofamScan. KofamScan then compares the protein sequence with the KEGG KO database, identifies matching KO entries, and assigns a corresponding KO number and KEGG pathway information to each protein sequence based on the alignment results, thus generating a dataset.

[0043] In some embodiments, the annotation file includes: a KO entry profile file, a mapping between a non-redundant genome and KO and KEGGPATHWAY, and a BLAST result of a non-redundant genome and NR.

[0044] In some embodiments, the process of preprocessing the annotation file, the sample grouping file, and the non-redundant genome profile file to obtain a sample dataset includes:

[0045] The description column in the KO entry profile file was deleted and the first column was set as the row name; the first two columns in the mapping between the non-redundant genome and KO and KEGG PATHWAY were extracted and duplicate rows were deleted; the domain classification in the BLAST results of the non-redundant genome and NR was deleted; the first column in the non-redundant genome profile file was set as the row name to obtain the sample data set.

[0046] Specifically, before performing element cycle analysis, the sample dataset was preprocessed as follows:

[0047] 1) Delete the "description" column in "ko" and set the first column to row names. Make sure all columns are numeric data types to avoid errors or NA values ​​(missing values) caused by character, factor, or logical data types.

[0048] 2) Eliminate duplicate rows based on the first two columns of "Gene" and extract the first two columns of "Gene".

[0049] 3) Delete the field classification in "tax".

[0050] 4) Set the first column of "abundance" to row names, ensuring that all columns are numeric. Keep the row names as part of the data and set this column as the first column in the dataset.

[0051] "ko" refers to the KO entry profile file, "group" refers to the sample grouping file, and "Gene" relates to the mapping between non-redundant genomes and KOs and KEGG PATHWAY in scattered metagenomics. This involves using annotation tools to functionally annotate gene sequences to obtain KO numbers. Then, by searching the KEGG PATHWAY database, KO numbers are associated with specific biological pathways. Ultimately, a database or table containing mappings between genes, KO numbers, and KEGG PATHWAY is created. "Mapping" ultimately refers to the process of linking gene information to its functional classification and biological pathway in the KEGG database. "Tax" includes BLAST results used to compare non-redundant genomes with NR. First, BLAST tools are used to align sequences in the non-redundant genome with sequences in the NR database to identify similar sequences. Significantly similar sequences are selected from the alignments based on parameters such as E-values ​​(expected values). Finally, based on these similar sequences, sequences in the non-redundant genome are assigned to specific species or taxa using classifiers or phylogenetic methods. "Abundance" contains the non-redundant genome profile file used in metagenomics.

[0052] Constructing a chemical element cycle analysis model, and obtaining a profile file of the chemical element based on the sample data set and the chemical element cycle analysis model;

[0053] In some embodiments, the chemical element cycle analysis model includes four cycle analysis branches, corresponding to the four chemical elements of carbon, nitrogen, phosphorus, and sulfur.

[0054] Specifically, after data preprocessing, the dataset was classified by elemental cycles, and profile files were generated using independent functions in four independent element modules: "Ccyc.Abundance", "Ncyc.Abundance", "Pcyc.Abundance", and "Scyc.Abundance". In addition, a set of functions for several cycling processes was included, including 20 KEGG positively selected (KO) entries related to seven different carbon cycle processes, 48 ​​KO entries related to 18 nitrogen cycle processes, 5 KO entries related to two phosphorus cycle processes, and 48 KO entries related to 15 sulfur cycle processes. Through these, functional genes and microorganisms related to the four elements were extracted. The four independent analysis modules include: carbon cycle, nitrogen cycle, phosphorus cycle, and sulfur cycle. Within each module, four different functional analyses are provided, including the organization and differential analysis of genes related to biogeochemical cycles, analysis of microorganisms with biogeochemical cycle-related genes, β-diversity analysis, and data visualization.

[0055] In some embodiments, the process of obtaining a profile file of a chemical element based on the sample data set and the chemical element cycle analysis model includes:

[0056] arranging KO entries related to corresponding chemical element cycles in the sample data set, extracting abundance information based on the KO entries, and merging the abundance information according to the chemical element cycle process to obtain a cycle abundance data set;

[0057] After performing a normality test on the circulating abundance data set, a significance analysis is performed and the analysis results are stored in a folder named Heatmap; wherein the significance analysis method is selected based on the number of species corresponding to the experimental conditions in the original experimental data and the normality test results;

[0058] Check the carbon cycle gene IDs of predefined categories to determine the existing categories; create carbon cycle pathway folders based on the existing categories, count the relevant microorganisms at different taxonomic levels and save them in the corresponding carbon cycle pathway folders;

[0059] β-diversity analysis was performed on the circulating abundance dataset to evaluate the similarities and differences of the microbial communities. The analysis results were visualized and stored in a folder named the same as the visualization method.

[0060] For example, first create four geochemical cycle folders Carbon, Nitrogen, Phosphorus, Sulfur, etc. through the dir.create() function, taking the carbon cycle as an example:

[0061] 1. Arrangement of geochemical cycle-related genes: First, arrange the sub-data related to the C cycle, such as KO number, and then extract the abundance information corresponding to the KO number in the carbon cycle through the function C.ko<-ko[rownames(ko)%in%C,], and then merge the KO abundance information according to the seven carbon cycle processes collected from the kegg database to obtain the carbon cycle abundance data set. Finally, save it as a txt file of KO abundance through the following function:

[0062] write.table(C.abundance,file="Results / Carbon / Gene / Abundance / C_cycle_gene_abun.txt", sep="\t", row.names=FALSE).

[0063] 2. Analysis of differences in genes related to biogeochemical cycles: Based on the carbon cycle abundance dataset, we first performed a Shapiro-Wilk normality test using the shapiro.test() function in the R language. Then, depending on the situation, we used parametric or non-parametric tests to determine if the distribution conformed to the normal distribution. Depending on the number of groups, for two groups, we used the t.test() / wilcox.test() functions in the R language stats package to perform a T test / Wilcoxon test. For more than two groups, we used the aov() / kruskal.test() functions in the stats package to perform an analysis of variance / Kruskal-Wallis test. Finally, the difference results were saved to the "Heatmap" folder using the following function:

[0064] write.table(result[[1]],"Results / Carbon / Gene / Heatmap / Diff_C_gene_test.txt", sep="\t", row.names=FALSE).

[0065] 3. Analysis of microorganisms with biogeochemical cycle-related genes: Use the function if(rowSums(C.abundance[,:ncol(C.abundance)])>0&Ccyc.h[]==1){ to check whether the input data contains carbon cycle gene IDs from predefined categories. A binary vector is returned to indicate whether each category exists (1) or not (0). Then, carbon cycle pathway (ACF, AR, ANCF, etc.) folders are created through dir.create. The number of folders is determined by the chemical element categories present in the sample. The relevant microorganisms at each level of kingdom, phylum, class, order, family, genus and species are counted through the corresponding community (Gene, tax, abundance, group) and saved in the corresponding subfolder.

[0066] 4. β-Diversity Analysis and Data Visualization: The carbon cycling microbial community abundances compiled in the previous step were combined with the vegdist function in the R package vegan to generate a Bray-Curtis matrix. By calculating the Bray-Curtis distance between samples, we can visually understand the similarities and differences between the microbial communities. Based on the Bray-Curtis distance matrix, we then used the Vegan package to convert the results into PCoA, PCA, and NMDS analysis results. These files are saved in a subfolder named "Distance." This folder also contains visualizations of the PCoA, PCA, and NMDS results, stored in subfolders named "PCoA," "PCA," and "NMDS," respectively.

[0067] For data visualization, custom-designed graphics using R code were used to plot elemental cycling patterns, vividly illustrating the differences in elemental cycling pathways between groups. Furthermore, PCA, PCoA, and NMDS plots were created for elemental cycling genes, graphically representing the beta diversity of these genes within each sample. Finally, heat maps of the abundance of microorganisms involved in each elemental cycling process at each taxonomic level were created, allowing users to quickly access relevant information about these microorganisms.

[0068] The profile files of chemical elements are classified and sorted according to the corresponding experimental processing conditions, and each classification is statistically analyzed and then visualized, and the statistical analysis results and visualization results are stored.

[0069] In some embodiments, the process of performing statistical analysis on each classification includes:

[0070] A normality test was performed on each classified data set. If it conformed to the normal distribution, a parametric test was selected for significance analysis. If it did not conform to the normal distribution, a non-parametric test was selected for significance analysis. Based on the results of the normality test and significance analysis, the differential changes in chemical element cycle genes were calculated. Envelope multivariate variance analysis, similarity analysis and multi-reaction envelope procedure were used to analyze the interactions between multiple chemical element cycle processes.

[0071] Specifically, after classifying the functional genes and microbial datasets under different treatments (i.e., different experimental conditions), the Shapiro-Wilk normality test was first performed using the shapiro.test() function in the R package. The statistic W was calculated based on the difference between the observed value of the data and the expected value of the normal distribution. Then, based on the significance level (α = 0.05), the critical value was calculated. If the W statistic is greater than the critical value, the null hypothesis cannot be rejected, indicating that the data may follow a normal distribution. If the W statistic is less than the critical value, the null hypothesis is rejected, indicating that the data do not follow a normal distribution.

[0072] Then, significance analysis was performed based on whether the data conformed to a normal distribution. For normal distributions, parametric tests were used; for non-parametric tests, non-parametric tests were used. The appropriate method was automatically selected based on the number of treatment groups. For two groups, the T-test / Wilcoxon test was performed using the t.test() / wilcox.test() functions in the R language stats package. For more than two groups, the ANOVA / Kruskal-Wallis test was performed using the aov() / kruskal.test() functions in the R language stats package. Next, differential changes in element cycling genes were calculated using either the T-test and Wilcoxon test or the ANOVA and Kruskal-Wallis test for differential gene expression analysis. Furthermore, multivariate statistical analysis was supported for all classified data. Considering interactions between multiple element cycling processes, methods such as envelope multivariate analysis of variance (Adonis), analysis of similarity (ANOSIM), and the multiple reaction envelope procedure (MRPP) were used, with significance analysis based on distance matrices. All statistical results were output in tabular form. The envelope multivariate analysis of variance (Adonis) is used to test whether two or more categorical factors significantly influence a multivariate response variable. This analysis is performed using the Adonis function in the vegan package. It takes a community data matrix as input and a factor vector as the sample grouping to analyze the impact of different groupings on the community data. Analysis of similarity (ANOSIM) focuses on the significance of differences between and within groups, calculating the Bray-Curtis distance between each pair of samples and calculating R and P values. The R value measures the relative magnitude of differences between and within groups, while the P value determines the significance of differences. The multiple reaction envelope procedure (MRPP) aims to search for groupings with the smallest intragroup distance, regardless of intergroup distance. It uses permutations to uniformly group all microbial communities into various possible combinations, constructs a statistic, δ, and then calculates the value of the statistic for each grouping and analyzes its distribution.

[0073] Specifically, the statistical results above were calculated and visualized. Heatmaps were used to display the abundance and fold-change values ​​of elemental cycling processes within each independent group. Elemental cycling pattern diagrams were used to demonstrate differences in elemental cycling pathways between treatments, as well as changes in pathway abundance. β-diversity results were output as PCA, PCoA, and NMDS plots of elemental cycling genes. Microbial abundance heatmaps were used to visualize the involvement of each taxonomic level in each elemental cycling process.

[0074] In some embodiments, element-related sub-files are constructed under the root folder according to the number of elements, and each sub-file has the same structure and stores statistical analysis results and visualization processing results.

[0075] The generated folders and files are in a specific structure in the final output directory:

[0076] "Result": This is the root folder, which contains four subfolders: "Carbon", "Nitrogen", "Phosphorus" and "Sulfur".

[0077] Each subfolder contains results related to the element cycle process and is divided into three main categories:

[0078] 1) "Beta diversity": This folder contains Bray-Curtis distance matrices for elemental cycling genes and the results of a multivariate analysis of variance (MANOVA) based on the Bray-Curtis distance matrix. These files are stored in a subfolder called "Distances." This folder also contains visualizations of PCoA, PCA, and NMDS, stored in subfolders called "PCoA," "PCA," and "NMDS," respectively.

[0079] 2) "Gene": This folder contains an overview of elemental cycling genes, located in a subfolder named "Abundance." It also contains line-of-sight analysis results and carbon cycle model diagrams, stored in a subfolder named "Cycle image." Additionally, this folder contains heatmap displays and variance analysis results, located in subfolders named "Heatmap."

[0080] 3) Host_relative_Group: This folder contains seven subfolders, each named after a specific elemental cycling process. Each subfolder contains a table and heat map showing the relative abundance of microorganisms involved in the corresponding elemental cycling process at different taxonomic levels (phylum, class, order, family, genus, and species).

[0081] like Figure 2 As shown, this embodiment provides a biogeochemical element cycle analysis system for a metagenomic dataset, comprising:

[0082] The data preparation module is used to obtain the original experimental data and obtain the annotation file sample grouping file and non-redundant genome profile file based on the original experimental data;

[0083] A data processing module is used to pre-process the annotation file, the sample grouping file and the non-redundant genome profile file to obtain a sample data set;

[0084] a cycle analysis module, configured to construct a chemical element cycle analysis model and obtain a profile file of the chemical element based on the sample data set and the chemical element cycle analysis model;

[0085] The statistical analysis module is used to classify and organize the profile files of chemical elements according to the corresponding experimental processing conditions, and perform statistical analysis on each classification and then perform visualization processing;

[0086] The storage module is used to store statistical analysis results and visualization processing results according to preset rules.

[0087] This embodiment also provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the program is executed by a processor, corresponding steps involved in a method for analyzing biogeochemical element cycles of a metagenomic dataset are implemented.

[0088] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A method for analyzing biogeochemical element cycles in a metagenomic dataset, characterized in that: include: Obtain original experimental data, and obtain annotation files, sample grouping files, and non-redundant genome profile files based on the original experimental data; Preprocessing the annotation file, the sample grouping file, and the non-redundant genome profile file to obtain a sample data set; The process of obtaining annotation files includes: Assembling the raw experimental data, performing gene prediction on the assembled long sequence fragments, and translating them to obtain amino acid sequences; comparing the amino acid sequences with the KEGG database and the NR database to obtain corresponding KO entries, and performing functional annotation and classification annotation on the amino acid sequences based on the KO entries to obtain annotation files; Constructing a chemical element cycle analysis model, and obtaining a profile file of the chemical element based on the sample data set and the chemical element cycle analysis model; The chemical element cycle analysis model includes four cycle analysis branches, corresponding to the four chemical elements of carbon, nitrogen, phosphorus and sulfur respectively; The process of obtaining a profile file of a chemical element based on the sample data set and the chemical element cycle analysis model includes: arranging KO entries related to corresponding chemical element cycles in the sample data set, extracting abundance information based on the KO entries, and merging the abundance information according to the chemical element cycle process to obtain a cycle abundance data set; After performing a normality test on the circulating abundance data set, a significance analysis is performed and the analysis results are stored in a folder named Heatmap; wherein the significance analysis method is selected based on the number of species corresponding to the experimental conditions in the original experimental data and the normality test results; Check the carbon cycle gene IDs of predefined categories to determine the existing categories; create carbon cycle pathway folders based on the existing categories, count the relevant microorganisms at different taxonomic levels and save them in the corresponding carbon cycle pathway folders; Perform β-diversity analysis on the circulating abundance dataset to assess the similarities and differences of the microbial communities, visualize the analysis results, and store the visualization results in a folder named the same as the visualization method; The profile files of chemical elements are classified and sorted according to the corresponding experimental processing conditions, and each classification is statistically analyzed and then visualized, and the statistical analysis results and visualization results are stored.

2. The method for analyzing biogeochemical element cycles of a metagenomic dataset according to claim 1, characterized in that: The annotation files include: KO entry profile file, mapping between non-redundant genome and KO and KEGG PATHWAY, and BLAST results between non-redundant genome and NR.

3. The method for analyzing biogeochemical element cycles of a metagenomic dataset according to claim 2, characterized in that: The process of preprocessing the annotation file, the sample grouping file, and the non-redundant genome profile file to obtain a sample data set includes: The description column in the KO entry profile file was deleted and the first column was set as the row name; the first two columns in the mapping between the non-redundant genome and KO and KEGG PATHWAY were extracted and duplicate rows were deleted; the domain classification in the BLAST results of the non-redundant genome and NR was deleted; the first column in the non-redundant genome profile file was set as the row name to obtain the sample data set.

4. The method for analyzing biogeochemical element cycles of a metagenomic dataset according to claim 1, wherein: The process of statistical analysis for each category includes: A normality test was performed on each classified data set. If it conformed to the normal distribution, a parametric test was selected for significance analysis. If it did not conform to the normal distribution, a non-parametric test was selected for significance analysis. Based on the results of the normality test and significance analysis, the differential changes in chemical element cycle genes were calculated. Envelope multivariate variance analysis, similarity analysis and multi-reaction envelope procedure were used to analyze the interactions between multiple chemical element cycle processes.

5. The method for analyzing biogeochemical element cycles of a metagenomic dataset according to claim 1, characterized in that: Element-related sub-files are constructed under the root folder according to the number of elements. Each sub-file has the same structure and stores statistical analysis results and visualization processing results.

6. A biogeochemical element cycle analysis system for a metagenomic dataset, for implementing the biogeochemical element cycle analysis method for a metagenomic dataset according to any one of claims 1 to 5, characterized in that: The data preparation module is used to obtain the original experimental data and obtain the annotation file sample grouping file and non-redundant genome profile file based on the original experimental data; A data processing module is used to pre-process the annotation file, the sample grouping file and the non-redundant genome profile file to obtain a sample data set; a cycle analysis module, configured to construct a chemical element cycle analysis model and obtain a profile file of the chemical element based on the sample data set and the chemical element cycle analysis model; The statistical analysis module is used to classify and organize the profile files of chemical elements according to the corresponding experimental processing conditions, and perform statistical analysis on each classification and then perform visualization processing; The storage module is used to store statistical analysis results and visualization processing results according to preset rules.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the corresponding steps involved in the method for analyzing the biogeochemical element cycle of a metagenomic dataset as claimed in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Method for acquiring relative abundance and activity of whole-course ammoxidation microorganisms from complex environment system based on macro omics technology

    CN112951330A

  • Method for analyzing nitrogen restriction ecosystem by combining isotopic tracing and metagenome

    CN117116341A