Method for analyzing functional information of marine microorganisms in marine samples

By constructing a species segmentation database and combining multiple search algorithms, the problem of false positives caused by the excessive size of the marine microbial metaproteome database was solved, and efficient and accurate functional information analysis of marine microorganisms was achieved.

CN116312820BActive Publication Date: 2026-03-31DALIAN INSTITUTE OF CHEMICAL PHYSICS CHINESE ACADEMY OF SCIENCES
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-07
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

The excessive size of the marine microbial metaproteome database leads to a high false positive rate, making it impossible to effectively analyze the functional information of marine microorganisms.

Method used

We constructed a database integration strategy, split the database according to species, combined multiple search algorithms for qualitative and quantitative proteomics analysis and functional analysis, used tools such as BLASTP, KEGG, and GO for protein sequence functional annotation, and combined sample-specific databases for data processing.

Benefits of technology

It improves the accuracy and coverage of protein identification, shortens the database search time, reduces false positive results, and enables precise analysis of the functions of marine microorganisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116312820B_ABST
    Figure CN116312820B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of functional information analysis method of marine microorganism in marine sample, based on liquid mass spectrometry obtains mass spectrum RAW file, utilize the peptide sequence obtained by macrogenomics sequencing, the database obtained by self-defined configuration or the database published in public is combined as macro proteome data search database.Combined with multiple sources of databases, such as databases published in public, databases obtained by self-defined configuration, and peptide sequences obtained by macrogenomics sequencing, as macro proteome data search database.The database is split according to the classification level of species, and the sub-database is reduced by iterative search method, and then the sub-database is combined, and the qualitative and quantitative analysis of macro proteome is carried out by using proteome search analysis software, and the function annotation software is used for function annotation based on public database.The proteome quantitative analysis software screens different environmental difference proteins, and deeply excavates the community composition and metabolic activity difference of different environmental microorganisms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for analyzing the functional information of marine microorganisms in marine samples. It is a method that integrates metagenomics and public databases, splits and optimizes databases according to species, performs qualitative and quantitative proteomic analysis, and analyzes metabolic activities and functions. Background Technology

[0002] Marine microorganisms play a vital role in biogeochemical cycles and in regulating global climate. Metaproteomics are crucial for revealing the functional mechanisms of microorganisms in marine ecosystems. Due to the complexity of marine microorganisms, less than 1% are culturable, and the protein sequences of many microorganisms are unknown. Metaproteomics databases are very large but incomplete. Therefore, the selection of databases and how to accommodate large datasets are essential for the accurate interpretation of metaproteomics data. There are three main methods for constructing metaproteomics databases: (1) high-throughput metagenomic sequencing or transcriptome databases; (2) public databases; and (3) metagenomic databases constructed from relevant species. Recently, the TARA oceans project collected 187 metatranscriptome and 370 metagenomic samples from 126 sampling sites worldwide, performed genome and transcriptome sequencing, and obtained the most comprehensive marine microbial reference gene set (OM-RGC.v2). The Malaspina expedition further extended the research depth to approximately 4,000 meters. The increasingly complete marine metagenomic datasets are of great significance for the application and development of metaproteomics. In the data-dependent acquisition (DDA) workflow, spectrogram matching (PSM) analysis is central to protein identification. PSM algorithm applications, such as SEQUEST and X! Tandem, have been used in metaproteomics research, but are not suitable for large datasets. Databases built for marine metaproteomics research typically contain millions of samples, which can impact the sensitivity of protein identification. Multi-step or cascaded searches of large databases can be utilized. Using multiple search algorithms can effectively increase the percentage of peptide matching. Applying ab initio search algorithms and spectrogram libraries can also improve the efficiency of peptide identification from metaproteomics samples. Developing precise analytical methods for protein data from marine microorganisms, especially complex in-situ marine microorganisms, is crucial for studying the function of marine microorganisms and their role in the marine environment at the molecular level. The composition and capacity of databases severely limit the precise analysis of marine metaproteomics data. Constructing a data analysis workflow for marine microbial proteins, including database assembly, species-based database splitting, qualitative and quantitative analysis of metaproteomics, and metabolic activity analysis, can improve the depth of protein identification, shorten lengthy database search times, and enable in-depth mining of marine microbial proteomics data. Summary of the Invention

[0003] To address the aforementioned problems, the present invention aims to provide a data processing method that integrates the construction and species-specific database search strategy, qualitative and quantitative analysis of macroproteomics, and species and functional analysis. The entire process is simple to operate, improves search efficiency while increasing the number of identified proteins, and effectively avoids the problems of high false positive rates or even inability to perform searches due to excessively large microbial databases. This method can achieve highly efficient and accurate analysis of marine microbial macroproteome samples.

[0004] To achieve this objective, the technical solution of the present invention is as follows:

[0005] 1. A certain volume of seawater is filtered onto a 0.22 μm membrane. Part of this is used for proteomics pretreatment, and the other part is used for deep sequencing of metagenomics or metatranscriptomics, translating it into protein sequences as part of a metaproteomics database. Based on the characteristics of the actual sample, download relatively complete databases such as uniprot, rn, or ensembl. Alternatively, previously published databases can be cited, such as the Global Marine Microbial Reference Gene Catalogue (OM-RGC) and the Malaspina Gene Database (M-GeneDB). Integrate these databases from multiple sources to obtain the most comprehensive database possible.

[0006] 2. For protein sequences without species annotations, BLASTP annotations were used based on the NCBInr database. The annotation results were then imported into MEGAN, and the lowest common ancestor algorithm was used for species attribution. The databases were then merged according to taxonomic levels such as class and order. If the number of sequences at a certain taxonomic level was large, they could be merged again within the same taxonomic level; otherwise, they could be split within the same taxonomic level, ultimately resulting in a database of appropriate size and number.

[0007] 3. Since the database size is still large after the metaproteome data is split, the MetaPro_IQ or SpectraCluster strategy is selected to reduce the database size, resulting in a sample-specific, non-redundant, small-capacity database. Software such as Pfind, Maxquant, or Proteome Discoverer is used for qualitative and quantitative analysis of the metaproteome.

[0008] 4. Functional annotation of protein sequences can be performed using BLASTP based on orthologous homology clusters (COGs), the Genome Encyclopedia (KEGG), and the Gene Ontology (GO). Additionally, functional annotation can also be performed using Diamond based on the EggNOG database.

[0009] 5. Using Persus software, a T-test was performed on the quantified proteins to identify differentially expressed proteins with a P-value < 0.05 and an abundance difference greater than 1.2 or 3-fold between different samples. The functions of these upregulated or downregulated proteins and their involvement in the KEGG pathway were analyzed. Sample-specific proteins were screened, and based on their metabolic activity and the biological processes they participated in, the important roles of microorganisms in the sample's environment were explored.

[0010] 6. After obtaining macroproteome data, based on known protein identification and functional information, and corresponding to different microbial species, and combined with protein abundance and environmental parameters, we can search for possible metabolic characteristics of microorganisms, further expand the interaction between microbial communities and the environment, and reveal the interaction between microbial communities and the environment.

[0011] 7. The developed data analysis method for marine microbial proteomics can be used for mixed samples of various bacteria and for in-situ analysis of marine macroprotein data. It integrates database construction, qualitative and quantitative analysis, and species and functional analysis, providing a highly accurate method for the precise analysis of marine macroproteinomics data.

[0012] The present invention has the following advantages:

[0013] 1. By selecting appropriate public databases, customizing databases, building databases based on the sample's own genomic or transcriptomic data, or integrating published databases with environmental conditions similar to those of the actual samples, and combining these with databases obtained from metagenomic sequencing that closely resemble the actual samples, the accuracy and coverage of metaprotein identification in marine samples from actual sites can be effectively improved. 2. Large-capacity databases are rationally divided according to species, and iterative search methods are employed to address the impact of excessively large databases.

[0014] This addresses the issue of protein identification, effectively shortening search time, improving search efficiency, increasing the number and accuracy of protein identifications, and reducing false positive results.

[0015] 3. Identify and annotate proteins to perform functional annotation, analyze species functional information, and combine sample environmental conditions to achieve functional analysis of microbial communities in samples from different environments.

[0016] 4. The system integrates database construction, data segmentation by species, macroprotein data search, functional annotation and bioinformatics analysis, and establishes a marine microbial data analysis process, enabling accurate analysis of marine microbial mixed samples and marine in-situ microbial macroprotein data. Attached Figure Description

[0017] Figure 1 A schematic diagram of the data processing workflow for marine microbial macroproteins.

[0018] Figure 2 Schematic diagram of marine station function analysis. Detailed Implementation

[0019] Example 1

[0020] Seawater from the euphotic zone at station S55 in the South China Sea, at a depth of 5m, was filtered onto a polycarbonate membrane. The membrane was then cut into small pieces using sterile scissors. 1g of the membrane was taken from 10L of water and added to 1ml of 1-dodecyl-3-methylimidazolium chloride ionic liquid. The mixture was prepared by sonicating guanidine hydrochloride (C12Im-Cl) in a water bath for 30 min, followed by sonication on ice for 30 min, centrifugation at 30000xg for 40 min, adding tris(2-carboxyethyl)phosphine to a final concentration of 50 mM and reacting in a 95°C water bath for 30 min. Silicon spheres with surface covalently bonded iodoacetic acid-N-succinamide ester (from patent CN106632877B, prepared in Example 1) were added and reacted overnight with shaking. The supernatant was discarded after centrifugation, and the surface of the silicon spheres was washed with 50% methanol (volume ratio 1:1) and 50 mM ammonium bicarbonate solution, respectively. Finally, 2 μg of trypsin was added and incubated at 37°C for 18 h. After centrifugation and collection of the supernatant, mass spectrometry data were acquired using liquid chromatography-mass spectrometry (LC-MS) to obtain a raw file, with the mass analyzer being an electrostatic field orbital trap.

[0021] Methods for analyzing the functional information of microorganisms based on preprocessed raw files, such as... Figure 1As shown, a suitable metagenomic database was constructed for samples collected at station S55 in the South China Sea at a depth of 5m. This included a metagenomic database obtained from genome sequencing, translated into a protein sequence database, and a downloaded global marine microbial internal reference gene catalog (OM-RGC, containing 40 million non-redundant protein sequences from viruses, prokaryotes, and microeukaryotes). This resulted in a large-capacity database containing 43.2 million protein sequences and species information. Based on the species information in the database, it was split into eight smaller databases according to phylum classification. The mass spectrometry raw files and the eight split protein databases were imported into Metalab software for iterative search (MetaPro_IQ) to reduce the database to smaller databases. These smaller databases were then integrated. Based on the database constructed in this way, protein quantification was performed on the raw mass spectrometry data using Pfind software. A total of 4960 protein sequences were identified. Without using the iterative database search, the search would have taken several days and only identified 2266 protein sequences. Based on the EggNOG database and using Diamond for functional annotation, combined with the corresponding species abundance of proteins, the identified marine macroproteins mainly belong to Actinobacteria, Proteobacteria, and Cyanobacteria. Their metabolic activities primarily include substance transport and photosynthesis. A key enzyme in photosynthesis, ribulose-1,5-bisphosphate carboxylase / oxygenase (RubisCO), was also identified. RubisCO catalyzes the first major carbon fixation reaction in the Calvin cycle of photosynthesis, converting atmospheric carbon dioxide into energy storage molecules in organisms, such as sucrose. Ribulose-1,5-bisphosphate carboxylase / oxygenase catalyzes the carboxylation of ribulose-1,5-bisphosphate with carbon dioxide or its oxidation with oxygen. Simultaneously, RuBisCO also enables RuBP to enter the photorespiration pathway.

[0022] Example 2

[0023] Preprocessing of metagenomic samples from the S55 subsurface at a depth of 3000m in the South China Sea was performed using the same procedure as in Example 1 above. Raw files were obtained using liquid chromatography-mass spectrometry (LC-MS), with an electrostatic field orbital trap used as the mass analyzer. Based on metagenomic databases and a Malaspina voyage with an average sampling depth of 3731m, the Malaspina gene database (M-GeneDB) was obtained. The database was split according to microbial taxonomic levels, resulting in 48 smaller databases. These 48 databases were further merged into 12 databases. The mass spectrometry raw files and the 12 split protein databases were imported into Metalab software. Spectra Clustering was used to obtain a reduced-capacity database, and the built-in Maxquant software was used for protein quantification, resulting in the quantification of 459 protein sequences. Based on the GO, COGs, and KEGG databases, BLASTP functional annotation revealed that Actinobacteria, Euryarchaeota, and Planctomycete were dominant, and enzymes related to dark carbon fixation were identified in the functional analysis.

[0024] Example 3

[0025] A mixture of marine bacterial samples, including Bacillus infantis, Bacillus pacificus, Pseudomonas abyssi, and Erythrobactermarinus, was added to 0.4 ml of 1-dodecyl-3-methylimida zolium chloride (C12Im-Cl), with a volume ratio of 6M guanidine hydrochloride to C12Im-Cl of 0.3:1.7. The mixture was placed on ice and sonicated for 10 min, then centrifuged at 16000 x g for 40 min. Other pretreatment processes were the same as in Example 1. After pretreatment, data were collected using liquid chromatography-mass spectrometry (LC-MS) to obtain raw files, with a time-of-flight (TOF) tube as the mass analyzer. Since the species were known, the database construction step could be omitted. A proteome sequence database containing functional information was directly downloaded from the Uniprot website. Metalab software was used to reduce the database size, and the built-in Proteome Discoverer software was used for identification and quantification. The database had complete functional information, eliminating the need for functional annotation of the identified proteins. A total of 5956 proteins were identified, of which 2459 belonged to *Bacillus infantis*, 1852 to *Bacillus pacificus*, 1195 to *Pseudomonas abyssi*, and 450 to *Erythrobacter marinus*. *Bacillus infantis* proteins were expressed in large numbers and at high abundance, exhibiting active metabolic activity and primarily involved in the transport of substances such as urea. *Erythrobacter marinus* proteins were expressed less frequently and were less competitive in mixed samples.

[0026] Example 4

[0027] The macroproteome samples from sites E1 (0m, 3000m), F2 (0m, 3500m), DC6 (0m, 2000m), and SEATS (0m, 3000m) in the South China Sea were pretreated. The specific pretreatment process was the same as in Example 1. Raw files were obtained by liquid chromatography-mass spectrometry (LC-MS), and the mass analyzer was a Fourier transform ion cyclotron resonance mass analyzer (FT-ICR). Based on metagenomic databases, M-GeneDB, and OM-RGC, the database was split according to phylum and taxonomic level, resulting in 25 smaller databases. These 25 smaller databases varied significantly in size. Several smaller databases with fewer protein sequences were merged. Larger databases, such as those in the phylum Proteobacteria, were further split according to the next lower class and taxonomic level, ultimately yielding 15 databases with suitable species and sizes. Iterative database searching was performed using Metalab software (MetaPro_IQ) to further obtain smaller databases. After merging the smaller databases, proteomic quantification was performed using Maxquant software. The obtained protein sequences were functionally annotated using BLASTP. Proteobacteria showed high abundance at all sites, as shown in the KO heatmap distribution at each site. Figure 2 As shown, KO abundance varied significantly across different sites and depths, revealing that microorganisms play different biological functions in different environments. Multiple enzymes involved in carbon and nitrogen cycling were discovered, demonstrating the crucial role of marine microorganisms in these cycles. Using Persus software, differentially expressed proteins were identified across different environments, revealing key metabolic proteins at each site. Surface Kelvin cycling activity was high, while deep-sea nitrogen cycling was active.

Claims

1. A method for analyzing functional information of marine microorganisms in a marine sample, characterized by: 1) collecting macroproteome data of the marine sample by liquid chromatography-mass spectrometry, collecting macroproteome data based on the liquid chromatography-mass spectrometry system, and generating a RAW file of the macroproteome; translating macrogenomic information obtained by macrogenomic sequencing of the marine sample into macroproteome sequences, and supplementing them with known databases in the prior art to obtain a macroproteome sequence database, including sequence information and their respective species information; 2) splitting the protein sequences obtained above into several sub-databases according to the biological classification of microorganisms, and using these sub-databases to perform an iterative search or spectral clustering data retrieval process using the RAW file to obtain a marine sample-specific exclusive database of the macroproteome, including macroproteome sequence information and their respective species information; 3) performing qualitative and quantitative analysis based on the exclusive database using protein retrieval software with the RAW file; 4) annotating the protein sequences based on the known functional database in the prior art using annotation software to obtain functional information of the marine sample. The mass spectrometer in the liquid chromatography-mass spectrometry includes one or more than two of an electrostatic field orbitrap, a time-of-flight tube (TOF), or a Fourier transform ion cyclotron resonance mass analyzer (FT-ICR). The obtained detailed macroproteome database is split into several sub-databases to reduce the library capacity of a single search; the splitting method splits the protein database according to the corresponding taxonomic characteristics of microorganisms at one or more than two classification levels of phylum, class, order, and family, and the obtained split databases are directly used for data retrieval or combined for data retrieval. For classification levels of class, order, family, and genus, different split databases of the previous classification level can be selected, and two or more of them can be combined as a sub-database of the previous classification level; or two or more typical split databases of the same previous classification level can be selected and combined as a sub-database of the previous classification level. The iterative search (MetaPro-IQ) method or the spectral clustering (Spectra Cluster) method is selected to obtain a reduced exclusive database; Pfind, Maxquant, or Proteome Discoverer protein search software is used to search based on the reduced exclusive database. The marine sample is one or more than two of marine seawater, marine seabed sediment, and marine microbial community. The marine microbial community includes a mixed solution of multiple marine microorganisms or marine microorganisms on a polyester fiber membrane after filtering seawater. Two or more different marine samples are analyzed, their quantitative results are imported into protein quantitative analysis software, and the difference proteins between them are obtained, the functions of the difference proteins are analyzed, and the functional composition of the microbial community in different marine samples is explored. ​ ​ ​ ​ ​ 2. The resolution method according to claim 1, characterized in that: ​ 3. The resolution method according to claim 1, characterized in that: ​ 4. The method of claim 3, wherein: ​ 5. The resolution method according to claim 1, characterized in that: ​ ​ 6. The data parsing method of claim 1, wherein: ​ 7. The resolution method according to claim 6, characterized in that: ​ 8. The resolution method according to claim 1, characterized in that: ​

Citation Information

Patent Citations

  • Preparation of Protein Solid-Phase Alkylation Reagent, Solid-Phase Alkylation Reagent and Its Application

    CN106632877B

  • Method for analyzing microbial population function by using metagenome data

    CN108804875A

  • Method for analyzing transcription and translation activity of functional genes of microbial community

    CN112342284A