Plastic degrading enzyme database and construction method thereof
By integrating multiple databases and literature data to build a PDEdb database, and applying advanced algorithms to predict three-dimensional structures and simulate enzyme combinations, the data fragmentation problem of plastic degradation enzyme research is solved, and efficient enzyme engineering transformation and plastic pollution control support is achieved.
Patent Information
- Application Number
- CN202510546277.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The research data on plastic degraded enzymes in the prior art is fragmented and the lack of system integration and update mechanisms, making it difficult for researchers to efficiently tap potential resources, and the three-dimensional structure analysis is insufficient, which affects the research on enzyme engineering transformation and catalytic mechanisms.
Integrate PlasticDB, PMBD, PMDB and related literature data to construct a PDEdb database, which contains protein sequences, annotation information and three-dimensional structural models of a variety of plastic degrading enzymes, and incorporates new research results through dynamic update mechanisms. ESMfold is used to predict three-dimensional structures, and smina simulates the docking of enzymes with plastic polymer molecules.
It has achieved comprehensive coverage and timeliness support for the information of degraded enzymes of various plastic types, quickly and accurately identify potential degraded proteins, provides high-precision structural prediction and binding mode simulation, and promoted the development of enzyme engineering transformation and plastic pollution control technology.
Smart Images

Figure CN120452559A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of enzyme engineering and bioinformatics, and in particular to a plastic-degrading enzyme database and a construction method thereof. Background Art
[0002] In today's world, plastic pollution has become an increasingly serious global environmental problem. Due to their resistance to natural degradation, plastic products continue to accumulate in the environment, causing extremely profound negative impacts on ecosystems. To address this challenge, microbial-mediated biodegradation technologies have gradually gained attention. Among them, plastic degrading enzymes (PDEs), with their efficient catalytic effects, play a central role in the decomposition of difficult-to-degrade plastics such as polyester materials. In recent years, scientists have conducted extensive research dedicated to discovering and understanding the functional mechanisms of PDEs, but current research data is highly fragmented.
[0003] In existing research, known plastic degradation gene and protein sequences are scattered across various literature and databases, lacking effective systematic integration and standardized analysis tools. This makes it difficult for researchers to efficiently tap into potential plastic degradation resources, thus limiting the functional analysis and engineering applications of novel enzyme systems. For example, while PlasticDB integrates plastic degradation enzyme sequences scattered across different studies to form a relatively systematic database of plastic degradation enzymes, its data coverage and update mechanism still need to be improved. PMBD (Plastic Microbial Biodegradation Database) collects microorganisms and genes related to plastic degradation, focusing on the genes and annotation information of various microorganisms that degrade plastics. However, it mainly focuses on data support at the biodegradation functional level. PMDB (Plastic Microbial Degradation Database) focuses on data resources related to the relationship between microorganisms and plastic degradation, integrating information on microbial plastic degradation from various sources and providing a certain degree of gene function annotation and classification, but it also lacks in the depth and breadth of its data.
[0004] Furthermore, while big data such as metagenomes and metatranscriptomes contain a wealth of information on environmental microbial degradation, existing technologies lack specific analytical methods for plastic degradation, making it impossible to quickly and accurately identify target gene resources in complex environments. Furthermore, the three-dimensional structure of plastic-degrading enzymes is a core basis for guiding enzyme engineering and catalytic mechanism research. However, a lack of structural prediction research on large-scale plastic-degrading proteins has made it difficult to optimize functional sites.
[0005] To this end, those skilled in the art have proposed a plastic degrading enzyme database and a method for constructing the same to solve the above problems. Summary of the Invention
[0006] In view of the deficiencies of the prior art, the present invention provides a plastic degrading enzyme database and a method for constructing the same, which solves the problems raised in the above background technology.
[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions: A plastic degrading enzyme database, comprising the following contents:
[0008] Integrates PlasticDB, PMBD, PMDB and related literature data, covering information on enzymes that degrade various types of plastics;
[0009] Contains protein sequences, annotation information, and three-dimensional structure models of various plastic-degrading enzymes;
[0010] Dynamic updates allow for timely inclusion of new research findings and discoveries.
[0011] A method for constructing a plastic degrading enzyme database, according to the plastic degrading enzyme database of claim 1, comprising the following steps:
[0012] Integrate protein sequences in PlasticDB, PMBD, and PMDB databases, as well as known plastic degradation genes reported in the literature, to construct the initial plastic degradation protein database PDEdb;
[0013] Collect various environmental metagenome and metatranscriptome data from self-tests and NCBI, first filter them by Fastp quality control, then assemble them using metaSpades or megahit, and then predict genes using prodigal;
[0014] Using the PDEdb protein dataset as a reference, plastic degradation protein sequences in the microbiome data were identified by blastp alignment when the similarity was ≥80% and the coverage was ≥70%;
[0015] CDhit was used to remove redundancy from plastic degradation protein sequences with a similarity threshold of 99% to obtain a non-redundant plastic degradation protein database;
[0016] ESMfold was used to predict the three-dimensional structure of plastic-degrading enzyme protein, and SMINA was used to dock the enzyme with plastic polymer molecules and simulate the binding mechanism.
[0017] Preferably, the various environmental metagenome and metatranscriptome data collected from self-tests and NCBI include but are not limited to environmental sample data such as plastic surfaces, sewage systems, drinking water systems, river water, and seawater.
[0018] Preferably, the protein sequences in the constructed initial plastic degradation protein database PDEdb are strictly screened and verified to ensure their accuracy.
[0019] Preferably, the non-redundant plastic degradation protein database obtained by removing redundancy from the plastic degradation protein sequences using CDhit provides a more streamlined and representative data set for subsequent structure prediction and molecular docking.
[0020] Preferably, when applying ESMfold to predict the three-dimensional structure of plastic degrading enzyme protein, an advanced algorithm based on a language model is used to achieve accurate prediction at the atomic level.
[0021] Preferably, when smina is used to dock the enzyme with the plastic polymer molecule, the binding mode between the enzyme and the plastic molecule is simulated.
[0022] Preferably, the constructed plastic degrading enzyme database can be dynamically updated to promptly incorporate new research results and discoveries, thereby maintaining the timeliness and advancement of the database.
[0023] The present invention provides a plastic-degrading enzyme database and a method for constructing the same. It has the following beneficial effects:
[0024] 1. This invention integrates databases such as PlasticDB, PMBD, and PMDB, as well as relevant literature data, to construct the PDEdb database, which covers information on enzymes that degrade various plastic types. Compared to existing technologies, this database not only integrates existing plastic-degrading enzyme data but also promptly incorporates new research results and discoveries through a dynamic update mechanism, ensuring the timeliness and advancement of its content. This comprehensive and systematic data resource provides strong support for the research of plastic-degrading enzymes, enabling researchers to obtain rich and accurate plastic-degrading protein sequences, annotation information, and three-dimensional structural models within a unified framework. This database lays a solid foundation for in-depth research on the functional mechanisms of plastic-degrading enzymes, enzyme engineering, and the development of efficient plastic pollution control solutions.
[0025] 2. This method collects data from various environmental samples collected from self-tests and NCBI, and then performs preprocessing steps such as quality control filtering, assembly, and gene prediction. Using the constructed PDEdb database as a reference, blastp comparisons can rapidly and accurately identify potential plastic-degrading protein sequences from complex microbiome data. This process addresses the existing challenges of low microbiome data utilization and a lack of specific analysis methods.
[0026] 3. Based on the PDEdb database, this paper applies the language model-based ESMfold algorithm to predict the three-dimensional structure of plastic-degrading enzyme proteins, achieving precise atomic-level predictions. Furthermore, SMINA is used to dock the enzyme with plastic polymer molecules, simulating the binding pattern between the enzyme and the plastic molecule. Compared with the existing technology, this paper not only provides rich structural information for plastic-degrading enzyme proteins, but also provides a key theoretical basis for understanding the catalytic mechanism of enzymes and optimizing their functional sites. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 This is a flow chart of the plastic degrading enzyme database construction method of the present invention. DETAILED DESCRIPTION
[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0029] Example 1: Construction and application of database
[0030] Please see the attached Figure 1 , an embodiment of the present invention provides a plastic degrading enzyme database, including the following contents:
[0031] Integrates PlasticDB, PMBD, PMDB and related literature data, covering information on enzymes that degrade various types of plastics;
[0032] Contains protein sequences, annotation information, and three-dimensional structure models of various plastic-degrading enzymes;
[0033] Dynamic updates allow for timely inclusion of new research findings and discoveries.
[0034] Specifically, PlasticDB can integrate plastic-degrading enzyme sequences scattered across different studies, providing researchers with query and download functions to facilitate scientific researchers in obtaining existing plastic-degrading enzyme-related data.
[0035] PMBD mainly collects microorganisms and genes related to plastic degradation, focusing on the genes and annotation information of various microorganisms that degrade plastics, and focusing on providing data support at the functional level of plastic biodegradation, which helps to understand the role and contribution of different microorganisms in the plastic degradation process.
[0036] PMDB focuses on data resources on the relationship between microorganisms and plastic degradation. It integrates information on microbial degradation of plastics from different sources, including microbial species, functional genes and degradation mechanisms, and provides a certain degree of gene function annotation and classification, providing rich background data and reference information for studying how microorganisms participate in plastic degradation.
[0037] Related Literature Data: In addition to the data in the aforementioned databases, we have also extensively collected and compiled known plastic-degrading genes reported in various scientific literature. These publications, from various research teams and academic journals, cover numerous research findings on plastic-degrading enzymes, including new enzyme discoveries, enzyme property studies, and enzyme mechanism exploration, enriching the database with more novel and cutting-edge information.
[0038] By integrating data from these diverse sources, the Plastic Degrading Enzyme Database comprehensively covers enzymes that degrade a wide range of plastic types. This includes enzymes that degrade common plastics such as polyethylene (PE), polypropylene (PP), and polystyrene (PS), as well as enzymes that degrade emerging or specialized plastics. For example, with the continuous emergence of new biodegradable plastic materials and the in-depth study of their degradation methods, the database also includes information on enzymes that degrade these new biodegradable plastic materials, providing researchers with a comprehensive and extensive information platform for plastic degrading enzymes, which will help promote the development and application of different plastic degradation technologies.
[0039] Protein sequences are the order in which the basic building blocks of proteins are arranged. For plastic-degrading enzymes, their protein sequences determine the enzyme's structure and functional properties. The protein sequence information included in the database details the amino acid sequences of various plastic-degrading enzymes. This sequence data was obtained through experimental determination, gene sequence prediction, and literature review. Researchers can use this protein sequence information for comparative analysis, homology modeling, and functional prediction, providing fundamental data support for discovering new plastic-degrading enzymes, improving the performance of existing enzymes, and understanding the mechanisms of enzyme action.
[0040] Annotations provide detailed descriptions of the various properties and characteristics of plastic-degrading enzymes, including but not limited to the enzyme's name, classification, microbial source, substrate specificity, and gene coding sequence. This annotation helps researchers quickly understand the basic characteristics and biological functions of each plastic-degrading enzyme, providing an important reference for experimental design, enzyme application development, and related basic research. Understanding enzyme substrate specificity can help screen for highly effective enzymes that degrade specific plastics.
[0041] Three-dimensional structural models: Three-dimensional structural models show the folding and conformation of plastic-degrading enzymes in space, which is crucial for understanding the enzyme's catalytic mechanism, substrate binding mode, and interaction with other molecules. The three-dimensional structural models provided in the database are mainly predicted based on advanced computational methods (such as the ESMfold algorithm based on language models). These three-dimensional structural models present key features of the enzyme in a visual way, such as the active center, binding pocket, and secondary structure. They provide data support for researchers to conduct research on the relationship between enzyme structure and function, perform enzyme engineering (such as site-directed mutagenesis, protein design, etc.), and design new plastic-degrading enzyme inhibitors or activators.
[0042] The plastic-degrading enzyme database is dynamically updated and continues to evolve as the field of plastic degradation research continues to develop. The R&D team regularly monitors the latest scientific research literature, academic conference results, and updates to relevant databases, promptly incorporating new research results and discoveries into the database. For example, when a new plastic-degrading enzyme is discovered and verified, its related protein sequence, annotation information, and three-dimensional structural model data are quickly organized and integrated into the database. For existing enzyme data, if more detailed property studies, more accurate structural determinations, or new application discoveries are made, the corresponding information in the database will also be promptly updated and supplemented.
[0043] This dynamic update mechanism ensures the database's content is up-to-date and advanced, consistently reflecting the latest advances and state-of-the-art research in plastic-degrading enzymes. This prevents researchers from using outdated or incomplete data, leading to research deviations and duplication of effort. Furthermore, the dynamically updated database provides more reliable and cutting-edge technical support for the practical application and development of plastic-degrading enzymes, helping to promote continuous innovation and development in plastic pollution control technologies and accelerate the transition of plastic-degrading enzymes from basic research to practical applications.
[0044] A method for constructing a plastic degrading enzyme database comprises the following steps:
[0045] Integrate protein sequences in PlasticDB, PMBD, and PMDB databases, as well as known plastic degradation genes reported in the literature, to construct the initial plastic degradation protein database PDEdb;
[0046] Specifically, integrating the protein sequences of known plastic degradation genes in the three databases of PlasticDB, PMBD, and PMDB as well as related literature is like collecting puzzle pieces scattered in various places and piecing them together into a more complete and clearer puzzle. Each database has its own focus and data advantages. For example, PlasticDB is good at integrating known plastic degrading enzyme sequences, PMBD pays more attention to the genetic information of microbial degradation of plastics, and PMDB focuses on the relationship between microorganisms and plastic degradation and related functional genes. By integrating them with the scattered and valuable known plastic degradation genes in the literature, the constructed initial plastic degradation protein database PDEdb can cover more comprehensive information related to plastic degradation enzymes. This not only greatly facilitates scientific researchers to obtain data in one stop, but also lays a solid data foundation for subsequent in-depth research on the various characteristics of plastic degrading enzymes and the development of new degradation technologies.
[0047] We collected various environmental metagenome and metatranscriptome data from our own tests and NCBI, first filtered through Fastp quality control, assembled using metaSpades or megahit, and then predicted genes using prodigal. We also collected various environmental metagenome and metatranscriptome data from our own tests and NCBI, including but not limited to data from samples of plastic surfaces, sewage systems, drinking water systems, river water, and seawater. The protein sequences in the initial plastic degradation protein database (PDEdb) were rigorously screened and verified to ensure their accuracy.
[0048] Specifically, we collected metagenomic and metatranscriptomic data from various environmental samples, including plastic surfaces, sewage systems, drinking water systems, river water, and seawater, from self-tests and NCBI (National Center for Biotechnology Information). These samples cover nearly every practical scenario where plastics might be present. Because the raw data inevitably contains some low-quality, impure, or incomplete sequences, we first used Fastp to perform quality control filtering on the data to retain high-quality, reliable sequences.
[0049] Next, use metaSpades or megahit, two powerful assembly software, to assemble the short, scattered sequence fragments into long sequence fragments, just like splicing a piece of paper. Figure 1 This makes the sequence more meaningful and complete. Then, using the highly efficient gene prediction tool prodigal, the assembled sequence is precisely identified with the potential protein-encoding gene regions, specifically those that may contain genes encoding plastic-degrading enzymes.
[0050] Furthermore, when constructing the initial plastic degradation protein database PDEdb, the protein sequences incorporated were rigorously screened and verified to ensure that each sequence had a reliable source and accurate characterization. This is equivalent to a "rigorous identity review" of each sequence before entering the database, excluding those uncertain or potentially erroneous sequences. This ensures the high quality of the data in the database, and researchers will feel more assured and confident when using it.
[0051] Using the PDEdb protein dataset as a reference, plastic degradation protein sequences in the microbiome data were identified by blastp alignment when the similarity was ≥80% and the coverage was ≥70%;
[0052] CDhit was used to remove redundancy from plastic degradation protein sequences, with a similarity threshold of 99%, to obtain a non-redundant plastic degradation protein database; the non-redundant plastic degradation protein database after CDhit was used to remove redundancy from plastic degradation protein sequences provided a more streamlined and representative dataset for subsequent structure prediction and molecular docking.
[0053] Specifically, CDhit is a commonly used tool for sequence deduplication. Here, it was used to dedupe plastic degradation protein sequences, with a similarity threshold set at 99%. This indicates that if two sequences have a similarity of 99% or higher, they are considered likely to be near-duplicates or highly similar variants. Only one representative sequence is retained, while the others are removed.
[0054] The benefit of this approach is that it significantly reduces the amount of data while retaining key, representative sequence features. This allows for higher computational efficiency in subsequent structure prediction and molecular docking, and the analysis results are more focused on sequences with truly unique characteristics and potential value, avoiding the interference and unnecessary waste of computational resources caused by large numbers of repetitive sequences.
[0055] ESMfold was used to predict the 3D structure of a plastic-degrading enzyme protein, and smina was used to dock the enzyme with a plastic polymer molecule and simulate the binding mechanism. ESMfold employed an advanced algorithm based on a language model to achieve precise atomic-level predictions. Smina was used to dock the enzyme with the plastic polymer molecule, simulating the binding pattern between the enzyme and the plastic molecule.
[0056] Specifically, ESMfold is an advanced algorithm based on a language model for predicting the three-dimensional structure of proteins. By learning the relationships between a large number of known protein sequences and structures, it can accurately predict protein folding and spatial conformation. When predicting the three-dimensional structure of plastic-degrading enzyme proteins, ESMfold's advantages lie in its high accuracy and efficiency. These include the following benefits:
[0057] ESMfold uses a deep learning approach based on language models. By studying massive amounts of protein sequence data, it can capture hidden patterns and regularities within protein sequences. This allows for accurate prediction of protein three-dimensional structures at the atomic level. This precise atomic-level prediction means the precise position of every atom in a protein can be determined, revealing detailed structural details and providing high-resolution structural information for subsequent research.
[0058] Compared with traditional protein structure prediction methods, ESMfold is more computationally efficient. It can quickly process large amounts of protein sequence data and generate prediction results in a relatively short time.
[0059] Smina is a molecular docking tool used to simulate the interactions between molecules. When studying the interaction between plastic-degrading enzymes and plastic polymers, Smina can simulate the binding mode between enzymes and plastic molecules. This includes the following effects:
[0060] Molecular docking can predict the binding mode between the active site of the plastic-degrading enzyme and the plastic polymer molecule, including the formation of different types of intermolecular forces such as hydrogen bonds, hydrophobic interactions, and van der Waals forces.
[0061] In summary, the application of these two advanced computational tools, ESMfold and smina, allows for in-depth study of the mechanisms of action of plastic-degrading enzymes, from the perspectives of protein structure prediction and molecular interaction simulation, respectively. This not only helps to reveal the functions and characteristics of plastic-degrading enzymes, but also provides important theoretical support and guidance for the development of efficient plastic-degrading enzyme preparations and plastic pollution control technologies.
[0062] The constructed plastic-degrading enzyme database can be dynamically updated to promptly incorporate new research results and discoveries, thereby maintaining the timeliness and advancement of the database.
[0063] Example 2: Construction and application of database
[0064] The build process is as follows:
[0065] Data integration: We integrated protein sequences from the PlasticDB, PMBD, and PMDB databases, and collected known plastic degradation genes reported in relevant literature to construct the initial plastic degradation protein database, PDEdb. This step ensured that the database's initial data sources were broad and representative.
[0066] Data Collection and Preprocessing: We collected various environmental metagenome and metatranscriptome data from our own tests and NCBI, including data from samples of plastic surfaces, sewage systems, drinking water systems, river water, and seawater. These data were first filtered using Fastp quality control to remove low-quality sequences and impurities. MetaSpades or MegaHit was then used to assemble short sequences into long fragments. Prodigal was then used to predict genes and identify potential coding regions.
[0067] Sequence alignment and identification: Using the PDEdb protein dataset as a reference, blastp alignment was performed, setting similarity ≥ 80% and coverage ≥ 70% as thresholds to identify plastic-degrading protein sequences in the microbiome data and screen out sequences with high similarity to known plastic-degrading enzymes.
[0068] De-redundancy processing: CDhit was used to deduplicate the plastic degradation protein sequences, and the similarity threshold was set to 99%. Highly similar sequences were removed to obtain a non-redundant plastic degradation protein database.
[0069] Structure prediction and molecular docking: ESMfold was used to predict the three-dimensional structure of the plastic-degrading enzyme protein, and SMINA was used to dock the enzyme with the plastic polymer molecule and simulate the binding mechanism.
[0070] This database has played a significant role in the research and application of plastic-degrading enzymes. Using the PDEdb database constructed by this invention, researchers can obtain rich and accurate protein sequences, annotations, and three-dimensional structural models within a unified framework, laying a solid foundation for in-depth research on the functional mechanisms of plastic-degrading enzymes, enzyme engineering, and the development of efficient solutions for plastic pollution control. By collecting data from diverse environmental samples, the utilization rate of microbiome data is increased, enabling the rapid and accurate identification of target gene resources in complex environments.
[0071] Example 3: Dynamic update and performance improvement of database
[0072] The dynamic update mechanism is as follows:
[0073] Continuous data collection: New plastic-degrading enzyme-related data are regularly collected from the latest research literature, PlasticDB, PMBD, PMDB databases, self-tests and NCBI.
[0074] Data integration and screening: The collected new data will be integrated into the existing PDEdb database, and the newly added protein sequences will be strictly screened and verified to ensure their accuracy.
[0075] Database update: Update the information in the database, including adding new protein sequences, updating annotation information, supplementing three-dimensional structure models, etc., while removing duplicate or outdated data to ensure that the database always remains timely and advanced.
[0076] Through a dynamic update mechanism, the database can promptly incorporate new research findings and discoveries, maintaining its timeliness and advancement. This not only continuously increases the amount of data in the database but also improves its quality and reliability. This dynamically updated database can more accurately reflect the latest advances in plastic-degrading enzyme research when providing research support to researchers, helping to accelerate the discovery and utilization of plastic-degrading enzyme resources and promote the development of plastic pollution control technologies.
[0077] The three-dimensional structural model predicted by this analytical method can demonstrate the spatial conformation of the plastic-degrading enzyme, providing a more intuitive basis for understanding the enzyme's catalytic mechanism. Furthermore, molecular docking results can reveal the interaction between the enzyme and plastic molecules, facilitating the development of more efficient plastic-degrading enzyme preparations, improving plastic degradation efficiency, and providing stronger technical support for addressing the problem of plastic pollution.
[0078] The above examples demonstrate that the construction method of the present invention offers significant advantages and benefits in integrating data resources, discovering potential plastic-degrading enzyme genes, predicting protein structures, and conducting molecular docking. The database's dynamic update mechanism ensures its timeliness and advancement, providing strong support for the research and application of plastic-degrading enzymes and promoting the development of technologies for plastic pollution control.
[0079] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A plastic degrading enzyme database, characterized in that: Includes the following: Integrates PlasticDB, PMBD, PMDB and related literature data, covering information on enzymes that degrade various types of plastics; Contains protein sequences, annotation information, and three-dimensional structure models of various plastic-degrading enzymes; Dynamic updates allow for timely inclusion of new research findings and discoveries.
2. A method for constructing a plastic degrading enzyme database, according to the plastic degrading enzyme database of claim 1, characterized in that: The following steps are involved: Integrate protein sequences in PlasticDB, PMBD, and PMDB databases, as well as known plastic degradation genes reported in the literature, to construct the initial plastic degradation protein database PDEdb; Collect various environmental metagenome and metatranscriptome data from self-tests and NCBI, first filter them by Fastp quality control, then assemble them using metaSpades or megahit, and then predict genes using prodigal; Using the PDEdb protein dataset as a reference, plastic degradation protein sequences in the microbiome data were identified by blastp alignment when the similarity was ≥80% and the coverage was ≥70%; CDhit was used to remove redundancy from plastic degradation protein sequences with a similarity threshold of 99% to obtain a non-redundant plastic degradation protein database; ESMfold was used to predict the three-dimensional structure of plastic-degrading enzyme protein, and SMINA was used to dock the enzyme with plastic polymer molecules and simulate the binding mechanism.
3. The method for constructing a plastic degrading enzyme database according to claim 2, characterized in that: The various environmental metagenome and metatranscriptome data collected from self-tests and NCBI include but are not limited to environmental sample data such as plastic surfaces, sewage systems, drinking water systems, river water, and seawater.
4. The method for constructing a plastic degrading enzyme database according to claim 2, characterized in that: The protein sequences in the initial plastic degradation protein database PDEdb were strictly screened and verified to ensure their accuracy.
5. The method for constructing a plastic degrading enzyme database according to claim 2, characterized in that: The non-redundant plastic degradation protein database obtained by removing redundancy from the plastic degradation protein sequences using CDhit provides a more streamlined and representative data set for subsequent structure prediction and molecular docking.
6. The method for constructing a plastic degrading enzyme database according to claim 2, characterized in that: When applying ESMfold to predict the three-dimensional structure of plastic-degrading enzyme protein, an advanced algorithm based on a language model is used to achieve accurate prediction at the atomic level.
7. The method for constructing a plastic degrading enzyme database according to claim 2, characterized in that: When smina is used to dock the enzyme with the plastic polymer molecule, the binding mode between the enzyme and the plastic molecule is simulated.
8. The method for constructing a plastic degrading enzyme database according to claim 2, characterized in that: The constructed plastic degrading enzyme database can be dynamically updated, and new research results and discoveries can be incorporated in a timely manner to maintain the timeliness and advancement of the database.
Citation Information
Patent Citations
Disease Associated Protein Database
CN109086574A
Method for analyzing microbial quorum sensing effect based on metagenome data
CN112365929A
Biochar treatment method for relieving soil microplastic pollution
CN115213212A
Method and device for identifying novel double-stranded DNA cytidine deaminase, computer readable storage medium and application
CN116721700A
Gene design method and platform based on AI improvement
CN118212972A