Systems and methods for bioremediation of pollutants

By creating a knowledge base and designing customized microbial communities, and utilizing the combination of genes and enzymes in the microbial communities, the problem of incomplete degradation of pollutants in existing technologies was solved, and complete degradation of pollutants and environmental safety were achieved.

CN114341987BActive Publication Date: 2025-09-26TATA CONSULTANCY SERVICES LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080041904.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-12
Filing Date
2020-04-11
Publication Date
2025-09-26
Estimated Expiration
2040-04-11

AI Technical Summary

Technical Problem

Existing bioremediation methods fail to effectively utilize the complete degradation pathways of microbial communities, resulting in incomplete degradation of pollutants and the accumulation of intermediates and by-products in the environment, affecting human health.

Method used

By creating a knowledge base, identifying and designing customized microbial communities, and utilizing the combination of genes and enzymes in the microbial communities, complete degradation of pollutants can be achieved, and intermediate compounds are further metabolized by environmental microorganisms.

Benefits of technology

Complete degradation of pollutants is achieved, the accumulation of intermediates and by-products in the environment is avoided, and the efficiency and safety of bioremediation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114341987B_ABST
    Figure CN114341987B_ABST
Patent Text Reader

Abstract

Environmental contamination with various pollutants is becoming a global health problem. Numerous methods are being used to bioremediate these pollutants. Methods and systems for one or more pollutants have been provided. A sample is collected from a site containing the pollutants. The pollutants are then isolated from the sample. Furthermore, a knowledge base of various types of degradants for these pollutants is created. This knowledge base is used to create a microbial profile. The microbial profile is then used to design a first microbial community and a second microbial community that together provide the genes, proteins, and enzymes required to degrade the pollutants. Finally, a mixture of the first microbial community and / or the second microbial community is applied to the site. The method further includes testing the efficacy of the applied mixture and reapplying the mixture.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is a Chinese national phase application based on and claims priority from International Patent Application No. PCT / IN2020 / 050346, filed on April 11, 2020, which in turn claims priority from Indian Provisional Patent Application No. 201921014894, filed on April 12, 2019. The entire contents of the aforementioned applications are incorporated herein by reference. Technical Field

[0003] Embodiments of the present application relate generally to the field of waste management, and more particularly, to methods and systems for bioremediation of pollutants by engineering microbial communities capable of completely degrading the pollutants. Background Art

[0004] Environmental contamination with a variety of pollutants, most of which are generated by industrial and agricultural practices, is becoming a global health problem. Global industrialization has increased the production of several products containing hazardous chemical ingredients, including plastics, pesticides, synthetic fertilizers, electronic waste, industrial waste, food additives, cleaning products, cosmetics, dyes, etc. Many of these products are released into the environment and eventually enter the human food chain, mainly through the gastrointestinal tract, but may also use other routes such as the airways or skin. Therefore, it is necessary to design methods to help remove pollutants from the environment and eliminate pollutants that enter the human system. A group of these pollutants are also called endocrine disrupting chemicals (EDCs) because they affect an individual's hormonal makeup and are associated with a variety of metabolic disorders. In addition, exposure to chemicals may have an impact on the microorganisms that live in our bodies, known as the "human microbiome."

[0005] Various physical and chemical methods are used to degrade pollutants. Existing physical and chemical methods (e.g., thermal oxidation, photooxidation) are slow and expensive to implement compared to the extent of debris accumulation. Most current bioremediation methods identify initial enzymes within microorganisms capable of degrading pollutants and classify these microorganisms as potential degraders. These methods do not consider the entire degradation pathway. However, in most cases, such enzymes are promiscuous, as they bind to a range of substrates and tend to appear in multiple copies within the microbial genome. Therefore, identifying an enzyme alone cannot guarantee that a pollutant can be completely degraded. This approach leads to an increase in false-positive results and misleading conclusions. Furthermore, the presence of key intermediates and their importance in the pathway, which are often overlooked by existing methods, must be determined. Many of these intermediates and byproducts (such as phthalic acid and bisphenol A (BPA)) remain in the environment, not only polluting the environment (soil, water, and air) but also affecting humans by entering the food chain and altering hormone levels in both males and females.

[0006] In addition, in order to overcome the shortcomings of physical and chemical methods, some bioremediation methods are also used for waste management. Bioremediation is the process of using naturally occurring or specially introduced microorganisms (such as bacteria, fungi, etc.) to degrade environmental pollutants in contaminated sites. Microorganisms have the excellent ability to degrade large amounts of organic compounds by consuming them as their main energy source and further absorbing them without releasing any harmful byproducts. Biodegradation methods have become a preferred choice for pollution management because they are safer, cheaper and more sustainable remediation methods than chemical and physical methods.

[0007] Most bioremediation approaches have been implemented to obtain the desired degradation products. However, most of these approaches have failed due to incomplete or missing information about the pathways that microorganisms use to degrade pollutants. These degradation pathways can be partially present in different microorganisms within a single community. A single microorganism may not be able to completely degrade a pollutant on its own, but when combined with other microorganisms that possess the remaining components of the degradation pathway, it may be able to do so. In other words, one group of microorganisms may be able to absorb and degrade an intermediate produced by another group of microorganisms. Summary of the Invention

[0008] Embodiments of the present disclosure provide technical improvements as solutions to one or more of the aforementioned technical problems identified by the inventors in conventional systems. For example, in one embodiment, a system for bioremediation of one or more contaminants is provided. The system includes a sample collection module, a contaminant separation and identification module, a processor, and a memory in communication with the processor. The sample collection module collects a sample from an environmental site containing one or more contaminants. The contaminant separation and identification module separates and identifies the one or more contaminants present in the sample. The memory is configured to perform the following steps: creating a knowledge base, wherein the knowledge base stores: information on one or more identified pollutants, information on complete degradation pathways and partial degradation pathways identified in microorganisms that can completely degrade one or more pollutants or partially degrade one or more pollutants, information on the corresponding environmental niches in which the microorganisms thrive, and a list of microorganisms from different environments that have specific complete / partial pollutant degradation pathways; by utilizing the information in the knowledge base, identifying a list of partial pollutant degraders and a list of complete pollutant degraders for each of the one or more pollutants identified in the sample, wherein the partial pollutant degrader refers to a microorganism that provides one or more subpathways and a corresponding set of genes, encoded proteins or enzymes for converting pollutants into intermediate compounds, and wherein multiple partial degraders provide in combination all subpathways for completely degrading pollutants identified in the collected sample, wherein the complete pollutant degrader has all subpathways and a corresponding set of genes, encoded proteins or enzymes within a single microorganism for degrading the pollutants identified in the collected sample. a combination of proteins or enzymes that degrade the pollutants; creating a microbial profile using information from a knowledge base, wherein the microbial profile includes information on one or more partial pollutant degraders and complete pollutant degraders that can degrade each of the one or more pollutants identified in the sample to different degrees of degradation, wherein the different degrees of degradation of the pollutants refer to degradation of the pollutants into different intermediate compounds or metabolites, and wherein the intermediate compounds or metabolites are determined by one or more end products, which are released by the degraders under the action of genes or proteins or enzymes corresponding to subpathways present in the genome of the degraders that degrade the pollutants, and wherein the intermediate compounds can be released into the environment and utilized by other microorganisms in the environment or can be taken up by the same degrading microorganisms; designing a first microbial community using the created microbial profile, the first microbial community including microorganisms that collectively provide the subpathways required for complete degradation of the one or more pollutants identified in the sample, and wherein the microorganisms can survive together in the same environmental niche from which the sample was collected;Using the created microbial profile, design a second microbial community, the second microbial community comprising microorganisms that collectively provide genes, proteins, and enzymes for subpathways required for partially degrading one or more pollutants identified in the collected sample into one or more desired intermediates, wherein the microorganisms forming the second microbial community are capable of surviving together in the environmental niche of the collected sample; applying a mixture of at least one or both of the first microbial community and the second microbial community to an environmental site containing one or more pollutants; examining the efficacy of the applied mixture in eliminating the one or more pollutants in the sample collected from the environmental site, wherein the efficacy is evaluated by isolating and identifying combinations of residual pollutants from the collected sample; and reapplying a new mixture to the environmental site, wherein the new mixture is prepared by adding a group of microorganisms capable of acting as partial degraders and combinatorially degrading the one or more pollutants in the collected sample.

[0009] In another aspect, a method for bioremediation of one or more pollutants is also provided. First, a sample is collected from an environmental site containing one or more pollutants. The one or more pollutants present in the sample are then separated and extracted. In the next step, a knowledge base is created. The knowledge base stores information about the identified pollutant(s), complete and partial degradation pathways identified in microorganisms that are capable of completely or partially degrading the one or more pollutants, information about the corresponding environmental niches in which the microorganisms thrive, and a list of microorganisms from different environments that possess specific complete / partial pollutant degradation pathways. A complete degradation pathway refers to a set of genes and / or proteins encoded by a microorganism in its genome that are responsible for completely degrading the pollutant into compounds that are safe for the environmental site or that can be absorbed by one or more other microorganisms in the environment. A partial degradation pathway in a microorganism refers to a set of genes or encoded proteins that constitute one or more subpathways, wherein a subpathway is a subset of the complete degradation pathway encoded within the microorganism's genome and degrades the pollutant into intermediate compounds that can be released by the microorganism into the environment and subsequently absorbed by another microorganism in the environment, wherein the other microorganism has another subpathway that metabolizes the released intermediate compounds. In the next step, a list of partial pollutant degraders and a list of complete pollutant degraders are identified for each of the one or more pollutants identified in the sample using information from the knowledge base. Partial pollutant degraders are microorganisms that provide one or more subpathways and a corresponding set of genes, encoded proteins, or enzymes that convert pollutants into intermediate compounds, and wherein multiple partial degraders can be combined to provide all subpathways for completely degrading the pollutants identified in the collected sample, wherein a complete pollutant degrader has a combination of all subpathways and a corresponding set of genes, encoded proteins, or enzymes within a single microorganism for degrading the pollutants identified in the collected sample. In addition, a microbial profile is created using information from the knowledge base. The microbial profile includes information on one or more partial pollutant degraders and complete pollutant degraders that are capable of degrading each of the one or more pollutants identified in the sample to different degrees of degradation. The different degrees of degradation of the pollutant refer to the degradation of the pollutant into different intermediate compounds or metabolites, and wherein the intermediate compounds or metabolites are determined by one or more final products, which are released by the degrader under the action of genes, proteins, or enzymes corresponding to the subpathways present in the genome of the degrader that degrades the pollutant. The intermediate compounds can be released into the environment and utilized by other microorganisms in the environment or can be taken up by the same microorganisms that degraded them.In the next step, the created microbial profile is used to design a first microbial community. This first microbial community includes microorganisms that collectively provide the subpathways required for complete degradation of one or more pollutants identified in the sample, and wherein the microorganisms are able to co-exist in the same environmental niche where the sample was collected. Similarly, the created microbial profile is used to design a second microbial community. This second microbial community includes microorganisms that collectively provide the genes, proteins, and enzymes required for the subpathways for partial degradation of one or more pollutants identified in the collected sample into one or more desired intermediate products, wherein the microorganisms forming the second microbial community are able to co-exist in the environmental niche where the sample was collected. In the next step, a mixture of at least one or both of the first and second microbial communities is applied to an environmental site containing one or more pollutants. Subsequently, the applied mixture is tested for its efficacy in eliminating the one or more pollutants in the sample collected from the environmental site. The efficacy is evaluated by isolating and identifying combinations of residual pollutants from the collected sample. Finally, a new mixture is re-applied to the environmental site, wherein the new mixture is prepared by adding a group of microorganisms that can act as partial degraders and degrade the one or more pollutants in the collected sample in combination.

[0010] On the other hand, a method for the bioremediation of one or more carbon-based pollutants is provided. Carbon-based pollutants refer to any pollutant molecules comprising one or more carbon atoms and possibly comprising one or more other atoms. In one embodiment, carbon-based pollutants may include polycyclic aromatic hydrocarbons (PAHs), polychlorinated biphenyls (PCBs), polyethylene terephthalate (PET) and carbon-based nanomaterials (CBNMs). Any other carbon-based pollutants are also included within the scope of the present disclosure. In one embodiment, the degradation of PET can be carried out by forming an intermediate compound (such as terephthalic acid (TPA)), which is further degraded into an intermediate (such as protocatechuic acid (PCA)) that is easily assimilated by microbial metabolism. In another embodiment, the bacterial degradation of CBNM is described. The initial step of the CBNM degradation in bacteria can be carried out by secretory bacterial peroxidases, and the intermediate produced in the process is found to be a cyclic aromatic hydrocarbon. Subsequent degradation is carried out by bacteria or microbial communities that can degrade aromatic hydrocarbons such as PAHs and biphenyls. In another embodiment, the degradation of PAHs (e.g., naphthalene, anthracene, and phenanthrene) may involve a set of coordinated enzymes and subpathways leading to the formation of intermediates that can be metabolized and digested by microorganisms, as described in detail in this disclosure. Similarly, in another embodiment, the degradation of PCBs into dehalogenated biphenyls and the degradation of dehalogenated biphenyls into intermediates that can be readily metabolized and digested by microorganisms have been described in this disclosure.

[0011] On the other hand, one or more non-transitory machine-readable information storage media are provided, which contain one or more instructions that, when executed by one or more hardware processors, cause the bioremediation of one or more pollutants. First, a sample is collected from an environmental site containing one or more pollutants. The one or more pollutants present in the sample are then separated and extracted. In the next step, a knowledge base is created. The knowledge base stores: information on one or more identified pollutants, information on complete degradation pathways and partial degradation pathways that can completely degrade one or more pollutants or partially degrade one or more pollutants identified in microorganisms, information on the corresponding environmental niches where microorganisms thrive, and a list of microorganisms with specific complete / partial pollutant degradation pathways from different environments. A complete degradation pathway refers to a set of genes on a microbial genome and / or proteins encoded by a microorganism, wherein the set of genes and / or encoded proteins is responsible for completely degrading pollutants into compounds that are safe for the environmental site or compounds that can be assimilated by one or more other microorganisms in the environment. A partial degradation pathway in a microorganism refers to a set of genes or encoded proteins that constitute one or more subpathways, where a subpathway is a subset of a complete degradation pathway encoded within the microorganism's genome, and the subpathway degrades the pollutant into an intermediate compound that can be released by the microorganism into the environment and subsequently taken up by another microorganism in the environment, where the other microorganism possesses another subpathway that metabolizes the released intermediate compound. In the next step, a list of partial pollutant degraders and a list of complete pollutant degraders are identified for each of the one or more pollutants identified in the sample, using information from a knowledge base. Partial pollutant degraders are microorganisms that provide one or more subpathways and a corresponding set of genes, encoded proteins, or enzymes that convert the pollutant into an intermediate compound. Multiple partial degraders can, in combination, provide all subpathways for the complete degradation of a pollutant identified in a collected sample. A complete pollutant degrader comprises a combination of all subpathways and a corresponding set of genes, encoded proteins, or enzymes within a single microorganism that degrades the pollutant identified in the collected sample. Furthermore, a microbial profile is created using information from the knowledge base. The microbial profile includes information on one or more partial and complete pollutant degraders capable of degrading each of the one or more pollutants identified in the sample to varying degrees of degradation. The varying degrees of degradation of the pollutant refer to the degradation of the pollutant into different intermediate compounds or metabolites, wherein the intermediate compounds or metabolites are determined by one or more end products released by the degrader under the action of genes, proteins, or enzymes corresponding to subpathways present within the genome of the degrader that degrade the pollutant. The intermediate compounds can be released into the environment and utilized by other microorganisms in the environment or can be absorbed into the same microorganisms that degrade the pollutant.In the next step, the created microbial profile is used to design a first microbial community, which includes microorganisms that collectively provide the subpathways required for completely degrading one or more pollutants identified in the sample, and wherein the microorganisms are able to co-exist in the same environmental niche where the sample was collected. Similarly, the created microbial profile is used to design a second microbial community, which includes genes, proteins, and enzymes that collectively provide the subpathways required for partially degrading one or more pollutants identified in the collected sample into one or more desired intermediate products, wherein the microorganisms forming the second microbial community are able to co-exist in the environmental niche where the sample was collected. In the next step, a mixture of at least one or both of the first and second microbial communities is applied to an environmental site containing one or more pollutants. Subsequently, the applied mixture is tested for its efficacy in eliminating the one or more pollutants in the sample collected from the environmental site. The efficacy is evaluated by isolating and identifying combinations of residual pollutants from the collected sample. Finally, a new mixture is re-applied to the environmental site, wherein the new mixture is prepared by adding a group of microorganisms that can act as partial degraders and degrade the one or more pollutants in the collected sample in combination.

[0012] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention, as claimed. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate exemplary embodiments and, together with the description, serve to explain the principles of the disclosure:

[0014] Figure 1 A block diagram of a system for bioremediation of contaminants according to an embodiment of the present disclosure is shown.

[0015] Figure 2A-2C is a flow chart illustrating the steps involved in the bioremediation of contaminants according to an embodiment of the present disclosure.

[0016] Figures 3A to 3C is a flow chart illustrating the steps involved in the creation of a knowledge base according to an embodiment of the present disclosure.

[0017] Figure 4 Various classes of contaminants that can be bioremediated using methods according to embodiments of the present disclosure are shown.

[0018] Figure 5 The proximal and distal active sites present in a bifunctional catalase-peroxidase in bacteria according to embodiments of the present disclosure are shown.

[0019] Figures 6A to 6C Schematic diagrams showing the degradation pathways of carbon-based pollutants PAH, PCB, and PET, respectively, according to embodiments of the present disclosure.

[0020] 7A to 7B is a flow chart illustrating the steps involved in the bioremediation of carbon-based pollutants according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] Exemplary embodiments are described with reference to the accompanying drawings. In the drawings, the left-most digit(s) of a reference number identifies the reference number of the drawing in which the reference number first appears. Wherever convenient, the same reference numerals are used throughout the drawings to identify identical or similar components. Although examples and features of the disclosed principles are described herein, modifications, variations, and other implementations are possible without departing from the scope of the disclosed embodiments. It is intended that the following detailed description be considered exemplary only, with the true scope being indicated by the following claims.

[0022] Glossary - Terms used in the examples

[0023] The expression "microorganism" or "organism" refers to a living organism, which may include bacteria, fungi, algae, protists, viruses, and the like.

[0024] In the context of the present disclosure, the expression "complete degrader" refers to a "complete pollutant degrader", whereas in the context of the present disclosure, the expression "partial degrader" refers to a "partial pollutant degrader".

[0025] In the context of the present disclosure, the expression "microorganism genome" refers to the microorganism genome as well as the corresponding protein and nucleotide sequences of the genome.

[0026] In the context of the present disclosure, the expression "degradation pathway" refers to the genetic mechanisms present in the genome of a microorganism for degrading / eliminating these various pollutants, wherein degradation refers to the conversion of pollutants into compounds that are assimilated into the metabolic machinery of the microorganism itself or that do not cause harm to the environment when released.

[0027] In the context of the present disclosure, the expression "key intermediate" refers to an intermediate formed during the degradation of a pollutant, which divides the pathway of pollutant degradation into its constituent sub-pathways.

[0028] Referring now to the accompanying drawings, and in particular to Figure 1 To the picture Figure 7B (wherein like reference numerals represent corresponding features consistently throughout the drawings), preferred embodiments are shown and described in the context of the following exemplary systems and / or methods.

[0029] According to an embodiment of the present disclosure, Figure 1The block diagram of the system 100 for the bioremediation of pollutants is shown. System 100 can identify different pollutants from any environmental contamination site (such as but not limited to soil, sediment, water, landfill, oil spill, etc.) and apply microbial communities to completely bioremediate pollutants. System 100 helps microbial communities to completely degrade pollutants and their intermediates without causing any harm to the environment. The complete degradation of pollutants refers to the conversion of pollutants into compounds that are harmless to the environment or that can be assimilated into compounds present in the microorganisms in a given environmental site. Complete degradation can be caused by a single microorganism or a microbial community / population that can cause the complete degradation of pollutants in combination. System 100 is using a method to determine the pollutant degradation potential in microorganisms by identifying genetic mechanisms in the form of degradation pathways / sub-pathways, which include genes and / or proteins and enzymes on the microbial genome that can perform one or more degradation reactions of a variety of pollutants, determine the final products obtained from these pathways / sub-pathways, and therefore determine the extent to which organisms degrade or change pollutants. Subpathways are defined so that these subsets of complete degradation pathways for pollutants and the genes / proteins / enzymes that comprise them exist independently in certain microorganisms and can metabolize the pollutants to release intermediates. These intermediates can be released into the environment by the microorganisms and absorbed by other microorganisms, which are then able to metabolize the intermediate compounds into other intermediates and release them. Combinations of different microorganisms can be designed so that this process can continue until the pollutant is completely degraded.

[0030] This information is used to create a back-end knowledge base. This knowledge base is then used to build customized microbial communities that, as a whole, provide a comprehensive view of the pollutant degradation potential, where each component microorganism has an enhanced ability to fully degrade the pollutant into products that can be assimilated into the environment or released into the environment without any harmful effects.

[0031] In the embodiment of system 100, can also design community in this way: the consortium (consortium) of microorganism is together pollutant degradation into intermediate compound or metabolite, and this intermediate / product obtained thus can be used for the application of multiple industry and other aspects.Have the microorganism that can pollutant degradation into the sub-pathway of different intermediate compound or metabolite is partial degradation agent.The combination of the partial microbial degradation agent of all sub-pathways in the degradation pathway can be provided jointly and can form the consortium / cocktail (cocktail) that can degrade pollutant completely in combination.System 10 can identify and set up the pollutant degradation potential of the microorganism that is present in environmental site, and it can improve the overall ability of community and design and can survive in this environment and can degrade the community of target pollutant.Various pollutants may comprise plastics, organic pollutants, inorganic pollutants, aromatic pollutants, gaseous and particulate pollutants, heavy metals, carbon-based nanomaterials and radioactive pollutants etc.

[0032] According to the embodiments of the present disclosure, Figure 1 As shown, the system 100 is composed of a sample collection module 102, a pollutant separation and identification module 104, a memory 106, and a processor 108. The processor 108 communicates with the memory 106. The processor 108 is configured to execute multiple algorithms stored in the memory 106. The memory 106 also includes multiple modules for performing various functions. The memory 106 includes a knowledge base module 110 and a community association module 112. Figure 1 As shown in the block diagram of , the system 100 also includes an administration module 114 and an efficacy module 116 .

[0033] According to an embodiment of the present disclosure, a sample collection module 102 is used to collect samples from a site. Samples can be collected from various pollutant sites, such as soil industrial wastelands, soil from textile wastewater discharge sites, and soil dumped with sewage sludge, sediments, water bodies, etc. Pollutants can also be collected from any other location affected by pollution. Samples are collected using site-specific methods. Sample collection methods vary depending on the type of environment / environmental niche (e.g., soil, sediment, and water) and the type of material being collected.

[0034] Soil samples between the two areas differed in absorptive properties, texture, density, moisture, the geological setting of the site, and the types and populations of microorganisms. Sampling depth varied depending on the type of sampling location (e.g., arable disturbed land vs. non-arable undisturbed land). Sample extraction was performed using tools such as augers, vehicle-mounted hydraulic augers, core barrels, trowels, brass sample sleeves, and solid tube samplers. Contaminants and contaminant fragments from the water column were collected using tools such as hydrographic bottles, manta nets, Neuston nets, and drift nets behind fixed or mobile vessels. The use of any other method for sample extraction is within the scope of the present disclosure.

[0035] Parameters such as the type of sampling location and its associated pH, salinity, temperature, and contaminant concentrations are recorded during sampling. These parameters are important because contaminant concentrations and microbial distribution can vary spatially or temporally. Similar sampling techniques can be applied to a variety of other environments with different types of contaminants.

[0036] According to an embodiment of the present disclosure, the system 100 also includes a pollutant separation and identification module 104. The pollutant separation and identification module 104 is configured to separate and identify one or more pollutants from the collected samples. There are different types of pollutants in the collected samples. The separation and identification of different types of pollutants are completed by physical methods or chemical methods. Pollutants are generally chemically inert and persistent, so physical methods are more often used than chemical methods. Classifying various pollutant entities in the sample usually constitutes the first step in pollutant identification and can be divided into categories such as optical or visual sensors, flotation technology, density measurement classification methods, etc. Based on the basic principles, different physical methods can also be used to identify different pollutants in the collected samples, for example, near-infrared sensors (NIR), electrostatic methods (such as friction electrostatic separators), hyperspectral imaging technology, pressurized fluid extraction (PFE), differential scanning calorimetry (DSC) and laser induced breakdown spectroscopy (LIBS) technology. The process of separating pollutants from sediment samples is completed by density separation. The use of any other method for pollutant separation is within the scope of the present disclosure.

[0037] Pollutant samples collected from different locations are filtered and subsequently subjected to wet peroxide oxidation (WPO), density separation, and gravimetric analysis to identify different types of micropollutants. The identification of pollutant types can be done using Fourier transform infrared (FTIR) spectroscopy, Fourier transform Raman spectroscopy (FT Raman), infrared or Raman spectroscopy, etc. Another method, "attenuated total reflectance" (ATR), is also good at identifying various pollutant components. Toxic substances, POPs, and chemicals used as additives to pollutant materials can be measured using methods such as chromatography, spectroscopy, etc. Heteroatoms in compounds (e.g., nitrogen, chlorine, sulfur, etc.) can be identified using laboratory heteroatom identification techniques, such as the Lassaigne method and the Beilstein test. Useful conclusions can be drawn from the results of these tests, and further, the results help to distinguish the types of pollutants in the query samples. Methods based on GC-MS (gas chromatography-mass spectrometry) and GC-ECD (gas chromatography-electron capture detection) can also be used to detect a variety of pollutants. Another aspect of detecting pollutants that affect human health involves immunoassays, including ELISA or cell-based assays. The separation and extraction of certain pollutants employs techniques such as repeated fractionation, ultrasonic / ultrasonic agitation, mechanical agitation, pressurized fluid extraction (PFE), and rotating disk sorptive extraction (RDSE). Certain pollutants, such as carbon-based nanomaterials (CBNMs), can be identified using electron microscopy. Other methods include optical detection methods, analyzing element ratios or isotopic signatures to determine the presence and type of CBNMs. The use of any other method for pollutant identification is within the scope of this disclosure.

[0038] The one or more pollutants (Pi) identified from the collected samples can be degraded through various pathways such as photo-oxidation, thermo-oxidative degradation, biodegradation, etc. In the present disclosure, the prospect of biodegradation is considered for complete and effective bioremediation of environmental samples from all pollutants.

[0039] According to an embodiment of the present disclosure, memory 106 includes a knowledge base module 110. The knowledge base module 110 creates a plurality of maps and matrices, as described below. All of these are referred to as a knowledge base together. The knowledge base module 110 is configured to create a knowledge base, wherein the knowledge base stores information on one or more pollutants identified, information on the complete degradation pathway and partial degradation pathway identified in the microorganism that can completely degrade the one or more pollutants or partially degrade the one or more pollutants, information on the corresponding environmental niche in which the microorganism is bred, and a column of microorganisms with specific complete / partial pollutant degradation pathways from different environments. A complete degradation pathway refers to a set of genes on a microorganism genome and / or proteins encoded by a microorganism, and the group of genes and / or encoded proteins or enzymes are responsible for completely degrading pollutants into environmentally friendly compounds or compounds that can be assimilated by one or more other microorganisms present in the environment. A partial degradation pathway in a microorganism refers to a set of genes or encoded proteins or enzymes that constitute one or more sub-pathways, and wherein a sub-pathway is a subset of the complete degradation pathway encoded in the genome of the microorganism. The sub-pathway degrades the pollutant into an intermediate compound that can be released by the microorganism into the environment and subsequently taken up by another microorganism in the environment, wherein the other microorganism has another sub-pathway that metabolizes the released intermediate compound.

[0040] The knowledge base serves as the backend data for building customized microbial communities for pollutant degradation. It also provides information on the microorganisms that degrade specific pollutants and the end products they form. It also provides information on specific microorganisms and the pollutants they can degrade—that is, the microorganism's overall pollutant degradation potential. Furthermore, the knowledge base provides information on the environmental conditions in which organisms evolve to survive or establish themselves.

[0041] According to an embodiment of the present disclosure, the knowledge base module 110 is configured to create a multidimensional pollutant pathway organism matrix (PPOM). The following steps are used for the pollutant pathway organism matrix. A pathway can be defined as a series of enzyme-catalyzed reactions, wherein the product formed in the previous reaction becomes the substrate for the subsequent reaction.

[0042] In the first step, a variety of isolated pollutants (P i In one example, for the separated pollutants (P i ) Generate the query string (Q in ): ['NameString'] + [(aerobic) or (anaerobic)] + [microorganism], where NameString = (P i)+[degradation or metabolism].

[0043] Query string (Q in ) is used as input to search selected literature search engines (such as Pubmed) and pathway databases (such as KEGG pathways / EAWAG (BBD / PPS) / MetaCyc). The result sets obtained from the literature search engines and pathway databases include as output a summary list A out and Pathway Results Set (PRS out ) and a list of organisms in which the pathway has been experimentally characterized ( out ). Using any other database to obtain pathway information and any other format of literature mining results (e.g., full text of published documents, literature reviews, etc.) are all within the scope of the present invention.

[0044] Further, for the degradation of each target pollutant P i The path P i Manual search of DP (which is a series of steps that converts a substrate contaminant into a final product or intermediate compound / metabolite) and the genes / enzymes involved in the process was performed using List A out And for each input query string Q in the previous step in This is done by manually managing and identifying the pathway result set obtained. All subpathways are then manually managed and identified. Subpathways are identified so that the product of one subpathway can serve as the initial substrate for the next subpathway. Therefore, all subpathways put together can completely degrade pollutants / compounds. The criteria involved in identifying subpathways may include: the presence of the subpathway (and its constituent genes / proteins) itself in the microbial genome; and the formation of a product that can be released into the environment and can be used by microorganisms in the community that have the next subpathway that can utilize this released product. This product is referred to herein as a "key intermediate metabolite" (CIM). All possible subpathways [P i SP1, P i SP2, P i SP3…P i SPn] together constitute the pollutant degradation pathway P i DP. The same can be explained with the help of the following example: Consider a hypothetical pathway P i DP, which includes the following steps. The enzymes catalyzing each reaction in the pathway are E1, E2…E9

[0045] E1 E2 E3 E4 E5 E6 E7 E8 E9

[0046] Subpathway P in this pathway i SP1, P i SP2 and P i SP3,

[0047] Among them, P i SP1=S1-->S4 is catalyzed by (E1, E2, E3) present in organism O1

[0048] P i SP2=S4-->S6 is catalyzed by (E4, E5) present in organic O2

[0049] P i SP3=S6-->S10 is catalyzed by (E6, E7, E8, E9) present in the organic O3.

[0050] And P i DP=P i SP1+P i SP2+P i SP3, O1, O2 and O3 accumulation provides pollutant P i The complete degradation pathway of P i DP.

[0051] These subpathways can be defined as complete degradation pathways P i DP and the subset of genes / proteins / enzymes that make it up, which are encoded within the genome of a microorganism and promote the biosynthesis of an intermediate (as described in S4, S6), which can be released by the microorganism into the environment and taken up by another microorganism in the environment that has a product pathway to utilize the released product.

[0052] Therefore, the pollutant P iA complete degradation pathway is a set of genes and / or proteins encoded by them on a microbial genome that are responsible for the complete degradation of a pollutant into a compound that is harmless to the environment or a compound that can be assimilated by a microorganism present in the environment. A partial degradation pathway in a microorganism is a set of genes and / or encoded proteins or enzymes that form one or more sub-pathways (a subset of the complete degradation pathway) encoded in the genome of a microorganism that is capable of degrading a pollutant into an intermediate compound that is taken up by another microorganism that has a sub-pathway for degrading the intermediate compound. This process can continue as a chain until a product is formed that can be assimilated by a microorganism present in the environment without releasing harmful substances into the environment. In summary, a sub-pathway is a subset of a complete degradation pathway encoded in a microbial genome, wherein a sub-pathway degrades a pollutant into an intermediate compound that can be released by the microorganism into the environment and subsequently taken up by another microorganism in the environment (206), wherein the other microorganism has another sub-pathway for metabolizing the released intermediate compound. The existence of all sub-pathways and corresponding genes / proteins / enzymes for degrading pollutants in a single microorganism makes the microorganism referred to as complete degradation agent or complete pollutant degradation agent (CPD), can degrade pollutants completely into compounds or metabolites that are harmless to the environment or assimilate the final product in its own metabolic process. On the contrary, partial degradation agent or partial pollutant degradation agent (PPD) refer to the microorganism that provides one or more sub-pathways and corresponding genes and / or encoded proteins or enzymes that pollutants are converted into intermediate compounds, and wherein, a plurality of partial degradation agents can provide all sub-pathways in combination, with the described pollutant identified in the sample collected to degrade completely. Degrade completely and refer to pollutant degradation into compounds or metabolites that are harmless to the environment or can be assimilated by the microorganism present in the environment of collecting sample or metabolites. Intermediate compounds or metabolites biosynthesized by these partial degradation agents can also be obtained separately using any industrial scale method or laboratory experimental procedure, and are used for multiple industrial and commercial applications.

[0053] Further, for each input query string Q in Obtained List A out , perform on the organism (O out ) and the range of the existence of the pathway. i SP1, P i SP2, P i SP3…P iSPn] conducted literature validation and identified "key intermediate metabolites" (CIMs). CIMs are metabolites formed during a pathway that can be released from one microorganism within a microbial community and taken up and utilized by another complementary microorganism within the community. This complementary microorganism can then utilize or metabolize the product as a nutrient source, a substrate for secondary metabolites, etc. This process can be continued similarly to a relay chain by a group of microorganisms acting as a consortium or community, such that the intermediate released by one group of microorganisms within the consortium is utilized / metabolized by another group of microorganisms within the consortium, and this sharing of the released intermediates continues until the released end product is environmentally benign or the end product is assimilated into microorganisms thriving in the environment where the sample was collected. Subsequently, a multidimensional pollutant pathway organism matrix (PPOM) was created, which includes the target pollutant, its complete degradation pathway, validated subpathways of the complete degradation pathway, experimentally characterized organisms in the pathway / subpathway, and literature or manually curated information on the environmental niche (e.g., soil, water, sediment, etc.) from which the organism was isolated.

[0054] According to an embodiment of the present disclosure, the knowledge base module 110 is also configured to create a genome-pathway master map (GPM). A GPM is a multidimensional map that provides information about which subpathways of a given pollutant degradation pathway are present in the genome of a candidate microorganism. The creation of a GPM depends on the following steps. Initially, using literature mining and manual curation techniques, a large number of microorganisms found in various environmental samples (e.g., soil, water surface, sediment, etc.) and present in multiple environmental niches are managed, and a database of the most abundant environmental bacteria-microorganisms (DEBGs) and the environments / environmental niches in which their corresponding microorganisms are known to thrive is created. The environmental niche in this disclosure refers to the environmental conditions where it is known that an organism exists or colonizes (obtained through literature mining) and evolves to survive. Any other source of information about microorganisms inhabiting different environments is within the scope of the present invention. Information about the gene sequences and locations of available, sequenced bacterial genomes is obtained from the National Center for Biotechnology Information (NCBI). These bacterial genomes have been annotated functionally to identify protein domains within each gene on the genome using a variety of methods, which may include but are not limited to gene homology (BLAST, etc.), HMM-based identification (protein family or PFAM database, etc.), position-specific scoring matrix (PSSM), etc. In one embodiment, a database PfamDB (or protein domain family database) containing HMMs corresponding to all protein domains can be obtained as taught in the PFAM database. It is also within the scope of the present disclosure to use any other protein domain or functional protein annotation method or database as PfamDB. It is fully within the scope of the present disclosure to use information on gene sequences and positions on available, sequenced bacterial genomes from any other source.

[0055] Identify the protein domain corresponding to each candidate enzyme in each of the subpathways and pathways listed in PPOM. In one embodiment, PfamDB is used as a database, and data collection based on Hidden Markov Model (HMM) is performed to candidate enzyme to obtain the functional protein domain information in each protein. Any other method can be used to obtain functional information. This information is used to create a map (i.e., pathway domain map (PDM)), which comprises all subpathways and their associated protein domains in the pathway of degrading specific pollutants, covering all microorganisms stored in DEBG. For each genomic sequence corresponding to the microorganism in DEBG, information about the position and genome arrangement of its constituent genes is listed. According to the genomic position of each bacterial genome obtained from NCBI, the gene list arranged in the order of genes on the genome is placed in the map called genome map (GM). GM also comprises information about the functional annotation of each of these genes in the form of the constituent protein domains of each gene.

[0056] In addition, for the key subpathways found in the hash PDM, the numerical value is a list of corresponding domains. The genes of pathways / subpathways are often found to occur near the genome of microorganisms and are referred to as gene clusters. In order to form functional gene clusters, the distance on the genome that the set of domains that form pathways or subpathways should be located at changes and is usually defined using artificial and literature-based management. In this embodiment, this distance is defined based on the number of genes in the genomic position (called window size) of the gene, and the domain should be located in this genomic position to indicate a gene cluster and therefore the existence of pathways / subpathways. In the genome atlas (GM), each relevant protein domain (pfam) of each key subpathway in the PDM is searched to find whether the protein domains that form subpathways appear together on the genome as a gene cluster, thereby being located within the window size defined on the genome. In this embodiment, a window of 20 genes upstream and downstream of the query protein domain (referring to any one protein domain in the subpathway) on the genome is used. The existence of other protein domains in the subpathway in the window (20 in this case) is recorded in the form of domain assignment based on gene name or pfam database. In some embodiments, the present invention provides the method for the generation of the sub-pathway of the present invention.Window size can change according to various factors, such as involved candidate pathway and domain.If the sub-pathway in the genome is contributed and appears at window size (such as, if the window size of 20 genes, the +20 and-20 of query protein domain) the quantity of the domain in crosses a threshold value (can change for each sub-pathway, and use document excavation and manual management to obtain), then think that this sub-pathway exists.Threshold value refers to the minimum threshold value number of the domain that needs to exist for confirming the existence of sub-pathway and the corresponding ratio of the sum of the domain corresponding with this sub-pathway in PDM.From the information recorded, create multidimensional matrix with genome name and approach / sub-pathway information of every kind of pollutant in one or more pollutants identified in the described sample collected for degrading.This is called as genomic pathway master collection of figures (GPM).

[0057] The GPM atlas is provided with a numerical value of 0 or 1 based on the first predefined criterion. The first predefined criterion is for each subpathway in the bacterial genome. If the corresponding number of the subpathway protein domains recorded in "PDM" does not occur or does not reach the threshold value in the window size of 20 genes in the genome, then a numerical value of 0 is assigned to the bacterial genome. If the number of subpathway protein domains is higher than the threshold value (defined as candidate pathway using literature mining and manual management) and is present in the window of 20 genes on the microbial genome, a numerical value of 1 is assigned. Depending on the system and candidate pathway, the window size can be variable. Finally, the results are verified according to the list of organisms from the PPOM matrix, wherein the pathway has been experimentally characterized to eliminate erroneous results.

[0058] The GPM map generated in the previous step provides information about which subpathways with protein domain data (pfam) exist in the genome and a numerical value of 1 or 0 for each subpathway. However, if there are protein domains whose number exceeds the subpathway threshold, it is sometimes impossible to determine whether there is a pathway in these microorganisms that can degrade a specific pollutant. Some protein domains corresponding to the enzyme encoded by the gene can be promiscuous and may have multiple copies on the bacterial genome, each copy involving different functions and binding different substrates. In this case, further verification is needed to annotate gene / protein function, which is done by performing active site analysis on the enzyme in this embodiment. Some protein domains belong to categories involved in multiple functions in microorganisms. In addition, these domains do not constitute part of a gene cluster or operon and therefore cannot be distinguished from their other homologs based on their genomic neighborhood. In order to understand the substrate specificity of these domains, the amino acid pattern corresponding to their active site needs to be considered.

[0059] According to an embodiment of the present disclosure, the knowledge base module 110 is further configured to create a genomic pathway enzyme map (GPE map). GPE maps and GPM maps help identify complete pollutant degraders and partial pollutant degraders.

[0060] Initially, for the chosen threshold, all subpathways with a value of 1 for at least one genome in the GPM map are filtered candidate pathways (CPs) for further validation by active site analysis. All subpathways with a value of 0 in the GPM are rejected. For each candidate pathway (CP) and its constituent enzymes, literature mining is performed to list specific patterns of active sites of the query enzymes in the candidate pathway. The set of patterns for each candidate enzyme of the hypothetical pathway (ECP) is P ecp .

[0061] Active site is the region of enzyme (protein), and this region combines with specific substrate to react and comprises the pattern that is called motif (motif).These motifs serve as identification sequence, and identification sequence helps to identify whether enzyme is functionally able to combine with substrate, and is all conserved in enzyme with similar function.In this embodiment of candidate enzyme (ECP), multiple sequence alignment (MSA) is carried out to all possible functionally similar enzymes.The list of all homologues of ECP is identified by sequence similarity method.Any other method of homologue identification is all within the scope of the present invention.Subsequently, these homologues of candidate enzyme are carried out multiple sequence alignment (MSA) to identify the conservative pattern of enzyme.These conservative patterns are verified in the literature, to assess its functional importance.The pattern that is located in active site and verified by literature is called P ecp Any other method for active site identification is within the scope of the present invention.

[0062] In addition, a genome-pathway-enzyme map (GPE map) is created. The GPE map stores the active site information of each enzyme corresponding to the catalysis of each step in each sub-pathway corresponding to the degradation pathway of one or more pollutants identified in the collected sample (about P ecp This information is obtained and recorded in the GPE profile for each bacterium (and its genome) contained in the DEBG. Based on a second predefined condition / criterion, a value of 0 or 1 is assigned to the GPE profile corresponding to each genome. The second predefined condition / criterion is that the value 1 is assigned to those enzymes for which the corresponding active site pattern is found, and the value 0 is assigned to those enzymes for which the pattern is not found.

[0063] Some pollutants cannot be absorbed by microorganisms due to their large size. Therefore, since the candidate enzyme may be secreted into the extracellular environment to degrade the polymer into its monomer or semi-degraded state, it is necessary to test the presence of signal peptides in the candidate enzyme ECP to determine its secretion ability. Before the monomers or semi-degraded intermediates of the pollutants are absorbed by the bacterial cells, some enzymes are secreted outside the bacterial cells to act on the pollutants. These secreted enzymes are identified by the presence of signal peptides in the protein sequence. In this example, the test was completed using the SignalP 4.1 server, which was further verified using literature mining. Any other method for identifying the secretion ability of enzymes is within the scope of the present invention. The value 1 is assigned to those enzymes in which secretion ability is found and the value 0 is assigned to those enzymes in which secretion ability is not found. Each P ecp All information is stored in GPE.

[0064] Thus, for example, the three subpathways (P i SP1, P i SP2 and P i SP3) composed of separated pollutants P i The given path P iDP, if all subpathways of a microorganism in the GPM map have a value of 1 and the enzyme components of these subpathways have a value of 1 in the GPE map, then the microorganism will be called a complete pollutant degrader (CPD). Microorganisms with genomes that have one or more subpathways (but not complete pathways) that degrade pollutants will be called partial pollutant degraders (PPD). Each PPD has a value of 1 in the GPM, and each enzyme corresponding to these subpathways present on the genome of the PPD has a value of 1. Multiple organisms marked as PPDs can provide various subpathways to achieve complete degradation of pollutants in combination and can jointly form a microbial consortium / community to completely degrade pollutants. In this case, one or more microorganisms will partially degrade the pollutant into CIM, which can be released into the environment and absorbed by another group of microorganisms, which can utilize / metabolize the CIM and degrade it into another CIM that can be released into the environment. This combined metabolic process involving different groups of microorganisms can continue until metabolites that can be completely assimilated by the microbial consortium are obtained or the metabolites will not cause harm to the environment when released by the microorganisms. Combinatorial utilization of compounds by bacterial consortia can also be designed so that degradation results in CIMs that can be used for other applications, including industrial and other commercial purposes. This type of consortium does not produce a final product that can be assimilated by the microorganisms and does not completely degrade the compound, but rather can produce intermediate compounds that can be isolated and used in a variety of industrial applications.

[0065] Therefore, the pollutant-degrading microbial community meets the following criteria:

[0066] 1. Indicates the presence of key functional domains of subpathways obtained by literature mining and manual curation.

[0067] 2. All subpathways for pollutant degradation are present in the genome corresponding to the microbial community, to achieve complete degradation or partial degradation to obtain intermediate products that can be redirected for multiple applications. All subpathways do not need to be present in a single microbial genome, but should be present in a "microbial genome community" to account for the effectiveness of the microbial community on pollutant degradation.

[0068] 3. The presence of an active site pattern in the enzyme that corresponds to the binding of a specific substrate involved in each reaction of the identified subpathway in the microorganism in which the subpathway is identified.

[0069] 4. If the degradation pathway requires extracellular digestion of the pollutant substrate, secretory capacity is present in the secretions involved in the reaction.

[0070] DEBG is also updated with identifier tags for complete and partial degraders of pollutants. In addition, a list of complete pollutant degraders is created, and a list of partial pollutant degraders is created. Finally, DEBG is updated with the tags PPD and CPD that make up the genome.

[0071] The multidimensional pollutant pathway organism matrix (PPOM), genome-pathway-enzyme map (GPE), genome-pathway-master map (GPM) and the most abundant environmental bacteria-microorganism database (DEBG) together constitute the knowledge base (main backend) in this embodiment. The knowledge base provides information about the sub-pathways for degrading pollutants and the intermediates that may be released into the environment by the organisms present in each genome. In addition, it also provides information about the total pollutant degradation potential of the genome. The knowledge base can be pre-created and stored in a memory for a group of well-known common multiple pollutants, which may include but are not limited to plastics, such as polyethylene terephthalate (PET), styrene, polyurethane, etc., polycyclic aromatic hydrocarbons (PAH), such as naphthalene, anthracene, pyrene, etc., and different congeners of polychlorinated biphenyls (PCBs), etc. The knowledge base module 110 can be used to further fill and expand the knowledge base with another set of multiple pollutants, which is not included in the pre-created knowledge base and can be identified in the environmental site where the sample is collected.

[0072] According to an embodiment of the present disclosure, the memory 106 also includes a community association module 112. The community association module 112 is used to create a microbial map, which includes information on one or more partial pollutant degraders and complete pollutant degraders that can degrade each of the one or more pollutants identified in the sample to the same degree of non-degradation. The information also includes the environmental site where the sample was collected. This information is collected using DEBG, GPM maps and GPE maps. The different degradation degrees of the pollutants refer to the degradation of pollutants into different intermediate compounds or metabolites, which are determined by the final products released by the degraders under the action of genes or proteins or enzymes corresponding to the sub-pathways present in the genome of the degraders that degrade the pollutants. These intermediate compounds can be released into the environment and utilized by other microorganisms in the environment, or can be assimilated by the same microorganism that performs such degradation (210);

[0073] The community association module creates / predicts a microbial community that as a whole has the functional ability to completely degrade the isolated pollutant. Organisms from the GPE matrix with a value of 1 are filtered to obtain enzymes for which active site analysis has been completed. This result set is RS1. Similarly, for the isolated pollutant (P i), the organisms from the GPM matrix with a value of 1 corresponding to their subpathways are filtered. This result set is RS2. Organisms with no active site pattern (value of 0) for the query enzyme in GPE are filtered out, and organisms without subpathways are filtered out. The combined results are the candidate organism result set (CRS com ), these candidate organism result sets (CRS com ) in combination have the functional potential to degrade each of the multiple pollutants separated. A subset of these organisms can be selected so that they partially degrade the compound / pollutant to an intermediate level, wherein the products formed can be used in a variety of industrial applications. A subset of these organisms can be selected so that they partially degrade the compound / pollutant to an intermediate level, wherein the products thus formed can be directed to a variety of industrial applications.

[0074] In addition, corresponding to CRS com The environmental information of each organism in DEBG is obtained from DEBG. This information is used to create a pollutant organism environment matrix (POEM). Matrix POEM is composed of pollutants, organisms (depending on the complete pathway or sub-pathway of existence) and the environment sampled by the organism that can carry out different degrees of degradation to this pollutant. This matrix representation can be used in combination as needed to degrade the organism of pollutants completely or partially. Therefore, for environmental samples, all combinations of organisms that can survive in the environment and functionally can partially or completely degrade target pollutants can be concocted, and customized microbial communities can be designed. Customized microbial communities can be made up of microorganisms, and wherein, the different sub-pathways in each microorganism can combine (cumulatively forming complete pollutant degradation pathways) to partially / completely degrade pollutants, even when single composition microorganism lacks this degradation ability.

[0075] In one embodiment, the above method can be used to design a microbial community with the function of degrading multiple pollutants. The minimum community required for the multiple pollutants in the contaminated site can be identified using a knowledge base. This minimum community includes a group of microorganisms that can survive under environmental conditions and has a sub-pathway corresponding to the complete degradation of multiple pollutants. Therefore, the microorganisms belonging to these communities can provide all the sub-pathways in combination to degrade each pollutant identified in the environmental contamination site completely. It should be understood that a microorganism may also be responsible for providing a group of sub-pathways for degrading more than one pollutant. The microbial community designed like this includes such microorganisms, that is, these microorganisms are the complete or partial degradation agents of one or more pollutants present in the sample collected, and the microbial community designed like this can be used to design the first microbial community, which includes the microorganisms that can co-exist in the environment and degrade one or more polluted environmental sites collected by the sample in combination. Degradation will depend on the presence of all sub-pathways (genes and / or proteins and enzymes corresponding to the sub-pathways) for degrading the pollutants identified in the environmental site in the genome of a group of microorganisms forming a consortium.

[0076] The methods described in the present invention can also be reused for applications other than bioremediation of pollutants. In one embodiment, the methods described can be used to produce compounds with commercial uses in the industrial field. For example, when bacteria bioremediate polyethylene terephthalate (PET), terephthalic acid (TPA) and ethylene glycol (EG) are produced as intermediates. Using these intermediates as raw materials, TPA is used in multiple industries in the industry that use polymers and polyesters, such as packaging, textiles, etc. Subsequently, another group of bacteria can be added to the microbial community, and these bacteria can convert EG into compounds for industrial use, such as glycolic acid, which is widely used in the cosmetics industry. Polyhydroxyalkanoates (PHA) are the most common type of bioplastics in industry, and they can also be produced using TPA as a raw material and by using a group of bacteria (such as those that can be obtained from the knowledge base) responsible for converting TPA into PHA to enhance the microbial community used to form TPA from PET. In another embodiment, the methods described here can be used to convert intermediates to recycle the parent compound for industrial use. For example, TPA and EG obtained from the bioremediation of PET contaminants can be used to make new PET polymers, which can then be used in industry to produce a variety of PET-based products. Thus, the intermediates formed as CIMs in this process can be isolated and used for a variety of industrial purposes. In these embodiments, the consortium is designed to degrade the contaminants / compounds only into intermediate compounds that can be isolated for further industrial use, rather than completely degrading the contaminants.

[0077] In another embodiment, the method described herein can be used to identify microorganisms that cause important industrial compound degradation and thus cause huge losses to industry. For example, asphalt used in road construction is sometimes degraded by bacteria, resulting in pinholes formed on the road surface, thereby hindering its structural stability. This bacterium with the pathway and corresponding sub-pathway for degrading asphalt can be identified and targeted using the method described in the present disclosure. The present invention described herein can be used for any other method of industrial application within the scope of the present disclosure. These microbial communities comprising a combination of partial degradation agents for the metabolism or degradation of one or more pollutants can be designed so that intermediate compounds biosynthesized by the partial degradation agents can be obtained and reused in multiple industrial applications. This can be used as a second microbial community to partially degrade pollutants to obtain intermediate compounds with industrial importance, which can meet but are not limited to industries such as packaging, automobiles, oil and gas, food and beverages, textiles, paints and lubricants.

[0078] According to an embodiment of the present disclosure, the system 100 further includes an application module 114. The application module 114 is configured to apply the designed custom microbial community to the environmental site where the sample was collected. The application results in complete and effective degradation of multiple pollutants from the site.

[0079] The application methods of bioremediation technology vary depending on the type of contaminated site, the degree of contamination, the location, the cost, and the environmental policies specific to the site. Different pollutants have been observed to contaminate various locations, such as soil, wastewater, industrial sludge, and aquatic environments (bodies of water, particularly oceans, lakes, and rivers). Application methods can be broadly divided into two categories: ex situ bioremediation and in situ bioremediation.

[0080] In the ex situ method of bioremediation, pollutants are excavated from the contaminated site, transported to another area for treatment, and then ultimately returned to the site after treatment.

[0081] In the case of in situ methods, bioremediation occurs at the contaminated site itself.

[0082] While ex situ methods are more effective, they are not economically viable when targeting larger contaminated areas. Although many types of bioremediation methods exist and are used industrially, in this example, two categories of bioremediation are disclosed below based on the contaminated site and the type of bioremediation.

[0083] Ex situ application: The methods used for ex situ bioremediation vary according to the phase of the contaminated material and can be divided into: (i) solid phase systems (using techniques such as land cultivation, soil heaping and composting) and (ii) slurry phase systems (involving the treatment of solid-liquid suspensions in bioreactors).

[0084] Solid phase systems can be used for large amounts of waste and require favorable conditions, such as moisture content for microbial growth, frequent aeration, mixing (mechanical and air mixing), pH, and inorganic nutrients. During land cultivation, contaminated soil is laid in a lined bed (to prevent leakage) and the soil is mixed regularly to provide nutrients and oxygen for the microorganisms. In the case of biopiling, contaminated soil samples are placed in piles on top of a bug vacuum pump. The vacuum pump maintains a steady flow of oxygen to keep the samples well aerated and adds nutrients to accelerate the bioremediation process. Each condition is monitored to ensure effective bioremediation.

[0085] In a slurry-phase system, contaminated solid material from the application site, along with microorganisms and water (all components formulated as a slurry), is introduced into a bioreactor. A bioreactor is a large vessel that converts the feedstock into various products through a series of biological reactions. Bioremediation in a bioreactor is one of the most common methods for treating contaminated soil / water. During this process, the bioreactor is maintained under optimal conditions for microbial growth, and contaminants in the feedstock (contaminated soil / water) are metabolized. The microorganisms added here are pollutant-degrading microorganisms or microbial communities identified through our pipeline. The feedstock can be any sample extracted from any contaminated site. The required biological process parameters (such as temperature, pH, agitation and aeration rates, substrate and inoculum concentrations) can be externally controlled, making it a preferred technology because it can effectively increase the bioremediation rate. After treatment, the treated soil / water can be returned to its original site. One advantage of bioremediation using a bioreactor is the use of engineered microbial communities. Because it is a closed system, the engineered microorganisms can be destroyed before the treated soil / water is returned, ensuring that the engineered microorganisms do not enter the ecosystem.

[0086] In situ application: In situ bioremediation techniques are relatively low cost compared to ex situ methods as they do not involve any excavation. However, the method does require complex equipment for improving microbial activity, the cost of its design as well as on-site installation adds to the expense of the process. In situ methods of bioremediation technology may be carried out naturally as in intrinsic bioremediation or with the help of some enhancements (bioventing, biospraying and phytoremediation). Microorganisms have the innate ability to degrade metabolites and consume them as a carbon source. Bioremediation methods that utilize and manage the existing capacity of naturally occurring microorganisms to degrade pollutants without applying any engineering steps to enhance the process are classified as intrinsic bioremediation. Bioventing involves the continuous supply of a steady flow of oxygen along with nutrients and water to the unsaturated (seepage) zones of the site to enhance the activity of inherent microorganisms to degrade pollutants in contaminated soil or water.

[0087] The microorganisms identified as pollutant degraders in this method can be used to achieve efficient bioremediation activity. The above-mentioned methods are some ways to bioremediate an environment contaminated by pollutants. Any other acceptable methods are within the scope of the present invention.

[0088] According to an embodiment of the present disclosure, system 100 also includes an efficacy module 116. In efficacy module 116, after the application of pollutant degradation microorganisms or microbial communities, the contaminated site must be evaluated at frequent intervals to check the presence of pollutant contamination and the rate of pollutant degradation. Any method discussed in pollutant identification module 104 can be used to evaluate the efficacy of the designed microbial community on the bioremediation of environmental pollutants. Any other accepted method for detecting the presence of pollutants is within the scope of the present disclosure. Based on the level of pollution existing after detection, the first microbial community and the second microbial community can be modified as necessary to accelerate the bioremediation of these pollutants. Modifications include but are not limited to: inoculating bacterial communities designed as multiple titers of the first microbial community and the second microbial community, modifying the applied bacterial community and strengthening it with additional CPD and PPD microorganisms that can survive in the environmental niche based on information stored in the knowledge base, adding necessary nutrients, better ventilation to the contaminated environment, etc., to further promote pollutant degradation.

[0089] In operation, Figures 2A to 2C A flow chart 200 is shown illustrating the steps involved in bioremediation of a contaminant. Initially, at step 202, a sample is collected from an environmental site containing one or more contaminants. At step 204, the one or more contaminants present in the sample are separated and identified.

[0090] In step 206, a knowledge base is created. The knowledge base stores information about one or more pollutants that have been identified, information about complete degradation pathways and partial degradation pathways identified in microorganisms that can completely degrade one or more pollutants or partially degrade one or more pollutants, information about the various environmental niches in which the microorganisms thrive, and a list of microorganisms from different environments that have specific complete / partial degradation pathways for pollutants. A complete degradation pathway refers to a set of genes on a microbial genome and / or proteins encoded by the microorganism, wherein this set of genes and / or encoded proteins is responsible for completely degrading the pollutants into compounds that are safe for the environment or compounds that can be assimilated by other microorganisms present in the environment. A partial degradation pathway in a microorganism refers to a set of genes or encoded proteins that constitute one or more sub-pathways. A sub-pathway is a subset of a complete degradation pathway encoded within a microbial genome, and a sub-pathway degrades the pollutant into an intermediate compound that can be released by the microorganism into the environment and subsequently absorbed by another microorganism in the environment, wherein the other microorganism has another sub-pathway that metabolizes the released intermediate compound.

[0091] In step 208, a list of partial pollutant degraders and a list of complete pollutant degraders are identified for each of the one or more pollutants identified in the sample using information from the knowledge base. A partial pollutant degrader is a microorganism that provides one or more sub-pathways and a corresponding set of genes, encoded proteins, or enzymes for converting pollutants into intermediate compounds. Multiple partial degraders can, in combination, provide all sub-pathways for completely degrading the pollutants identified in the collected sample. A complete pollutant degrader is a combination of all sub-pathways and a corresponding set of genes, encoded proteins, or enzymes for degrading pollutants in the collected sample in a single microorganism.

[0092] In step 210, the information in the knowledge base is utilized to create a microbial atlas.The microbial atlas comprises the information of one or more (each pollutant degradation in one or more pollutants identified in the sample can be degraded to different degradation degrees) in a partial pollutant degradation agent and a complete pollutant degradation agent.The different degradation degrees of pollutant refer to that pollutant is degraded into different intermediate compounds or metabolites, and intermediate compounds or metabolites are determined by the final product, and this final product is discharged by the degradation agent under the effect of the gene corresponding to the sub-pathway existing in the genome of the degradation agent for degrading pollutants or protein or enzyme.Intermediate compounds can be released into the environment and utilized by other microorganisms in the environment, or can be assimilated into the same microorganism that carries out this degradation.

[0093] In step 212, the created microbial profile is used to design a first microbial community, the first microbial community comprising microorganisms that collectively provide the required subpathways for complete degradation of one or more pollutants identified in the sample, and wherein the microorganisms are able to collectively survive in the same environmental niche from which the sample was collected.

[0094] In step 214, a second microbial community is designed using the created microbial profile. The second microbial community includes microorganisms that collectively provide genes, proteins, and enzymes for sub-pathways required to partially degrade one or more contaminants identified in the collected sample into one or more desired intermediate products. The microorganisms forming the second microbial community are capable of surviving together in the environmental niche from which the sample was collected.

[0095] At step 216, at least one of the first microbial community and the second microbial community, or a mixture of both, is applied to the environmental site containing the one or more contaminants. The application method of the bioremediation technology depends on the type of contaminated site, the degree of contamination, the location, the cost, and the site-specific environmental policies. Application methods can be broadly categorized into two types, namely, ex situ bioremediation and in situ bioremediation, as described above.

[0096] In step 218, the efficacy of the applied mixture in eliminating one or more pollutants in the sample collected from the environmental site is checked. The efficacy is evaluated by separating and identifying the residual pollutant combination from the sample collected. In step 220, a new mixture is reapplied to the environmental site. A new mixture is prepared by adding a group of microorganisms that can act as partial degradation agents and degrade one or more pollutants in the sample collected in combination. The previously applied mixture can also be enhanced by adding other microorganisms that can act as partial degradation agents and degrade one or more of the pollutants identified in the sample collected in combination.

[0097] According to an embodiment of the present disclosure, a flowchart 300 for creating a knowledge base is shown as follows: Figures 3A to 3CAs shown. Initially, in step 302, literature mining techniques are used to identify degradation pathways for a plurality of isolated pollutants. Literature mining also generates information about a set of microorganisms in which the pathways have been experimentally characterized and the environmental niches from which these microorganisms are isolated. Similarly, in step 304, a plurality of subpathways are also identified within the degradation pathways that lead to partial / complete utilization / assimilation of the plurality of isolated pollutants. In one embodiment, each of the plurality of subpathways is present in the genome of a different microorganism referred to as a partial pollutant degrader (PPD), and the products formed by each of the plurality of subpathways are released into the environmental site and metabolized or taken up by other microorganisms inhabiting the environment. In another embodiment, each of the one or more subpathways is present in the genome of a microorganism referred to as a complete degrader (CPD);

[0098] In step 306, a pollutant pathway organism matrix (PPOM) is created using the identified degradation pathways for each of the one or more identified pollutants, multiple subpathways of the degradation pathways, a set of organisms in which the degradation pathways are characterized, and information about the corresponding one or more environmental niches from which the set of organisms was isolated based on literature mining and manual curation. In step 308, a rich environmental bacteria-microbe database (DEBG) is created using literature mining techniques. The DEBG contains information about all microorganisms and the different environmental niches in which microorganisms thrive.

[0099] In step 310, a pathway domain map (PDM) is created from a pre-created protein family database (pfamDB), wherein the protein domains included in the PDM correspond to genes / proteins that constitute multiple subpathways, each of which is included in the degradation pathways created for a plurality of pollutants. In step 312, a genome map (GM) is created, wherein the genome map provides a list of genes / proteins (ordered by their respective genomic locations) in a microorganism and the constituent protein domains within these genes / proteins. In the next step 314, the microorganism genome stored in the DEBG is searched for the presence of protein domains included in the PDM for each of the multiple subpathways of all pathways listed in the PPOM to determine the occurrence of these subpathways on the genome; wherein the genome map GM is used as a database for this search, and wherein if the number of domains in the genome listed in the PDM that provide a subpathway appears within a gene window size on the genome and exceeds a predefined threshold, then the subpathway from the PDM is considered to be present.

[0100] In step 316, for each of the one or more pollutants identified in the collected sample, a genomic pathway master map (GPM) is created using the microorganism name corresponding to the microorganism genome in the DEBG and information about the presence or absence of multiple pathways and multiple subpathways on the genome, wherein the GPM map has a value of 0 or 1 based on a first predefined criterion, and wherein the GPM provides information about all subpathways of a given pollutant degradation pathway present in each of the microorganism genomes listed in the GPM.

[0101] At step 318, a genomic pathway enzyme profile (GPE) is created for each of the one or more pollutants identified in the collected sample, wherein the GPE includes the names of all microorganisms listed in the DEBG and information about the active sites of each enzyme involved in each step of multiple subpathways on each genome, wherein the GPE profile has a value of 0 or 1 based on a second predefined criterion. The pollutant pathway organism matrix (PPOM), the GPE profile, the GPM profile, and the DEBG together form a knowledge base.

[0102] According to embodiments of the present disclosure, the system 100 can be used for bioremediation of a variety of pollutants, including but not limited to: plastics (polyethylene terephthalate (PET), styrene, etc.), rubber, pesticides, synthetic fertilizers, electronic waste, industrial waste, food additives, cleaning products, cosmetics, dyes, etc. Figure 4 In another embodiment, in addition to removing contaminants, the system 100 can also be used to reuse the products and intermediates of the process for various other applications in the industry in which they are included.

[0103] In one embodiment of the present disclosure, system 100 can be applied to the bioremediation of carbon-based pollutants, such as carbon-based nanomaterials (CBNMs), PAHs, PCBs, and PET. Any other carbon-based pollutants are within the scope of the present invention. In the case of CBNM degradation, the method can follow a two-step process to determine whether the bacterial community or bacteria is capable of degrading CBNM, such that the pollutant is converted into a biologically and environmentally safe substance or is completely assimilated into bacteria present in the environment from which the CBNM is isolated as a pollutant. The first step is to determine the presence of key secreted peroxidases in the bacteria or bacterial community. Any other enzymes capable of degrading CBNM are also within the scope of the present invention. The second step is to identify the presence of aromatic degradation capabilities (such as PAHs, monocyclic aromatic hydrocarbons (SAHs), and PCBs). Bacteria are known degraders of PAHs and PCBs. The method searches for the presence of genetic mechanisms in the bacterial genome that are critical for PAH and PCB degradation. The method hypothesizes that if these two characteristics are present in the bacteria or can be provided in combination by members of the bacterial community, they are capable of completely degrading CBNM.

[0104] CBNM is of the order of 1 nm to 1000 nm (although most of it falls within the range of 10 nm to 100 nm) and therefore cannot be internalized by bacteria with a size of 2 μm (i.e., 2000 nm) for degradation. Therefore, the initial degradation of CBNM may occur outside the bacterial cells, mainly through peroxidases. However, it is not clear what bacterial peroxidases may degrade nanomaterials. In this embodiment, the bacterial enzyme catalase-peroxidase (kat) is identified as a potential peroxidase that may be able to degrade CBNM. Additionally, it has been identified that the presence of secreted catalase-peroxidase in bacterial communities or bacteria is crucial for the degradation of CBNM.

[0105] A variety of eukaryotic peroxidases have been experimentally shown to degrade CBNM, among which the plant-secreted peroxidase horseradish peroxidase (HRP) can degrade various types of CBNM, such as single-walled carbon nanotubes (SWCNTs), graphene oxide (GO), reduced graphene oxide (RGO) and multi-walled carbon nanotubes (MWCNTs), etc. In the present disclosure, it is assumed that bifunctional secreted catalase-peroxidase may have the ability to degrade CBNM in bacteria. Any other enzyme capable of degrading CBNM is within the scope of the present invention. Although catalase-peroxidase is a prokaryotic peroxidase, it has a high structural similarity to HRP. It is known that there are distal active sites and proximal active sites in catalase-peroxidase, and CBNM ligands can bind to the proximal or distal active sites of the enzyme, such as Figure 5 As shown. The proximal and distal active site cavities are arranged with various aromatic amino acid residues, such as tryptophan (Trp), tyrosine (Tyr) and phenylalanine (Phe), as well as various other polar residues such as arginine (Arg). These amino acid residues may help to stabilize the binding of CBNM in the proximal and distal active sites. In addition, the distal active site of the enzyme is connected to the central heme cavity (central heme cavity) through a series of non-polar aromatic amino acid residues. The electron transfer from the heme active site to the distal cavity bound to CBNM may occur through the aromatic amino acid bridge in the enzyme, especially via the electron hopping of the W176 residue. These results indicate that secreted bifunctional catalase-peroxidase may be a bacterial peroxidase capable of degrading CBNM.

[0106] The bifunctional catalase-peroxidase belongs to Family III of the peroxidase superfamily. The primary function of this enzyme is to scavenge H₂O₂, thereby protecting bacterial cells from oxidative stress. Catalase-peroxidase may undergo the following reactions in the presence of CBNM contaminants.

[0107] CBNM+H2O2 Oxidized-CBNM+H2O2

[0108] 2H2O2 O2+2H2O

[0109] In the present disclosure, it has been shown that bifunctional catalase-peroxidase can degrade CBNM, provided that they are secreted outside the bacterial cell. This allows them to access the larger CBNM for degradation. Since catalase-peroxidase is highly conserved in all bacteria containing the enzyme, similar results are expected for all enzyme homologs present in other bacteria. CBNM is converted into intermediates such as PAHs or PCBs, which are also pollutants and require further degradation to achieve complete degradation of CBNM. PAHs and PCBs themselves are part of various industrial wastes and are responsible for environmental pollution. Therefore, it is also necessary to remove these compounds from the environment.

[0110] PAHs are organic compounds containing multiple aromatic rings composed of carbon and hydrogen, such as Figure 6A As shown. In this disclosure, low molecular weight PAHs (LMW PAHs), such as naphthalene, anthracene, and phenanthrene, were analyzed according to the described methods. Degradation methods for other PAHs are also within the scope of this disclosure. Using the literature mining methods used in our study, it was found that LMW PAHs are biodegradable and generally favor aerobic degradation through oxygen-mediated metabolism, followed by dehydrogenases and subsequent ring cleavage by dioxygenases to form TCA cycle intermediates that are readily absorbed by organisms. Pathway literature mining was performed in literature search engines (such as PubMed) and pathway databases (such as KEGG pathways / EAWAG (BBD / PPS) / MetaCyc). The query string could be PAH (e.g., naphthalene) degradation + aerobic + bacteria. Using the above mechanism, pathways, their corresponding enzymes, key intermediates, model organisms, and gene clusters were identified. Clusters of genes (encoding their enzymes) involved in each subpathway were searched across all genomes. Organisms harboring these clusters were further validated by the presence of structurally significant patterns found in the active sites of the enzymes. Multiple bacterial genomes or bacterial consortia that possess all subpathways and active site patterns are considered to be PAH degraders. Figure 6A The naphthalene degradation pathway shown includes two sub-pathways, namely, NSP1 and NSP2, in which i) naphthalene is converted to salicylic acid (NSP1), and ii) salicylic acid can be further converted to catechol (NSP2), which is ultimately degraded to compounds that can be assimilated by the bacterial genome through the cat gene cluster, which is evolutionarily conserved in bacteria. Similarly, tricyclic anthracenes are degraded through the combination of two sub-pathways, ASP1 and ASP2. Figure 6AAs shown, the former subpathway (ASP1) involves a series of enzymatic steps: i) conversion of anthracene to 2,3-dihydroxynaphthalene, followed by ii) conversion of the resulting compound to salicylic acid, which is further degraded via the catechol pathway (ASP2). Salicylic acid, an intermediate degraded via catechol metabolism, forms a common key intermediate (CIM) in the degradation of naphthalene and anthracene. The degradation pathway of phenanthrene, another tricyclic PAH analyzed in our study, is divided into three subpathways: i) phenanthrene is converted to 1-hydroxy-2-naphthoate, regulated by its gene cluster, to form phthalate (PSP1); ii) phthalate is degraded to 3,4-dihydroxybenzoic acid (protocatechuic acid) via a subpathway regulated by the pth gene cluster and ultimately assimilated into bacterial metabolism via benzoate degradation (PSP2). iii) phenanthrene can be degraded via a subpathway that forms 1,2-dihydroxynaphthalene, which is then further degraded via naphthalene metabolism (PSP3).

[0111] Other aromatic intermediates formed during the degradation of CBNM are biphenyls, which are the reduced (dehalogenated) forms of PCBs. The reductive dechlorination of highly chlorinated biphenyls to less chlorinated biphenyls involving the rdhABR gene cluster has been included as a subpathway (PcSP1), as Figure 6B As shown, the degradation of low-chlorinated biphenyl derivatives is caused by specialized communities of organohalide-respiring bacteria under anaerobic conditions. Under aerobic conditions, low-chlorinated biphenyl derivatives are further degraded to 2-hydroxypenta-2,4-dienoate and benzoate via the biphenyl dioxygenase (bphA) activity, which involves the bphABCD gene cluster and is known as the upstream pathway (PcSP2). 2-Hydroxypenta-2,4-dienoate is further degraded to pyruvate via a different, well-studied gene cluster, bphEFG, known as the downstream degradation pathway (PcSP3). A few organisms possess both PcSP2 and PcSP3 as a complete gene cluster for biphenyl degradation, designated PcSP4. These pathways, specific for PCB degradation, lead to the formation of intermediates such as benzoate. Benzoate degradation proceeds via the catechol or benzoyl-CoA metabolic pathways, designated PcSP5 and PcSP6, respectively.

[0112] According to an embodiment of the present disclosure, the system 100 can also be explained with the help of the following example of polyethylene terephthalate (PET). The above method can be used to infer the degradation ability of bacteria on PET-contaminants (polyethylene terephthalate).

[0113] Bacterial degradation of PET involves three major subpathways: (a) hydrolysis of PET to its monomers, such as bis(hydroxyethyl) terephthalate (BHET), mono(hydroxyethyl) terephthalate (MHET), or terephthalic acid (TPA); (b) conversion of MHET to TPA; and (c) reduction of TPA to protocatechuic acid (PCA). After a thorough literature search, the PETase enzyme from Ideonella sakeinsis was considered in this analysis because it has efficient activity in breaking down the polymer into its monomers. In this example, PSI-BLAST was used to identify its distant homologs. Other methods for collecting these homologs are also within the scope of this disclosure. This PETase exists as a monomer and structurally belongs to the α / β hydrolase superfamily, which is strictly conserved among all esterase proteins (such as lipases and cutinases). Similar to other α / β hydrolases, the enzyme PETase possesses a conserved catalytic triad S131-H208-D177 and a serine hydrolase motif Gly-x1-Ser-x2-Gly at the active site. However, the presence of two intramolecular disulfide bridges (DS1 and DS2) formed near the catalytic center constitutes a unique feature of PETase, while other hydrolases have only one disulfide bridge. Given the presence of the DS1 disulfide bridge and the conserved pattern within the α / β hydrolase superfamily, further filtering for distant homologs was performed. The sequences of these distant homologs were aligned, and their multiple sequence alignment was used to create a hidden Markov model (PET-Pfam) of PETase. PETase is a secreted protein and is known to be secreted into the extracellular environment, providing simple and optimized PET accessibility. The secretion ability of potential genes with PET-Pfam was confirmed by using the SignalP 4.1 server to identify signal peptides in the query genes.

[0114] The conversion of TPA to PCA is a two-step process involving two enzymes. The final output of this step (i.e., PCA) is a functionally important intermediate that is found to be conserved in bacteria. Therefore, the degradation of PET involves the conversion of PET by the action of the enzyme PETase, resulting in the formation of its constituent monomers (e.g., TPA), and is labeled PeSP1. The action of PETase is the rate-limiting step and is labeled as the specific pathway PeSP1. The subpathway for the degradation of TPA to form protocatechuic acid (PCA) is labeled PeSP2. PCA is degraded to form acetyl-CoA and is regulated by the pcaIJFHGBL gene cluster. This subpathway is common to the degradation of PET and even phenanthrene and is labeled PeSP3 (e.g., Figure 6C shown).

[0115] Bacteria that possess PETases but lack the TPA subpathway are potential PET degraders. Conversely, bacteria that possess the TPA subpathway but lack the PET-to-TPA conversion are potential TPA degraders. A bacterial population that can collectively provide all subpathways (TPA to PCA and PCA to catechol) and possesses PETases with the appropriate active sites and signal peptides is considered a complete PET-degrading microbial community. Bacteria that possess the subpathways, active site patterns, and signal peptides are considered complete PET degraders. A single microorganism can possess all subpathways within it, or a microbial community can provide subpathways that collectively and effectively degrade the pollutant. Due to size limitations, a list of organisms maintained by the inventors can be used to create microbial mixtures that include different combinations of microorganisms that provide each identified subpathway. This list is available to examiners upon request. In another embodiment, these mixtures will result in the complete degradation of PET into products that can be assimilated by the bacteria. Additionally, microbial mixtures can be designed to include organisms that can degrade PET to different intermediate levels, where the resulting products can have various industrial applications. For example, a provided list of organisms can be used to create a microbial consortium composed of multiple organisms capable of degrading PET to TPA and EG. This microbial consortium forms two intermediates (TPA and EG), which can be isolated and used in a variety of industrial applications.

[0116] According to an embodiment of the present disclosure, the system 100 may be explained with the help of examples of various tables utilized in a knowledge base.

[0117] Pollutant Pathway Organism Matrix (PPOM): This matrix includes a list of pollutants, the degradation pathways and identified subpathways that degrade the corresponding pollutants, the experimentally characterized organisms in these pathways, and the environmental niches from which these organisms are identified. Table 1 shows the format of the PPOM. A sample PPOM for a pollutant PET is shown in Table 2.

[0118] Table 1: Format of PPOM

[0119]

[0120] Table 2: Sample PPOM for contaminant PET

[0121]

[0122]

[0123] Database of Enriched Environmental Bacteria-Microorganisms (DEBG): This database includes the microbial genomes reported to date and the environments from which these microorganisms were isolated. Table 3 shows a sample DEBG.

[0124] Table 3: Abundant Environmental Bacteria - Microbial Database (DEBG) samples

[0125]

[0126] Protein Domain Database (PDM) from pfamDB: This matrix is ​​populated with protein domains that are associated with gene components in the subpathways / pathways listed in the PPOM for each pollutant. PfamDB includes a large database that provides information on different protein domains or functional annotations for various proteins. PDMs are made by searching for these domains listed in PfamDB in the desired genes. Table 4 shows a prototype of the PDM, and Table 5 shows a sample PDM for a pollutant PET.

[0127] Table 4: PDM prototype

[0128]

[0129] Table 5: Sample PDM for contaminant PET

[0130]

[0131] Genome Map (GM): stands for Genome Map, which provides gene list information based on the genomic position order of each bacterial genome and the protein domain composition or function information of each gene. An example of GM is shown in Table 6.

[0132] Table 6: Sample genome maps

[0133]

[0134] Genomic Pathway Master Map (GPM): This provides the genome name and corresponding pathway / subpathway information for each of one or more contaminants, where the GPM map has a value of 0 or 1 based on a first predefined criterion, where the criterion is to search for pathway-specific protein domains within a window of 10 adjacent genes and assign a value of 1 if the domain is present above a threshold and a value of 0 if it is absent, for all subpathways in the genome. Table 7 shows a prototype of the GPM. Table 8 shows an example GPM for PET contaminants.

[0135] Table 7: GPM prototype table

[0136]

[0137] Table 8: Examples of GPM of PET Contaminants

[0138]

[0139] Genomic Pathway Enzyme Map (GPE): GPE provides information on the active sites of each enzyme corresponding to a step in multiple sub-pathways of the genome and the presence of signal peptides for each enzyme in the GPE. Table 9 shows the prototype of GPE and Table 10 shows an example of GPE for PET pollutants.

[0140] Table 9: GPE prototype table

[0141]

[0142] Table 10: Examples of GPE for PET contaminants

[0143]

[0144] The prototype POEM matrix represents the subpathways of each degradation pathway for each pollutant and information about the organisms that contain these subpathways obtained using the methods discussed in this disclosure. The matrix also shows the complete / partial pollutant degradation capacity of each organism and the environmental niche from which these organisms were isolated. Information about each pollutant identified in the collected samples forms part of the POEM matrix.

[0145] Table 11: Sample POEM Matrix

[0146]

[0147] Possible consortia derived from the POEM matrix can be obtained based on the presence of subpathways, and microorganisms possessing these subpathways should be able to survive in the same environmental niche from which the samples were collected, as shown in Tables 12, 13, and 14. The first consortium for complete degradation of pollutant 1 in environment 1 can be obtained as follows.

[0148] Table 12: Combination 1 for Degrading Pollutant 1 in Environment 1

[0149] Sub-pathway 1 environment degradation Sub-pathway 2 environment degradation Organism 1 / Genome 1 Environment 1 part Organism 5 / Genome 5 Environment 1 part Organism 1 / Genome 1 Environment 1 part Organism 6 / Genome 6 Environment 1 part

[0150] It should be noted that the consortium that design is made up of the microorganism that only has sub-pathway 1, it will include organism 1 and organism 2, in this case, consortium will stop in organism 1 or organism 2 interior sub-pathway effect after release / produce on the intermediate and carry out pollutant degradation.The product intermediate of so releasing can be reused in multiple industrial applications.Therefore, comprise and show the consortium that has sub-pathway 1 organism (being organism 1 and 2 in this case) and can form the second consortium, it discharges the intermediate that can be used for industrial applications.Similarly, can decipher the consortium that is used to degrade the pollutant 1 in other environments.

[0151] Table 13: Degradation of pollutants 1 in the environment 3 by the combination 2

[0152] Organism 3 / Genome 3 Environment 3 part Organism 7 / Genome 7 Environment 3 part

[0153] Table 14: Combinations of 3 for Degrading Pollutants 1 in Environment 4

[0154] Organism 4 / Genome 4 Environment 4 part Organism 8 / Genome 8 Environment 4 part

[0155] Additionally, some examples of prototypes of POEM matrices for partial and complete degraders of PET are provided in Table 15.

[0156] Table 15: Prototypes of POEM matrices with several examples of partial and complete degraders of PET

[0157]

[0158] Possible combinations of subpathways 1 and 2 that provide complete degradation of PET based on the POEM matrix are shown in Tables 16 and 17 for soil and marine sediments, respectively.

[0159] Table 16: Possible combinations of subpathways 1 and 2 that provide complete degradation of PET based on the POEM matrix for several examples in soil.

[0160] Environmental preference: soil

[0161]

[0162]

[0163] Table 17: Possible combinations of subpathways 1 and 2 that provide complete degradation of PET based on the POEM matrix for several examples of sediments from marine environments

[0164] Environmental preference: Sediments in marine environments

[0165]

[0166] The consortium can use the following criteria for forecasting:

[0167] - There are subpathways 1 and 2 for the complete degradation of PET to PCA, where PCA can ultimately be

[0168] Bacterial assimilation

[0169] - The strains comprising the consortium should be isolated from or able to survive in the environmental niche from which the sample was obtained.

[0170] The method discussed in this example was used to identify the potential of bacteria to degrade major industrial pollutants such as polyethylene terephthalate (PET), polycyclic aromatic hydrocarbons (PAHs), and polychlorinated biphenyls (PCBs), as well as emerging pollutants such as carbon-based nanomaterials (CBNMs). Using this method for bioremediation of various other pollutants is also within the scope of the present invention. According to the method described in this example, the complete degradation of PET, CBNMs, PAHs, and PCBs involves multiple subpathways as described below.

[0171] For example, complete PET degradation involves the PETase enzyme subpathway, which converts PET to its constituent monomers, such as TPA. Candidate bacterial families involved in the PETase subpathway and the TPA to PCA subpathway include one or more of the bacterial families shown in Table 18. It should be understood that in this context, family refers to a taxonomic classification according to Linnaean taxonomy, and in this disclosure refers to strains of microorganisms in a given family that possess the genes / proteins / enzymes of the corresponding subpathway. Any other bacterial family with the potential to degrade PET is included within the scope of this disclosure.

[0172] Table 18: List of candidate bacterial families corresponding to various pathways of PET degradation

[0173]

[0174]

[0175]

[0176] According to an embodiment of the present disclosure, the system 100 is further configured to identify key enzymes for CBNM degradation. Key peroxidases are identified for the initial degradation of CBNM, wherein the presence of the enzyme is essential for CBNM degradation to occur. The key peroxidases form the initial step of the enzymatic degradation reaction for CBNM, and the intermediates formed are degraded in subsequent steps discussed further below:

[0177] First, literature mining techniques were performed to identify enzymes capable of degrading CBNM. i )Generate a query string (Q in The pollutants here are any type of CBNM, such as SWCNT, MWCNT, GO, RGO, etc. in ) is used as input to mine the selected literature search engines such as PubMed and pathway databases (e.g. KEGG pathway / EAWAG (BBD / PPS) / MetaCyc). The result set obtained from the literature search engine includes the summary A as output outList of enzymes for degrading CBNM (E out ), wherein the degradation of CBNM was characterized experimentally.

[0178] In the next step, the key bacterial enzymes (E bac ). Create a list of all potential bacterial enzyme candidates for CBNM degradation (E list ), compare the enzyme (E out These candidate bacterial enzymes were identified to have similar out The protein domain structure of enzymes. out The factors for comparison between these candidate enzymes include similarity at the protein and nucleotide sequence levels, comparison at the protein structure level, and similarity of residues forming the active site. out For comparison, for each E list Members are assigned points. out The enzyme with the greatest similarity was selected and considered as a potential bacterial candidate enzyme capable of degrading CBNM (E bac ).

[0179] In addition, to degrade large molecular weight CBNM, key bacterial enzymes (E bac ) need to be secreted outside the bacteria. bac ) secretion capacity and the presence in the extracellular region of bacteria. In one embodiment, E bac The presence of secretion ability in the enzyme was determined by two methods, which involved analysis of the presence of an N-terminal signal peptide and analysis of leader-free secretion ability based on the amino acid composition of the enzyme. bac ) of each genome Ge of strain S, analyze E bac The secretion capacity of the protein was evaluated and a score was derived for each secretion method tested (the presence of an N-terminal signal peptide resulted in a score of D, while secretion without a leader peptide resulted in a score of SP). bac The D score and SP score determine E bac secretion potential. Only those E bac The secretion potential is above the threshold score SO thre The bacterial species S was considered as a potential CBNM degrader. This secretion score was higher than that of So thre The bacterial species is called S NM In one embodiment, a threshold score of 0.79 was considered (So thre ), but it can vary depending on the method used and the enzyme system analyzed. Any other method of analyzing the secretion capacity of an enzyme is within the scope of the present disclosure.

[0180] According to an embodiment of the present disclosure, a candidate bacterial enzyme (E bac ) were identified as secreted bifunctional catalase-peroxidases and as a candidate bacterial family containing secreted bifunctional catalase-peroxidases (S NM ) were identified. Candidate bacterial families containing secreted catalase-peroxidase and involved in CBNM degradation are listed in Table 19. It should be understood that in this context, family refers to a taxonomic classification according to Linnaean taxonomy, and in this disclosure refers to microbial strains within a given family that have genes / proteins / enzymes for the corresponding subpathways. Any other bacterial family capable of degrading CBNM is within the scope of this disclosure.

[0181] Table 19 shows the detailed candidate bacterial families for CBNM degradation

[0182]

[0183] Since catalase-peroxidase (S NM ) during the degradation of CBNM, the intermediates formed showed structural similarities with PAHs, PCBs, and SAHs, which are aromatic compounds. For this reason, in the next step of CBNM degradation, the bacterial species S NM The presence of aromatic degradation in the bac Each bacterial species S NM , identifying the presence of aromatic hydrocarbon (such as PAH and biphenyl) degradation capability, as discussed in further detail below.

[0184] According to the embodiments of the present disclosure, the PAHs included in this study include low molecular weight PAHs, such as A) naphthalene, B) anthracene, and C) phenanthrene. Any other PAH is within the scope of our invention. The degradation pathway of each PAH pollutant is divided into multiple subpathways. PAH degradation involves a naphthalene to salicylic acid subpathway, an anthracene to dihydroxynaphthalene subpathway, a catechol to acetyl-CoA subpathway, a phenanthrene to phthalic acid subpathway, a phthalic acid to dihydroxybenzoic acid subpathway, and a phenanthrene to dihydroxynaphthalene subpathway. Candidate bacterial families involved in the naphthalene to salicylic acid subpathway include one or more of the bacterial families detailed in Table 20A. Any other bacterial family capable of degrading naphthalene to salicylic acid is within the scope of the present disclosure.

[0185] Candidate bacterial families involved in the anthracene to dihydroxynaphthalene subpathway are described in detail in Table 20A. It should be understood that in this context, family refers to a taxonomic classification according to the Linnaean classification system and, in the present disclosure, refers to microbial strains within a given family that possess the genes / proteins / enzymes for the corresponding subpathway. Any other bacterial family capable of degrading anthracene to dihydroxynaphthalene is within the scope of this disclosure.

[0186] Candidate bacterial families involved in the catechol to acetyl-CoA subpathway are detailed in Table 20 A. Any other bacterial family capable of degrading catechol to acetyl-CoA is within the scope of the present disclosure.

[0187] Candidate bacterial families involved in the phenanthrene to phthalate subpathway are described in detail in Table 20 B. Any other bacterial family capable of degrading phenanthrene to phthalate is within the scope of the present disclosure.

[0188] Candidate bacterial families involved in the phthalate to dihydroxybenzoate subpathway are described in detail in Table 20 B. Any other bacterial families capable of degrading phthalate to dihydroxybenzoate are within the scope of the present disclosure.

[0189] Candidate bacterial families involved in the phenanthrene to dihydroxynaphthalene subpathway are described in detail in Table 20B. Any other bacterial family capable of degrading phenanthrene to dihydroxynaphthalene is within the scope of the present disclosure.

[0190] Tables 20A and 20B show detailed candidate bacterial families for PAH degradation.

[0191] Table 20A: List of candidate bacterial families corresponding to various pathways for PAH degradation

[0192]

[0193]

[0194] Table 20B: List of candidate bacterial families corresponding to various pathways for PAH degradation

[0195]

[0196]

[0197] According to embodiments of the present disclosure, PCB degradation involves a subpathway of PCB to biphenyl, a subpathway of biphenyl to acetyl-CoA / pyruvate, a subpathway of biphenyl to 2-hydroxypenta-2,4-dienoate, a subpathway of 2-hydroxypenta-2,4-dienoate to acetyl-CoA / pyruvate, a subpathway of benzoate to acetyl-CoA via catechol, and a subpathway of benzoate to acetyl-CoA via benzoyl-CoA.

[0198] Candidate bacterial families involved in the PCB to biphenyl subpathway are described in detail in Table 21A. It should be understood that, in this context, family refers to a taxonomic classification according to the Linnaean classification system and, in this disclosure, refers to microbial strains within a given family that possess the genes / proteins / enzymes for the corresponding subpathway. Any other bacterial family capable of degrading PCBs to biphenyls is within the scope of this disclosure.

[0199] Candidate bacterial families involved in the biphenyl to acetyl-CoA / pyruvate subpathway are detailed in Table 21 A. Any other bacterial family capable of degrading PCBs to biphenyl and then to acetyl-CoA / pyruvate is within the scope of the present disclosure.

[0200] Candidate bacterial families involved in the biphenyl to 2-hydroxypenta-2,4-dienoate subpathway are described in detail in Table 21 A. Any other bacterial family capable of degrading biphenyl to 2-hydroxypenta-2,4-dienoate is within the scope of the present disclosure.

[0201] Candidate bacterial families involved in the 2-hydroxypenta-2,4-dienoate to acetyl-CoA / pyruvate subpathway are described in detail in Table 21 B. Any other bacterial family capable of degrading 2-hydroxypenta-2,4-dienoate to acetyl-CoA / pyruvate is within the scope of the present disclosure.

[0202] Candidate bacterial families involved in the benzoate via catechol to acetyl-CoA subpathway are detailed in Table 21 B. Any other bacterial family capable of degrading benzoate via catechol to acetyl-CoA is within the scope of the present disclosure.

[0203] Candidate bacterial families involved in the benzoate via benzoyl-CoA to acetyl-CoA subpathway are detailed in Table 21 B. Any other bacterial family capable of degrading benzoate via benzoyl-CoA to acetyl-CoA is within the scope of the present disclosure.

[0204] Table 21A and Table 21B show detailed candidate bacterial families for PCB degradation.

[0205] Table 21A: List of candidate bacterial families corresponding to various pathways for PCB degradation

[0206]

[0207]

[0208] Table 21B: List of candidate bacterial families corresponding to various pathways for PCB degradation

[0209]

[0210]

[0211] In operation, 7A to 7B A flowchart 700 illustrates steps involved in bioremediation of carbon-based contaminants. First, at step 702, a sample is collected from a site containing a plurality of contaminants. At step 704, the plurality of contaminants are separated from the sample. At step 706, one or more types of the plurality of contaminants present in the separated sample are identified, wherein the plurality of contaminants may be, but are not limited to, carbon-based contaminants, such as polycyclic aromatic hydrocarbon (PAH)-based contaminants, polychlorinated biphenyl (PCB)-based contaminants, monocyclic aromatic hydrocarbon (SAH)-based contaminants, or carbon-based nanomaterial (CBNM)-based contaminants.

[0212] In the next step 708, if the identified pollutant is a carbon-based nanomaterial (CBNM), it is degraded using peroxidase, wherein the degradation results in the production of oxidized carbon-based nanomaterials, wherein the oxidized CBNM is one type of carbon-based pollutant, which may lead to the generation of intermediates (including PAH, PCB, SAH, etc.).

[0213] In next step 710, create knowledge base, the information of the pollutant of this knowledge base storage identification, its degradation pathway and from the association of organism and specific pollutant degradation pathway of different environments.This knowledge base also comprises the peroxidase that is used for CBNM degradation and from the organism with this peroxidase of different environments.Further, in step 712, create microbial community, it has the functional ability of the pollutant of complete degradation separation as a whole.In step 714, use the microbial community that is created in the field, to be used for the bioremediation of carbon-based pollutants.In step 716, check that the mixture that is used is to the removal efficacy of eliminating one or more pollutants from the sample that collects in environmental field, and carry out efficacy evaluation by separating and identifying residual pollutant combination from the sample that collects.Finally, in step 718, by adding as partial degradation agent and degrading a group of microorganisms of one or more pollutants that identify in the sample that collects in combination, re-use new mixture in this environmental field.

[0214] According to embodiments of the present disclosure, system 100 can also be illustrated using the example of CBNM degradation. Specific bacterial families have been shown to degrade CBNM. Current analysis suggests that the hypothesis that catalase-peroxidase is the primary enzyme responsible for CBNM degradation is correct. This research and the corresponding analysis are described in detail below.

[0215] Research Overview: As discussed, the initial step in CBNM degradation involves the presence of a bifunctional catalase-peroxidase (katG enzyme) secreted outside the bacterial cell. The redox reaction of CBNM catalyzed by this enzyme produces aromatic intermediates, including but not limited to: various PAH and PCB compounds, such as naphthalene, acenaphthene, and biphenyl; as well as many monocyclic aromatic compounds, such as phthalic acid, salicylic acid, and benzoic acid. These aromatic intermediates are then further degraded by enzymes essential for aromatic hydrocarbon-degrading bacteria.

[0216] Based on E out CBNM degradation ability of the presence of: As discussed in the methods, enzymes capable of degrading CBNM E out Identified as a bifunctional catalase-peroxidase (katG). out (katG) bacterial species E bac The identification of katG suggests that many bacterial genera do possess the enzyme katG. Protein sequences were obtained to analyze the secretion capacity of specific enzymes. The D score (for the presence of an N-terminal signal peptide) and the SP score (for leader-peptide-free secretion) were determined accordingly. For katG enzymes for which no N-terminal signal peptide was detected using SignalP software, the possibility of leader-peptide-free secretion was detected using SecretomeP software, and those with D and SP scores exceeding a threshold score of 0.5 were considered. thre Therefore, we can conclude that many bacterial genera, such as Pseudomonas, Labrys sp., and Stenotrophomonas, do contain the necessary secreted bifunctional catalase-peroxidase to initiate the first step in CBNM degradation.

[0217] Microbial community mixtures: Contaminated sites often include a mixture of pollutants, and carbon-based pollutants and compounds (such as CBNMs, PAHs, PCBs, etc.) are very common at these sites. Effective bioremediation of contaminated samples therefore requires a combination of multiple organisms that are capable of degrading each pollutant type to achieve complete degradation of these pollutants. Bacteria are known to live in multi-species communities and exhibit extensive interactions within and between species and have the extraordinary ability to degrade a large number of organic compounds by consuming them as their primary energy source and further assimilating them without releasing any harmful byproducts. Therefore, in order to achieve complete degradation of graphene oxide, a microbial community mixture that includes CBNM-degrading bacterial genera as well as other microorganisms capable of higher-order aromatic degradation needs to be identified and applied to the contaminated site.

[0218] According to embodiments of the present disclosure, system 100 can also be illustrated using the example of Labrys sp. WJW. Labrys sp. WJW has been shown to degrade CBNM, particularly graphene oxide (GO). Current analysis suggests that the hypothesis that catalase-peroxidase is the primary enzyme responsible for CBNM degradation is correct. This research and the corresponding analysis are described in detail below.

[0219] Research Overview: As discussed, the initial step in CBNM degradation involves the presence of a bifunctional catalase-peroxidase (katG) enzyme secreted outside the bacterial cell by Labrys sp. WJW. The redox reaction of CBNM catalyzed by this enzyme produces aromatic intermediates, including but not limited to various PAH and PCB compounds such as naphthalene, acenaphthene, and biphenyl, as well as numerous monocyclic aromatic compounds such as phthalic acid, salicylic acid, and benzoic acid. These aromatic intermediates are then further degraded by enzymes in Labrys sp. WJW bacteria essential for aromatic hydrocarbon degradation.

[0220] Case Study of GO Degradation by the Novel Bacterial Species Labrys sp. WJW: In this study, a novel strain of the bacterium Labrys sp. WJW, isolated from soil, was shown to utilize GO as a sole carbon source under laboratory conditions. Mass spectrometry analysis of the degradation process revealed that many of the intermediates produced during this process are aromatic hydrocarbons. Furthermore, microarray analysis revealed that many aroma-degrading genes in Labrys sp. WJW were upregulated during this process, indicating that these intermediates are degraded by Labrys sp. WJW.

[0221] Based on E out CBNM degradation ability of the presence of: As discussed in the methods, enzymes capable of degrading CBNM E out It was identified as a bifunctional catalase-peroxidase (katG). out (katG) bacterial species E bac The identification of the protein indicated that Labrys sp. WJW indeed possesses the katG enzyme. The protein sequence was obtained to analyze the secretion capacity of the specific enzyme. The D score (for the presence of an N-terminal signal peptide) and the SP score (for leader-peptide-free secretion) were determined accordingly. While no N-terminal signal peptide was detected using SignalP software, the possibility of leader-peptide-free secretion of the katG enzyme from Labrys sp. WJW was detected using SecretomeP software, with an SP score of 0.80 (on a scale of 0 to 1), which far exceeds the threshold score of 0.5. thre Therefore, we can say that Labrys sp. WJW indeed contains the necessary secreted bifunctional catalase-peroxidase to initiate the first step of CBNM degradation.

[0222] Aromatic Hydrocarbon Degradation Capacity: The aromatic hydrocarbon degradation capacity of Labrys sp. WJW was determined as discussed in the Methods. It was determined that Labrys sp. WJW possesses a complete gene cluster solely for benzoic acid degradation and therefore may not be able to degrade higher aromatic hydrocarbons released as part of the intermediate mixture.

[0223] Microbial community mixtures: Contaminated sites often contain a mixture of pollutants, and carbon-based pollutants and compounds (such as CBNMs, PAHs, PCBs, etc.) are very common at these sites. Effective bioremediation of contaminated samples therefore requires a combination of multiple organisms that are capable of degrading each pollutant type to achieve complete degradation of these pollutants. Bacteria are known to live in multi-species communities and exhibit extensive interactions within and between species and have the extraordinary ability to degrade a large number of organic compounds by consuming them as their primary energy source and further assimilating them without releasing any harmful byproducts. Therefore, in order to achieve complete degradation of graphene oxide, a microbial community mixture that includes Labrys sp. WJW as well as other microorganisms capable of higher-order aromatic degradation needs to be identified and applied to the contaminated site.

[0224] The aromatic hydrocarbon degradation capabilities of various bacteria were analyzed to determine their ability to degrade intermediates formed during CBNM degradation, which show structural similarities to PAH and PCB intermediate compounds. Furthermore, PAHs and PCBs are potent pollutants in their own right and must be degraded into harmless byproducts.

[0225] Existing literature indicates that low-molecular-weight (LMW) PAHs, including hydrocarbons with fewer than four fused benzene rings (e.g., naphthalene, anthracene, and phenanthrene), are biodegradable and typically undergo aerobic degradation. Detailed genomic analysis of bacterial genomes has been conducted to determine their ability to degrade these PAHs. Naphthalene is a PAH that generally favors aerobic degradation by bacteria under oxygen-mediated metabolism, followed by dioxygenase ring cleavage to form TCA cycle intermediates that can be readily assimilated by the organism.

[0226] The naphthalene degradation pathway can be divided into two subpathways: the conversion of naphthalene to salicylic acid (NSP1) and the degradation of salicylic acid via catechol (NSP2), both of which are controlled by the Lys-R regulatory factor. Pathway analysis, including literature mining, manual curation, and comparison with the model organism Pseudomonas stutzeri, helped identify gene clusters involved in the naphthalene-to-salicylic acid and salicylic acid-to-acetyl-CoA subpathways. The Pfam database was searched for domain information corresponding to each gene. In this method, a hidden Markov model-based approach was used to identify the presence of this cluster in bacterial genomes using tools such as HMMER. A window of 20 genes upstream and downstream of the query gene was searched for the presence of the gene cluster. Bacterial genomes (e.g., Polaromonas naphthalenivorans CJ2, Novosphingobium aromaticivorans DSM 12444, Celeribacter indicus, etc.) harbor both this gene cluster and the Lys-R regulatory factor, present in context within their genomes. These organisms are potential naphthalene degraders (PAHs) and require further validation of their active site patterns.

[0227] Using literature mining, we identified a specific pattern of active sites for naphthalene dioxygenases (NDOs) involved in the initial attack on the ring aromatic structure. Multiple sequence alignment (MSA) was used to determine the conservation of the active sites of the enzymes involved. All potential naphthalene degraders (e.g., Polaromonas naphthalenivorans CJ2, Celeribacter indicus, etc.) were further validated by searching for the presence of residues critical for enzymes involved in naphthalene degradation. It was observed that species such as Polaromonas naphthalenivorans CJ2 and Acidovorax sp. P4 retained these important active site residues (e.g., V-209, N-297, F-352). Therefore, bacterial organisms possessing this gene cluster and its regulators and having the specific active site pattern for NDO can be designated as true naphthalene degraders.

[0228] PCBs are treated using a pathway similar to the naphthalene degradation described above. Biphenyl and its less chlorinated forms are produced by the reductive dehalogenation of highly chlorinated PCBs under anaerobic conditions. The rdhABR gene cluster is involved. The domain of the RdhB gene was used to identify the dehalogenation potential of bacterial genomes. Our method was used to identify bacteria such as Dehalococcoides mccartyi, Sulfurospirillum multivorans, and Desulfitobacterium dehalogenans. Clusters for the upstream pathway degradation of biphenyl were identified in organisms such as Acidovorax sp. KKS102, Azoarcus sp. CIB, Celeribacter indicus, and Comamonas testosteroni TK102. The intermediates formed at the end of the upstream pathway are further degraded within the same organism or transported out and degraded by another organism via the downstream pathway. Some of the potential organisms that contain the downstream pathway found in our analysis are listed as Acidovorax JS42, Acidovorax KKS102, etc. Many of these bacteria contain both pathways for complete PCB degradation.

[0229] Therefore, a mixture of microbial communities capable of degrading CBNM as well as other higher and lower aromatic compounds would promote the complete and efficient degradation of carbon-based pollutants.

[0230] Application of the Microbial Mixture: A culture of the aforementioned microbial mixture can be added to a given soil sample contaminated with CBNM. While any other method is within the scope of the present invention, an internal application method is used herein. In this process, the aforementioned microbial mixture is added to the soil along with essential nutrients specified for the soil sample (e.g., nutrient broth containing beef extract). The sample is further aerated and fully hydrated to ensure that the microbial mixture reaches a logarithmic growth phase to promote contaminant bioremediation.

[0231] Efficacy of the applied microbial mixture: The efficacy of the applied microbial mixture is evaluated by isolating and identifying the residual contaminant combination from the collected sample and reapplying a new mixture to the environmental site. The new mixture is prepared by adding a group of microorganisms that can act as partial degraders and degrade in combination one or more contaminants identified in the collected sample.

[0232] The description describes the subject matter herein to enable those skilled in the art to make and use the embodiments. The scope of the subject embodiments is defined by the claims and may include other modifications that occur to those skilled in the art. Such other modifications are intended to be within the scope of the claims if they have similar elements that do not differ from the literal language of the claims or if they include equivalent elements that do not differ substantially from the literal language of the claims.

[0233] The embodiments of the present disclosure herein address the unresolved problem of degradation of pollutants that have a serious impact on the environment. Accordingly, the present embodiments provide a method and system for complete bioremediation of one or more pollutants.

[0234] It should be understood that the scope of protection extends to programs and computer-readable devices having messages therein; when the program is run on a server or mobile device or any suitable programmable device, such computer-readable storage device contains program code means for implementing one or more steps of the method. A hardware device can be any type of programmable device, for example, including any type of computer (such as a server or personal computer, etc.) or any combination thereof. The device can also include a device that can be, for example, a hardware device (such as an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA)) or a combination of hardware and software devices (such as an ASIC and an FPGA), or at least one microprocessor and at least one memory with a software module located therein. Therefore, the device can include hardware devices and software devices. The method embodiments described herein can be implemented in hardware and software. The device can also include software devices. Alternatively, the embodiments can be implemented on different hardware devices, for example, using multiple CPUs.

[0235] The embodiments described herein may include both hardware and software elements. Software-implemented embodiments include, but are not limited to, firmware, resident software, microcode, and the like. The functions performed by the various modules described herein may be implemented by other modules or combinations of other modules. For the purposes of this description, a computer-usable or computer-readable medium may be any device capable of containing, storing, communicating, propagating, or transmitting a program for use by or in connection with an instruction execution system, apparatus, or device.

[0236] The steps shown are used to explain the exemplary embodiments shown, and it should be expected that ongoing technological developments will change the way specific functions are performed. These examples are presented herein for illustrative purposes rather than limiting purposes. In addition, for ease of description, the boundaries of the functional components have been arbitrarily defined here. As long as the specified functions and their relationships are properly performed, alternative boundaries can be defined. Based on the teachings contained herein, alternatives (including equivalents, extensions, variations, deviations, etc. of those described herein) are obvious to those skilled in the art. Such alternatives fall within the scope of the disclosed embodiments. In addition, words such as "including," "having," "having," and "comprising" and other similar forms are equivalent in meaning and are open-ended because the one or more items following any of these words are not intended to exhaustively list such items or be limited to the listed items. It must also be noted that, as used herein and in the appended claims, the singular forms "a" and "the" include plural references unless the context clearly dictates otherwise.

[0237] In addition, embodiments consistent with the present disclosure may be implemented using one or more computer-readable storage media. A computer-readable storage medium refers to any type of physical memory that can store information or data that is readable by a processor. Thus, a computer-readable storage medium can store instructions for execution by one or more processors, including instructions for causing a processor to perform steps or stages consistent with the embodiments described herein. The term "computer-readable medium" should be understood to include tangible objects and exclude carrier waves and transient signals, i.e., non-transient. Examples include random access memory (RAM), read-only memory (ROM), volatile memory, non-volatile memory, hard drives, CD ROMs, DVDs, flash drives, disks, and any other known physical storage media.

[0238] It is intended that the disclosure and examples be considered as exemplary only, with the true scope of the disclosed embodiments being indicated by the claims.

Claims

1. A method (200) for bioremediation of one or more pollutants, the method comprising: collecting a sample from an environmental site containing the one or more pollutants (202); separating and identifying the one or more contaminants present in the sample (204); Create a knowledge base, wherein the knowledge base stores: information on the one or more pollutants identified, Information on complete and partial degradation pathways identified in microorganisms that are capable of completely degrading the one or more pollutants or partially degrading the one or more pollutants, Information about the respective environmental niches in which the microorganisms thrive, and List of microorganisms from different environments with specific complete / partial pollutant degradation pathways, The specific complete degradation pathway refers to a group of genes on the genome of a microorganism and / or proteins encoded by the microorganism, wherein the group of genes and / or encoded proteins are responsible for completely degrading the pollutants into compounds that are safe for the environmental site or compounds that can be assimilated by one or more other microorganisms present in the environment. wherein the partial degradation pathway in the microorganism refers to a set of genes or encoded proteins that constitute one or more subpathways, wherein a subpathway is a subset of the complete degradation pathway encoded within the genome of the microorganism, and the subpathway degrades the pollutant into an intermediate compound that can be released by the microorganism into the environment and subsequently taken up by another microorganism in the environment, wherein the other microorganism has another subpathway that metabolizes the released intermediate compound, wherein the knowledge base includes a pollutant pathway organism matrix PPOM, a genomic pathway enzyme GPE map, a genomic pathway master map GPM, and an enriched environmental microbial database DEBG for each of the one or more pollutants identified in the collected sample, wherein the GPE map includes the names of the microorganisms listed in the DEBG and information about the active sites of each enzyme involved in multiple subpathways on each microbial genome (206), The step of creating the knowledge base further includes: identifying one or more degradation pathways and corresponding genes / proteins in a microorganism using literature mining techniques, wherein the one or more pathways degrade the one or more pollutants isolated and identified, wherein the literature mining is further performed to identify a group of microorganisms in which the one or more degradation pathways are characterized, wherein the literature mining is performed to obtain information regarding the environmental niche in which the microorganism exists and is isolated; Identifying a plurality of subpathways in the degradation pathway that completely or partially degrade the one or more separated pollutants, wherein genes and / or proteins or enzymes corresponding to each of the plurality of subpathways are encoded by the genome of a single microorganism or the genomes of multiple microorganisms, and products formed by each of the plurality of subpathways are released into the environmental site, wherein the products are metabolized by the single microorganism or taken up by one or more other microorganisms present in the environment, wherein the other multiple microorganisms have the ability to metabolize the products; creating a pollutant pathway organism matrix (PPOM) using information regarding the identified degradation pathways for each of the one or more identified pollutants, the plurality of subpathways of the degradation pathways, the set of microorganisms characterizing the degradation pathways, and information based on literature mining and manual curation regarding one or more corresponding environmental niches from which the set of microorganisms were isolated; Using the literature mining technology to create an enriched environmental microbial database DEBG, wherein the DEBG contains information about microorganisms and different environmental niches in which the microorganisms thrive; Creating a pathway domain map (PDM) from a pre-created protein family database (pfamDB), wherein the protein domains contained in the PDM correspond to genes / proteins constituting the plurality of subpathways, and the plurality of subpathways include each degradation pathway present in the PPOM created for the one or more pollutants; Creating a genome map GM, wherein the genome map includes information about all microbial genomes, wherein the information also includes a list of genes in the microorganisms sorted according to their respective genomic positions and the component protein domains encoded by these genes; searching the microbial genomes stored in the DEBG for the presence of protein domains contained in the PDM of each of the plurality of subpathways of all pathways listed in the PPOM to determine the presence of these subpathways on the genome, wherein the search is performed using the genome map GM as a database, and wherein if the domains listed in the PDM in the genome providing the subpathway of the PDM appear within a window size of genes on the genome and the number exceeds a predefined threshold, then the subpathway from the PDM is considered to be present; creating, for each of the one or more pollutants identified in the collected sample, a genome pathway master map (GPM) using the microorganism name corresponding to the microbial genome in the DEBG and information about the presence or absence of the plurality of pathways and the plurality of subpathways on the genome, wherein the GPM map has a value of 0 or 1 based on a first predefined criterion, wherein the GPM provides information about all subpathways of a given pollutant degradation pathway present in each of the microbial genomes listed in the GPM; and Creating a genomic pathway enzyme (GPE) profile for each of the one or more contaminants identified in the collected sample, wherein the GPE profile includes the names of all microorganisms listed in the DEBG corresponding to all microbial genomes reported by the National Center for Biotechnology Information (NCBI), and information about the active site of each enzyme involved in each step of a plurality of subpathways on each genome, wherein the GPE profile has a value of 0 or 1 based on a second predefined criterion; By utilizing the information in the knowledge base, a list of partial pollutant degraders and a list of complete pollutant degraders are identified for each of the one or more pollutants identified in the sample, wherein the partial pollutant degraders are microorganisms that provide one or more subpathways and a corresponding set of genes, encoded proteins, or enzymes for converting pollutants into intermediate compounds, wherein a plurality of partial degraders provide in combination all subpathways for completely degrading the pollutants identified in the collected sample, wherein a complete pollutant degrader has a combination of all subpathways and a corresponding set of genes, encoded proteins, or enzymes within a single microorganism for degrading the pollutants identified in the collected sample (208); creating a microbial profile using information from the knowledge base, wherein the The microbial profile includes information on one or more of the partial pollutant degraders and the complete pollutant degraders capable of degrading each of the one or more pollutants identified in the sample to different degrees of degradation, wherein the different degrees of degradation of the pollutants refer to degradation of the pollutants into different intermediate compounds or metabolites, wherein the intermediate compounds or metabolites are determined by one or more end products released by the degraders under the action of genes or proteins or enzymes corresponding to the subpathways present in the genome of the degraders for degrading the pollutants, wherein the intermediate compounds can be released into the environment and utilized by other microorganisms in the environment or can be assimilated into the same microorganisms that perform such degradation (210); designing a first microbial community using the created microbial profile, the first microbial community comprising microorganisms that collectively provide subpathways required for complete degradation of the one or more contaminants identified in the sample, wherein the microorganisms are capable of surviving together in the same environmental niche from which the sample was collected (212); designing a second microbial community using the created microbial profile, the second microbial community comprising microorganisms that collectively provide genes, proteins, and enzymes for subpathways required for partially degrading the one or more contaminants identified in the collected sample into one or more desired intermediates, wherein the microorganisms forming the second microbial community are capable of surviving together in the environmental niche from which the sample was collected (214); applying a mixture of at least one of the first microbial community and the second microbial community, or both, to an environmental site containing the one or more pollutants (216); examining the efficacy of the applied mixture in eliminating one or more contaminants in a sample collected from the environmental site, wherein the efficacy is evaluated by separating and identifying a combination of residual contaminants from the collected sample (218); and A new mixture is reapplied to the environmental site, wherein the new mixture is prepared by adding a group of microorganisms capable of acting as partial degraders and combinatorially degrading the one or more pollutants in the collected sample (220).

2. The method according to claim 1, wherein The first predefined criteria are for each subpathway in the microbial genome: If the protein domain corresponding to the subpathway recorded in "PDM" does not exist or does not reach the threshold within the predefined gene window, a value of 0 is assigned, and If the subpathway protein domain recorded in "PDM" is present and above the threshold within the predefined gene window, a value of 1 is assigned.

3. The method according to claim 2, wherein: The threshold is determined based on literature mining and manual curation and corresponds to a threshold minimum number of domains and domains that need to be present to confirm the presence of the subpathway in the microbial genome, wherein the threshold is defined as the required fraction of the total number of domains in the PDM corresponding to this subpathway, and The predefined gene window is defined using manual curation, and the gene window describes the distance in terms of the number of genes based on genomic position, wherein the domains in the gene window can be considered to constitute a subpathway, wherein the genes encoding protein domains forming the subpathway appear together on the microbial genome and are therefore located within the gene window size defined on the genome.

4. The method according to claim 1, wherein The second predefined criteria for each enzyme corresponding to each step of the subpathway identified in the microbial genome were: a value of 1 is assigned to the plurality of enzymes for which the active site pattern of the enzyme is found, the active site pattern of the constituent enzymes of each candidate pathway being a pattern specific to the active site of the enzyme obtained by literature mining, wherein the active site pattern is identified by a motif that serves as a signature sequence that helps identify whether the enzyme is functionally capable of binding to a substrate and is conserved among all enzymes with similar functions, and wherein the motif is identified by performing a multiple sequence alignment (MSA) between all possible homologs of functionally similar enzymes to identify conserved amino acid patterns of enzymes validated in the literature to assess functional importance, and In case the active site pattern of the enzyme was not found, a value of 0 was assigned.

5. The method according to claim 1, further comprising: Each enzyme was tested for the presence of secretory ability, and in the GPE profile, a value of 0 was updated for the absence of secretory ability or a value of 1 was updated for the presence of secretory ability.

6. The method according to claim 1, wherein The desired intermediates refer to a group of intermediates obtained during the partial degradation of the one or more identified pollutants.

7. The method according to claim 1, wherein The knowledge base stores information about one or more pollutants, wherein some of these pollutants include different homologues of polyethylene terephthalate (PET), styrene, polyurethane, polycyclic aromatic hydrocarbons (PAH), polychlorinated biphenyls (PCB), or carbon-based nanomaterials (CBNM).

8. The method according to claim 7, wherein: Complete CBNM degradation involves the presence of a bifunctional catalase-peroxidase Kat enzyme and a subpathway for the degradation of one or more intermediates formed following the catalytic action of the Kat enzyme, wherein the intermediates include one or more of PAH, PCB, and monocyclic aromatic hydrocarbon (SAH) degradation products formed as a result of the catalase-peroxidase catalytic action on CBNM.

9. The method according to claim 7, wherein: The complete PET degradation involves a subpathway of PETase, followed by the conversion of terephthalate TPA to protocatechuate PCA, and candidate bacterial families involved in these subpathways for PET degradation include one or more of the following: Polycystaceae, Burkholderiaceae, subclass and order of undetermined Burkholderiales, Alteromonadaceae, Oceanospirillaceae, Pseudomonadaceae, or Vibrioaceae, Comamonadaceae, Bacillaceae, Bradyrhizobiaceae, Burkholderiaceae, Ichthyorhizobactaceae, Sphingomonasaceae, Hyphomicrobacteriaceae, Pseudomonadaceae, Pseudonocardaceae, Oxalobacteraceae, Rhizobiaceae, Nocardiaceae, Rhodocyclaceae, or Streptomycetaceae for the terephthalate to protocatechuate subpathway, and For the protocatechuate to acetyl-CoA subpathway, Actinomycetes, Caulobacteraceae, Oxalobacteraceae, Streptomycetaceae, Micrococcaceae, Rhizobacteraceae, Myxococcaceae, Nocardiaceae, Brucellaceae, Nocardiaceae, Oceanospirillaceae, Kininococcaceae, Pseudonocardiaceae, Actinopolyspora, Neurospora, Xanthomonadaceae, Mycomicrobiaceae, Microbacteriaceae, Alcaligenesaceae, Geodermatophilaceae, Burkholderiaceae, Enterobacteriaceae, Halomonasaceae, Moraxellaceae, Dietziaceae, Leafbacteriaceae, Sphingomonasaceae, Rhodospirillaceae, Micromonosporaceae, Comamonadaceae, Pseudomonadaceae, Aeromonadaceae, Alteromonadaceae, Auriculariaceae, Phage Cellulose, Neisseriaceae, Deinococciaceae, Nocardiaceae, Vibrio, Kiloniellaceae, Gordoniaceae, Listeriaceae, Bacillaceae, Xanthobacteriaceae, Rhodobacteriaceae, Bungellaceae, Bradyrhizobacteriaceae, Saprospiraceae, Sphingobacteriaceae, Thermus, Clostridiaceae, Flavobacteriaceae, Brevibacteriaceae, Corynebacterium, Beyerinkiaceae, Methylobacteraceae, Sporobacteriaceae, Granulosae, Saccharomyces, Bacillaceae 1, Tenuispora, Globobacteriaceae, unclassified Betaproteobacteria, unclassified Burkholderiales, unclassified Flavobacteriales, Yersiniaceae, or Vicinamibacteraceae.

10. The method according to claim 7, wherein: The complete degradation of PAH also includes: The degradation of naphthalene includes two sub-pathways, namely, the degradation of naphthalene to salicylic acid, followed by the degradation of salicylic acid to acetyl-CoA via catechol. Anthracene degradation, which is divided into a subpathway that converts anthracene to dihydroxynaphthalene, followed by a subpathway that degrades salicylic acid to acetyl-CoA via the catechol metabolic pathway, and Degradation of phenanthrene, which involves a subpathway from phenanthrene to phthalic acid, a subpathway from phthalic acid to dihydroxybenzoic acid, and a subpathway from phenanthrene to dihydroxynaphthalene, and candidate bacterial families involved in the subpathway for degradation of each type of PAH include one or more of the following: Comamonadaceae, Rhizobiaceae, Alteromonadaceae, Bacillaceae, Alcaligenesaceae, Bradyrhizobiaceae, Burkholderiaceae, Rhodobacteraceae, Erythrobacteraceae, Ichthyorhizobacteriaceae, Gordoniaceae, Oceanospirillaceae, Aurantiaceae, Oxalobacteraceae, Leafbacteraceae, Mycobacteraceae, Sphingomonasaceae, Hyphomicrobacteriaceae, Neisseriaceae, Pseudomonadaceae, unclassified Rhizobiales, Nocardiaceae, Rhodocyclaceae, or Streptomycetaceae for the naphthalene to salicylic acid subpathway, Comamonadaceae, Alcaligenesaceae, Actinomorphaceae, Bacillaceae, Geodermatophilaceae, Sphingomonasaceae, Bradyrhizobiaceae, Burkholderiaceae, Caulobacteraceae, Rhodobacteraceae, Frankiaceae, Alteromonadaceae, Leafbacteraceae, Mycobacteraceae, Hyphomicrobacteraceae, Erythrobacteraceae, Pseudomonadaceae, Rhizobiaceae, Nocardiaceae, Streptomycetaceae, or an undetermined subclass of Gammaproteobacteria for the anthracene to dihydroxynaphthalene subpathway, Comamonadaceae, Moraxellaceae, Alcaligenesaceae, Alkanedinomycetaceae, Alicyclobacillusaceae, Exothiorhodospirillaceae, Actinomorphaceae, Leafbacteraceae, Neisseriaceae, Rhodocyclaceae, Pseudomonadaceae, Bacillaceae, Thiotrichozoaceae, Bradyrhizobactaceae, Burkholderiaceae, Clostridiaceae, Rhodobacteraceae, Corynebacteraceae, Oxalobacteraceae, Rickettsiaceae, Frankiaceae, Gordoniaceae, Intersporaceae, Enterobacteriaceae, Kinococcus, Alteromonadaceae, Auriculariaceae, Mycobacteriaceae, Nakamurellaceae, Nocardiaceae, Nocardiaceae, Sphingomonasaceae, Micrococcaceae, Rhizobiaceae, Streptomycetaceae, Sulfobacteraceae, or Thermomonasaceae for the catechol to acetyl-CoA subpathway, Comamonadaceae, Alteromonadaceae, Leafbacteraceae, Bacillaceae, Bradyrhizobiaceae, Burkholderiaceae, Erythrobacteraceae, Oxalobacteraceae, Mycobacteraceae, Sphingomonasaceae, Hyphomicrobacteriaceae, Micrococcaceae, Pseudomonadaceae, Rhizobiaceae, Nocardiaceae, or Streptomycetaceae for the phenanthrene to phthalate subpathway, Acetobacteraceae, Comamonadaceae, Alcaligenesaceae, Bacillaceae, Bradyrhizobactaceae, Brucellaceae, Burkholderiaceae, Halomonasaceae, Coulorhabditisaceae, Corynebacteraceae, Frankiaceae, Gordoniaceae, Leafbacteriaceae, Rhodobacteraceae, Mycobacteriaceae, Nocardiaceae, Sphingomonasaceae, Rhodobacteraceae, Oxalobacteraceae, Pseudonocardaceae, Pseudoalteromonasaceae, Pseudomonadaceae, Rhizobacteraceae, Nocardiaceae, Alteromonadaceae, Streptomycetaceae, Eucystisaceae, Rhodospirillaceae, Enterobacteriaceae, or an undetermined Gammaproteobacteriaceae for the phthalate to dihydroxybenzoate subpathway, and Comamonadaceae, Bacillaceae, Bradyrhizobiaceae, Burkholderiaceae, Caulobacteraceae, Oxalobacteraceae, Rhodocyclaceae, Frankiaceae, Halomonasaceae, Immundisolibacteraceae, Rhodobacteraceae, Alteromonadaceae, Oceanospirillaceae, Leafbacteraceae, Mycobacteraceae, Nocardiaceae, Sphingomonasaceae, Hyphomicrobacteriaceae, Enterobacteriaceae, Erythrobacteraceae, Pseudonocardiaceae, Micrococcaceae, Pseudomonadaceae, Rhizobiaceae, or Streptomycetaceae for the subpathway from phenanthrene to dihydroxynaphthalene.

11. The method according to claim 7, wherein: The complete degradation of PCBs involves: a subpathway of reductive dehalogenation of highly chlorinated PCBs to biphenyl; followed by a subpathway of converting biphenyl to 2-hydroxypenta-2,4-dienoate, which is further degraded via a downstream pathway to form pyruvate and acetyl-CoA; and a subpathway of the intermediate formed, which is converted to acetyl-CoA via a benzoyl-CoA / catechol pathway, and candidate bacterial families involved in these subpathways for PCB degradation include one or more of the following: Dehalogenobacteriaceae, Peptococcus or Campylobacteraceae for the PCB to biphenyl subpathway, Comamonadaceae, Alcaligenesaceae, Alkanodesaceae, Rhodocyclaceae, Bacillaceae, Burkholderiaceae, Connesiaceae, Corynebacteriumaceae, Erythrobacteraceae, Frankiaceae, Aurantiaceae, Beyerlinckiaceae, Mycobacteriumaceae, Sphingomonasaceae, Paenobacteraceae, Hyphomicrobacteriaceae, Pseudoalteromonasaceae, Pseudomonadaceae, Pseudonocardiaceae, Xanthomonadaceae, Rhizobacteriaceae, Nocardiaceae, Alteromonadaceae, Kininococcus, or Tropical Spongobacteraceae for the biphenyl to acetyl-CoA / pyruvate subpathway Comamonadaceae, Alcaligenesaceae, Rhodocyclaceae, Bacillaceae, Bradyrhizobiaceae, Burkholderiaceae, Rhodobacteraceae, Corynebacteraceae, Immundisolibacteraceae, Beyerlinckiaceae, Mycobacteraceae, Nocardiaceae, Sphingomonasaceae, Pseudomonadaceae, Rhizobiaceae, Nocardiaceae, or Streptomycetaceae for the biphenyl to 2-hydroxypenta-2,4-dienoate subpathway, Comamonadaceae, Rhizobiaceae, Alcaligenesaceae, Alicyclobacillusaceae, Neisseriaceae, Rhodocyclaceae, Bacillaceae, Paenibacillusaceae, Burkholderiaceae, Oxalobacteraceae, Gordoniaceae, Rhodospirillaceae, Aurantiaceae, Mycobacteraceae, Sphingomonasaceae, Rhodobacteraceae, Kininococcus, Pseudomonadaceae, Pseudonocardiaceae, Nocardiaceae, Streptomycetaceae, Neurocystaceae, or an undetermined subclass of Gammaproteobacteria for the 2-hydroxypenta-2,4-dienoate to acetyl-CoA / pyruvate subpathway, Moraxellaceae, Rhizobacteriaceae, Alteromonadaceae, Actinomorphaceae, Micrococcaceae, Burkholderiaceae, Oxalobacteraceae, Geodermatophilaceae, Gordoniaceae, Halomonasaceae, Xanthomonadaceae, Methylobacteriaceae, Mycobacteriaceae, Aeromonadaceae, Rhodobacteraceae, Comamonadaceae, Neisseriaceae, Pseudomonadaceae, Pseudonocardiaceae, Sphingomonasaceae, or Vibrioaceae for the benzoate to acetyl-CoA subpathway via catechol, and Comamonadaceae, Moraxellaceae, Alcaligenesaceae, Rhodocyclaceae, Burkholderiaceae, Polycystaceae, Oxalobacteraceae, Labilitrichaceae, subclass, order, family, and undetermined Oceanospirillales for the benzoate to acetyl-CoA subpathway via benzoyl-CoA.

12. The method according to claim 1, wherein A profile of microorganisms that can survive in the environmental site from which the sample was obtained and that can degrade the pollutant to different degrees is created using information from the knowledge base, wherein a first matrix is ​​created from microorganisms with a GPE matrix value of 1 and a second matrix is ​​created using microorganisms from pollutants (P i ), and a third matrix is ​​created by the set of candidate organism results having a value of 1 for the subpathways in the GPM and the corresponding enzymes in the GPE, wherein information about the environmental niches in which the microorganisms in the third matrix reproduce can be obtained from the DEBG, wherein the information is used to create a pollutant organism environment matrix POEM, the POEM including: each of the one or more pollutants identified in the sample, organisms capable of degrading each pollutant to varying degrees depending on the presence of a complete pathway or subpathway, and the environments in which the organisms are isolated and reproduced.

13. A system (100) for bioremediation of one or more contaminants, the system (100) comprising: a sample collection module (102) for collecting a sample from an environmental site containing the one or more pollutants; a contaminant separation and identification module (104) for separating and identifying the one or more contaminants present in the sample; Processor (108); A memory (106) in communication with the processor, wherein the processor is configured to perform the following steps: Create a knowledge base, wherein the knowledge base stores: information on the one or more pollutants identified, Information on complete and partial degradation pathways identified in microorganisms that are capable of completely degrading the one or more pollutants or partially degrading the one or more pollutants, Information about the respective environmental niches in which the microorganisms thrive, and A list of microorganisms from different environments with specific complete / partial pollutant degradation pathways, wherein the specific complete degradation pathway refers to a set of genes on the microorganism genome and / or proteins encoded by the microorganism, wherein the set of genes and / or encoded proteins are responsible for the complete degradation of pollutants into compounds that are safe for the environmental site or compounds that can be assimilated by one or more other microorganisms present in the environment, wherein the partial degradation pathway in the microorganism refers to a group of genes or encoded proteins that constitute one or more subpathways, wherein a subpathway is a subset of the complete degradation pathway encoded in the genome of the microorganism, and the subpathway degrades the pollutant into an intermediate compound, which can be released by the microorganism into the environment and subsequently taken up by another microorganism in the environment, wherein the other microorganism has another subpathway that metabolizes the released intermediate compound, wherein the knowledge base includes a pollutant pathway organism matrix PPOM, a genomic pathway enzyme GPE map, a genomic pathway master map GPM and an enriched environmental microbial database DEBG for each of the one or more pollutants identified in the collected sample, wherein the GPE map includes the names of the microorganisms listed in the DEBG and information about the active sites of each enzyme involved in multiple subpathways on each microbial genome; The step of creating the knowledge base further includes: identifying one or more degradation pathways and corresponding genes / proteins in a microorganism using literature mining techniques, wherein the one or more pathways degrade the one or more pollutants isolated and identified, wherein the literature mining is further performed to identify a group of microorganisms in which the one or more degradation pathways are characterized, wherein the literature mining is performed to obtain information regarding the environmental niche in which the microorganism exists and is isolated; Identifying a plurality of subpathways in the degradation pathway that completely or partially degrade the one or more separated pollutants, wherein genes and / or proteins or enzymes corresponding to each of the plurality of subpathways are encoded by the genome of a single microorganism or the genomes of multiple microorganisms, and products formed by each of the plurality of subpathways are released into the environmental site, wherein the products are metabolized by the single microorganism or taken up by one or more other microorganisms present in the environment, wherein the other multiple microorganisms have the ability to metabolize the products; creating a pollutant pathway organism matrix (PPOM) using information regarding the identified degradation pathways for each of the one or more identified pollutants, the plurality of subpathways of the degradation pathways, the set of microorganisms characterizing the degradation pathways, and information based on literature mining and manual curation regarding one or more corresponding environmental niches from which the set of microorganisms were isolated; Using the literature mining technology to create an enriched environmental microbial database DEBG, wherein the DEBG contains information about microorganisms and different environmental niches in which the microorganisms thrive; Creating a pathway domain map (PDM) from a pre-created protein family database (pfamDB), wherein the protein domains contained in the PDM correspond to genes / proteins constituting the plurality of subpathways, and the plurality of subpathways include each degradation pathway present in the PPOM created for the one or more pollutants; Creating a genome map GM, wherein the genome map includes information about all microbial genomes, wherein the information also includes a list of genes in the microorganisms sorted according to their respective genomic positions and the component protein domains encoded by these genes; searching the microbial genomes stored in the DEBG for the presence of protein domains contained in the PDM of each of the plurality of subpathways of all pathways listed in the PPOM to determine the presence of these subpathways on the genome, wherein the search is performed using the genome map GM as a database, and wherein if the domains listed in the PDM in the genome providing the subpathway of the PDM appear within a window size of genes on the genome and the number exceeds a predefined threshold, then the subpathway from the PDM is considered to be present; creating, for each of the one or more pollutants identified in the collected sample, a genome pathway master map (GPM) using the microorganism name corresponding to the microbial genome in the DEBG and information about the presence or absence of the plurality of pathways and the plurality of subpathways on the genome, wherein the GPM map has a value of 0 or 1 based on a first predefined criterion, wherein the GPM provides information about all subpathways of a given pollutant degradation pathway present in each of the microbial genomes listed in the GPM; and Creating a genomic pathway enzyme (GPE) profile for each of the one or more contaminants identified in the collected sample, wherein the GPE profile includes the names of all microorganisms listed in the DEBG corresponding to all microbial genomes reported by the National Center for Biotechnology Information (NCBI), and information about the active site of each enzyme involved in each step of a plurality of subpathways on each genome, wherein the GPE profile has a value of 0 or 1 based on a second predefined criterion; identifying, by utilizing the information in the knowledge base, a list of partial pollutant degraders and a list of complete pollutant degraders for each of the one or more pollutants identified in the sample, wherein the partial pollutant degraders are microorganisms that provide one or more subpathways and a corresponding set of genes, encoded proteins, or enzymes for converting a pollutant into an intermediate compound, wherein a plurality of partial degraders in combination provide all subpathways for completely degrading the pollutants identified in the collected sample, and wherein a complete pollutant degrader has a combination of all subpathways and a corresponding set of genes, encoded proteins, or enzymes within a single microorganism for degrading the pollutant identified in the collected sample; creating a microbial profile using the information from the knowledge base, wherein the microbial profile includes information on one or more of the partial pollutant degraders and the complete pollutant degraders that are capable of degrading each of the one or more pollutants identified in the sample to different degrees of degradation, wherein the different degrees of degradation of the pollutants refer to degradation of the pollutants into different intermediate compounds or metabolites, wherein the intermediate compounds or metabolites are determined by one or more end products released by the degraders under the action of genes, proteins, or enzymes corresponding to the subpathways present in the genome of the degraders that degrade the pollutants, and wherein the intermediate compounds can be released into the environment and utilized by other microorganisms in the environment or can be assimilated into the same microorganisms that perform such degradation; using the created microbial profile to design a first microbial community, the first microbial community comprising microorganisms that collectively provide subpathways required for complete degradation of the one or more contaminants identified in the sample, and wherein the microorganisms are capable of surviving together in the same environmental niche from which the sample was collected; designing a second microbial community using the created microbial profile, the second microbial community comprising microorganisms that collectively provide genes, proteins, and enzymes for subpathways required for partially degrading the one or more contaminants identified in the collected sample into one or more desired intermediates, wherein the microorganisms forming the second microbial community are capable of surviving together in the environmental niche from which the sample was collected; an application module (114) for applying a mixture of at least one of the first microbial community and the second microbial community, or both, to an environmental site containing the one or more pollutants; an efficacy module (116) for examining the efficacy of the applied mixture in eliminating one or more pollutants in a sample collected from the environmental site, wherein the efficacy assessment is performed by separating and identifying a combination of residual pollutants from the collected sample; and The application module (114) is used to re-apply a new mixture to the environmental site, wherein the new mixture is prepared by adding a group of microorganisms that can act as partial degradants and degrade the one or more pollutants in the collected sample in combination.

14. One or more non-transitory machine-readable information storage media containing one or more instructions that, when executed by one or more hardware processors, result in: collecting samples from environmental sites containing the one or more pollutants; separating and identifying the one or more contaminants present in the sample; Create a knowledge base where The knowledge base stores: information on the one or more pollutants identified, Information on complete and partial degradation pathways identified in microorganisms that are capable of completely degrading the one or more pollutants or partially degrading the one or more pollutants, Information about the respective environmental niches in which the microorganisms thrive, and List of microorganisms from different environments with specific complete / partial pollutant degradation pathways, The specific complete degradation pathway refers to a set of genes on the genome of a microorganism and / or proteins encoded by the microorganism, wherein the set of genes and / or the encoded proteins are responsible for completely degrading the pollutants into compounds that are safe for the environmental site or compounds that can be assimilated by one or more other microorganisms in the environment. wherein the partial degradation pathway in the microorganism refers to a group of genes or encoded proteins that constitute one or more subpathways, wherein a subpathway is a subset of the complete degradation pathway encoded in the genome of the microorganism, and the subpathway degrades the pollutant into an intermediate compound, which can be released by the microorganism into the environment and subsequently absorbed by another microorganism in the environment, wherein the other microorganism has another subpathway that metabolizes the released intermediate compound, wherein the knowledge base includes a pollutant pathway organism matrix PPOM, a genomic pathway enzyme GPE map, a genomic pathway master map GPM and an enriched environmental microbial database DEBG for each of the one or more pollutants identified in the collected samples, wherein the GPE map includes the names of the microorganisms listed in the DEBG and information about the active sites of each enzyme involved in multiple subpathways on each microbial genome, The step of creating the knowledge base further includes: identifying one or more degradation pathways and corresponding genes / proteins in a microorganism using literature mining techniques, wherein the one or more pathways degrade the one or more pollutants isolated and identified, wherein the literature mining is further performed to identify a group of microorganisms in which the one or more degradation pathways are characterized, wherein the literature mining is performed to obtain information regarding the environmental niche in which the microorganism exists and is isolated; Identifying a plurality of subpathways in the degradation pathway that completely or partially degrade the one or more separated pollutants, wherein genes and / or proteins or enzymes corresponding to each of the plurality of subpathways are encoded by the genome of a single microorganism or the genomes of multiple microorganisms, and products formed by each of the plurality of subpathways are released into the environmental site, wherein the products are metabolized by the single microorganism or taken up by one or more other microorganisms present in the environment, wherein the other multiple microorganisms have the ability to metabolize the products; creating a pollutant pathway organism matrix (PPOM) using information regarding the identified degradation pathways for each of the one or more identified pollutants, the plurality of subpathways of the degradation pathways, the set of microorganisms characterizing the degradation pathways, and information based on literature mining and manual curation regarding one or more corresponding environmental niches from which the set of microorganisms were isolated; Using the literature mining technology to create an enriched environmental microbial database DEBG, wherein the DEBG contains information about microorganisms and different environmental niches in which the microorganisms thrive; Creating a pathway domain map (PDM) from a pre-created protein family database (pfamDB), wherein the protein domains contained in the PDM correspond to genes / proteins constituting the plurality of subpathways, and the plurality of subpathways include each degradation pathway present in the PPOM created for the one or more pollutants; Creating a genome map GM, wherein the genome map includes information about all microbial genomes, wherein the information also includes a list of genes in the microorganisms sorted according to their respective genomic positions and the component protein domains encoded by these genes; searching the microbial genomes stored in the DEBG for the presence of protein domains contained in the PDM of each of the plurality of subpathways of all pathways listed in the PPOM to determine the presence of these subpathways on the genome, wherein the search is performed using the genome map GM as a database, and wherein if the domains listed in the PDM in the genome providing the subpathway of the PDM appear within a window size of genes on the genome and the number exceeds a predefined threshold, then the subpathway from the PDM is considered to be present; creating, for each of the one or more pollutants identified in the collected sample, a genome pathway master map (GPM) using the microorganism name corresponding to the microbial genome in the DEBG and information about the presence or absence of the plurality of pathways and the plurality of subpathways on the genome, wherein the GPM map has a value of 0 or 1 based on a first predefined criterion, wherein the GPM provides information about all subpathways of a given pollutant degradation pathway present in each of the microbial genomes listed in the GPM; and Creating a genomic pathway enzyme (GPE) profile for each of the one or more contaminants identified in the collected sample, wherein the GPE profile includes the names of all microorganisms listed in the DEBG corresponding to all microbial genomes reported by the National Center for Biotechnology Information (NCBI), and information about the active site of each enzyme involved in each step of a plurality of subpathways on each genome, wherein the GPE profile has a value of 0 or 1 based on a second predefined criterion; identifying, by utilizing the information in the knowledge base, a list of partial pollutant degraders and a list of complete pollutant degraders for each of the one or more pollutants identified in the sample, wherein the partial pollutant degraders are microorganisms that provide one or more subpathways and a corresponding set of genes, encoded proteins, or enzymes for converting a pollutant into an intermediate compound, wherein a plurality of partial degraders in combination provide all subpathways for completely degrading the pollutants identified in the collected sample, and wherein a complete pollutant degrader has a combination of all subpathways and a corresponding set of genes, encoded proteins, or enzymes within a single microorganism for degrading the pollutant identified in the collected sample; creating a microbial profile using information from the knowledge base, wherein the microbial profile includes information on one or more of the partial pollutant degraders and the complete pollutant degraders that are capable of degrading each of the one or more pollutants identified in the sample to different degrees of degradation, wherein the different degrees of degradation of the pollutants refer to degradation of the pollutants into different intermediate compounds or metabolites, wherein the intermediate compounds or metabolites are determined by one or more end products released by the degraders under the action of genes, proteins, or enzymes corresponding to the subpathways present in the genome of the degraders that degrade the pollutants, and wherein the intermediate compounds can be released into the environment and utilized by other microorganisms in the environment or can be assimilated into the same microorganisms that perform such degradation; designing a first microbial community using the created microbial profile, the first microbial community comprising microorganisms that collectively provide subpathways required for complete degradation of the one or more contaminants identified in the sample, wherein the microorganisms are capable of surviving together in the same environmental niche from which the sample was collected; designing a second microbial community using the created microbial profile, the second microbial community comprising microorganisms that collectively provide genes, proteins, and enzymes for subpathways required for partially degrading the one or more contaminants identified in the collected sample into one or more desired intermediates, wherein the microorganisms forming the second microbial community are capable of surviving together in the environmental niche from which the sample was collected; applying at least one of the first microbial community and the second microbial community, or a mixture of both, to an environmental site containing the one or more pollutants using an application module (114); examining the efficacy of the applied mixture in eliminating one or more pollutants in a sample collected from the environmental site using an efficacy module (116), wherein the efficacy assessment is performed by separating and identifying a combination of residual pollutants from the collected sample; and A new mixture is reapplied to the environmental site by using an application module (114), wherein the new mixture is prepared by adding a group of microorganisms capable of acting as partial degraders and combinatorially degrading the one or more pollutants in the collected sample.