Deep neural networks for biodegradability
A computer-based method using graph representation models predicts biodegradability properties of substances, addressing the challenge of non-degradable waste by enabling the design of biodegradable substances with improved decomposition capabilities.
Patent Information
- Application Number
- PCT/EP2024/088243
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-12-20
- Publication Date
- 2025-06-26
AI Technical Summary
The challenge is to develop a reliable and standardized method for generating biodegradability properties of substances, particularly functional chemical compounds, to address the issue of non-degradable waste and environmental pollution.
A computer-implemented method using data-driven graph representation models and property models to generate biodegradation properties. This involves creating substance-specific graph representations from numeric graph representations associated with atoms and bonds, and then using these representations to predict biodegradability properties.
The method enables more accurate and standardized prediction of biodegradability properties, facilitating the design and production of biodegradable substances that can effectively decompose and reduce environmental pollution.
Smart Images

Figure EP2024088243_26062025_PF_FP_ABST
Abstract
Description
[0001] Deep neural networks for biodegradability
[0002] TECHNICAL FIELD
[0003] The disclosure relates to the field of biodegradable substances or molecules, such as small molecules or large molecules or formulations using machine learning. Disclosed are methods, apparatuses, computer elements, substances or uses for generating biodegradation properties of one or more substances.
[0004] TECHNICAL BACKGROUND
[0005] Generally, functional chemical compounds referring, for example, to small molecules, are widely used in industrial and / or daily use products due to their broad range of application properties. The use of functional chemical compounds encompasses amongst others coatings, personal care products, washing detergents, lubricants, packages and foams. However, this widely spread application leads on the other hand to a huge amount of waste containing the used functional chemical compounds. Non-degradable waste is a problem when disposing in a non-designated environment. Especially, build-ups of chemicals in the environment, like a phosphate build up leading to an algae bloom, are undesired. Thus, there is not only a need for functional chemical compounds that decompose.
[0006] SUMMARY OF THE INVENTION
[0007] In one aspect disclosed is a method, in particular a computer-implemented method, for generating at least one biodegradability property characterizing at least one substance, the method comprising the steps of: - providing at least one substance specific graph representation generated by providing numeric graph representation(s) to at least one data-driven graph representation model configured or trained to map numeric graph representation(s) associated with at least with atoms and bonds of the at least one substance to at least one substance specific graph representation associated with at least atoms, bonds, and correlations related to atoms and / or bonds;
[0008] - generating at least one biodegradation property by providing the at least one substance specific graph representation to the at least one data-driven property model configured ortrained to map the at least one substance specific graph representation to at least one biodegradation property;
[0009] - providing the at least one generated biodegradation property for characterizing the biodegradability of the at least one substance.
[0010] In another aspect disclosed is an apparatus for generating at least one biodegradability property characterizing at least one substance, the method comprising the steps of:
[0011] - an input interface configured to provide at least one substance specific graph representation generated by providing numeric graph representation (s) to at least one data-driven graph representation model configured or trained to map numeric graph representation(s) associated with at least with atoms and bonds of the at least one substance to at least one substance specific graph representation associated with at least atoms, bonds, and correlations related to atoms and / or bonds;
[0012] - a property generator configured to generate at least one biodegradation property by providing the at least one substance specific graph representation to the at least one data-driven property model configured or trained to map the at least one substance specific graph representation to at least one biodegradation property;
[0013] - an output interface configured to provide the at least one generated biodegradation property for characterizing the biodegradability of the at least one substance.
[0014] In another aspect disclosed is a use of the biodegradation property and associated substance structure generated and selected according to the methods disclosed herein or by the apparatus disclosed herein to produce a biodegradable substance with the selected substance structure. In another aspect disclosed is a computer element, such as a computer readable storage medium, a computer program or a computer program product, comprising instructions, which when executed by a computing node or a computing system, direct the computing node or computing system to perform the methods disclosed herein.
[0015] Any disclosure and embodiments described herein relate to the methods, the apparatuses, the surfactants, the uses and the computer elements. Advantageously, the benefits provided by any of the embodiments and examples equally apply to all other embodiments and examples.
[0016] EMBODIMENTS
[0017] In the following, embodiments of the present disclosure will be outlined by ways of embodiments and / or example. It is to be understood that the present disclosure is not limited to said embodiments and / or examples.
[0018] The methods, the methods, apparatuses, computer elements, biodegradable substances or uses disclosed herein allow for more reliable and standardized generation of biodegradation property of biodegradable substances.
[0019] The substance may include or be a functional chemical compound, such as a small molecule, a large molecule, such as a polymer or an oligomer, a formulation, or combinations thereof.
[0020] The functional chemical compound may include or be any functional chemical compound. A functional chemical compound generally is a chemical compound with application properties for a technical purpose, i.e. fulfilling a function in a chemical product. For example, a chemical compound providing an UV-protection in a sunscreen substance, is a functional chemical compound. Functional chemical compounds can comprise an active ingredient, i.e. an ingredient that provides the functionality of the functional chemical compound. Moreover, an active ingredient can provide a biological activity of the functional chemical compound. For example, an active ingredient can refer to an antifungal, aroma chemical, UV absorber, food additive, vitamin, nutrient, dye, surfactant. Functional chemical compounds are generally characterized by their chemical structure. However, also different chemical structures can be present in one functional chemical compound. For example, a functional chemical compound may be composed of a component being a molecule undergoing tautomerism, protonation or deprotonation, or the like. A functional chemical compound may be composed by more than one stereoisomer. Hence, a functional chemical compound composing of one type of molecule may be associated with one or more chemical structures, e.g. one protonated structure and one uncharged structure or two stereoisomers. Thus, generally, functional chemical compounds may include all molecules related to one chemical formula by means of (de)protonation, isomerization such as tautomerism and stereo isomerization. Moreover, a functional chemical compound may refer to an arbitrary functional chemical compound describable with one or more chemical structures. In an embodiment, chemical structures may be associated with one chemical formula.
[0021] Preferably, the functional chemical compound consists of a small molecule. Preferably, the functional chemical compound has a molecular mass below 10000 g / mol. More preferably the chemical compound has a molecular weight of less than 800 g / mol, even more preferably of less than 400 g / mol. Further, it is preferred that the functional chemical compound is present in the environment in a form that allows to completely describe the molecules of the functional chemical compound using simple structural formulas, that contain the relevant information. A simple molecular structure refers to molecules that can be unambiguously described by covalent bindings between the atoms of the molecule. Examples, where this is not the case, are e.g. systems with dynamic equilibria between several forms like monomer and oligomers as in the case of several inorganic acids, or ionic species with very localized charge that strongly interacts with a solvent, e.g. via hydrogen bonding. Preferably, the functional chemical compound has at least one of the following properties: Having an effect on a living organism, being suitable for influencing the structure or, being suitable for influencing the functioning of a living organism. In an embodiment, the functional chemical compound comprises at least one of the following functional groups: Ether group, hydroxyl group, peroxo group, hydroperoxide groups, carboxyl group, carboxyl group derivative, carbonyl group, amine group, imine group, hydrazine group, urea group, urethane group, thiourethane group, nitril group, azide group, azo group, cyanate group, isocyanate group, isocyanide group, pyridine group, alkane group, alkene group, alkyne group, phenyl group, ketone group, thioketone group, aldehyde group, thioaldehyde group, acetal group, ketal group, oxime group, hydrazone group, nitro group, nitroso group, thiol group, sulfide group, disulfide group, sulfonic acid derivative, sulfin ic acid derivative, sulfate group, sulfate group derivative, sulfone group, sulfoxide group, sulfhydryl group, sulfide group, phosphorane group, phosphate group, phosphatic acid derivative, phosphonate, phosphine group, silane group, silazane group, silicone group, borate group, borane group, halogenide group or a combination thereof. Preferably, the functional chemical compound corresponds to one of the following compound classes: Carboxyl group derivative, ether group, amine group, hydroxy group, carbonyl group, alkane group, alkene group, benzene derivative, pyridine derivate, halogenide group. Generally, a large molecule may include or be a polymer and / or an oligomer. The polymer or oligomer may include one or more subgroups, wherein all subgroups together form the polymer or oligomer. For example, a subgroup can refer to a part of the polymer or oligomer, wherein the subgroups are linked together successively along a chain or network to form the polymer. Preferably, the subgroups of the polymer refer to repeating units that describe a part of the polymer which when repeated produces the polymer chain. However, in some cases, a subgroup can also refer to a single part of the polymer or oligomer that is not repeated. The subgroups may comprise parts that are repeated, for example, a subgroup of a polymer can comprise a repeating core also present in other subgroups and further additional parts that are not repeated and present in other subgroups. The subgroups may include at least one of polymerized monomer or oligomer fragments. The subgroups may include polymerized monomers. In this context, polymerized monomers refer to monomers after their polymerization sometimes also called “mer unit” or “mer”. In particular, polymerized monomers do not refer to monomers, i.e. raw materials, as present in a reaction mixture before polymerization, but refer to repeating units derived from monomers that have been changed during or after the polymerization.
[0022] The formulation may be any formulation including more than two chemical and / or biological components. A formulation comprises of at least two components that may include any chemical and / or biological entity. For example, the component may include small molecules, polymers or the like. However, the components can themselves also be more complex chemical products.
[0023] The biodegradability property may characterize the at least one substance with respect to biodegradability behavior. The biodegradability property may relate to the classification of biodegradable and / or non-biodegradable substances, potentially in relation to one or more habitat type(s) and / or conditions(s). The habitat type may relate to the environment or substrate with respect to which the biodegradation process or the biodegradability is of interest. The habitat type may include soil, compost, sewage water, aquatic system or other suitable environments or substrates of interest. The habitat condition may relate to the condition under which the biodegradation process or the biodegradability is of interest. This may include nutrient concentration of the habitat, temperature, pH, pO2, ionic condition, toxicity or other suitable conditions that influence the biodegradation process.
[0024] The biodegradability property may relate to a quantity characterizing the time evolution of the biodegradation process. The biodegradability property may relate to a measurement quantity characterizing the time evolution of the biodegradation process. The biodegrada- bility property may include a reference quantity as measured under reference measurement conditions. Reference quantities and reference measurement conditions may be provided by OECD, ASTM, ISO Standards or other applicable Standards such as the ones cited in the context of Fig. 1 .
[0025] The at least one substance specific graph representation in contrast to the numeric graph representation(s) is generated by at least one data-driven graph representation model based on at least one graph neural network architecture. The numeric graph representation may be generated based on pre-defined feature vectors that encode the substance structure and / or substance composition as case may be. As a result, the numeric graph representation^) associated with at least with atoms and bonds of the at least one substance may be generated based on pre-defined feature vectors for at least the atoms and bonds of the substance. The feature vectors may represent nodes and edges corresponding to the substance structure and / or substance composition as case may be and the feature vectors may form the numeric graph representation. The substance specific graph representation on the other hand is associated with at least atoms, bonds, and correlations related to atoms and / or bonds. In other words, the data-driven graph representation model may generate a representation of the substance that takes correlations related to atoms and / or bonds into account and thus takes specifics related to the substance structure and / or composition including correlations between structure or composition elements into account.
[0026] In one embodiment at least one substance specification associated with the at least one substance is provided, wherein the at least one specification is mapped to corresponding numeric graph representation(s) associated with at least the atoms and bonds of the at least one substance.
[0027] In another embodiment the at least one data-driven graph representation model is trained by generating persubstance graph representation numeric graph representation(s) related to one or more augmented substance(s), wherein the augmented substances include at least one synthetic change in the numeric graph representation associated at least with atoms, bonds, atom-bond combinations or substance components.
[0028] In another embodiment the at least one data-driven graph representation model is trained using a contrastive loss function that depends on a distance between numeric graph representation^) of different substances and / or augmented substance(s). In another embodiment the data-driven property model is trained on a training dataset including at least one biodegradation property measured for one or more substance(s) and substance specific graph representation(s) generated by at least one data-driven graph representation model.
[0029] In another embodiment at least one substance type relating to a small substance, a large molecule, such as a polymer or oligomer and / or a formulation is provided to the at least one data-driven graph representation model and / or the at least one data-driven property model and / or is provided for selecting the at least one data-driven graph representation model and / or the at least one data-driven property model.
[0030] In another embodiment at least one measurement type relating to the measurement of the biodegradability property is provided to the at least one data-driven graph representation model and / or the at least one data-driven property model and / or is provided for selecting the at least one data-driven graph representation model and / or the at least one data-driven property model.
[0031] In another embodiment based on the substance specific graph representation associated with the at least one substance having the generated biodegradation property, one or more similar substance specific graph representation(s) are generated based on a representation distance measure, wherein one or more similar substance specific graph representation^) and corresponding substance(s) are provided.
[0032] In another embodiment a substance map including multiple substance specific graph representation^) associated with multiple substances and optionally corresponding generated biodegradation properties is generated based on a representation distance measure. The substance map including a map of the substances and optionally corresponding generated biodegradation properties with respect to the representation distance measure may be provided.
[0033] In another embodiment one or more habitat condition(s) are provided that influence the one or more biodegradation properties, wherein the at least one data-driven property model is trained on substance specific graph representation(s) and related one or more biodegradation properties dependent on one or more habitat condition(s). The one or more habitat condition(s) may be mapped to numeric habitat condition(s). The numeric one or more habitat condition(s) may be fused with the substance specific graph representation(s). The fused substance specific graph representation(s) may be provided to the at least one data- driven property model. Fused in this context may contain any vector or matrix operation such as concatenating, adding, dot products or other arithmetic vector or matrix transformations.
[0034] In another embodiment one or more habitat condition(s) are provided that influence the one or more biodegradation properties. The at least one data-driven graph representation model may be trained on numeric graph representation(s) and related one or more habitat condition(s). The one or more habitat condition(s) may be mapped to numeric habitat condition^). One or more numeric habitat condition(s) may be fused with the numeric graph representation(s). The fused numeric graph representation(s) may be provided to the at least one data-driven graph representation model.
[0035] In another embodiment the biodegradation property relates to at least one substance structure associated with a non-biodegradable or biodegradable substance including at least one functional chemical compound, at least one large molecule, at least one formulation, or combinations thereof or at least one small molecule, at least one polymer, and / or at least one formulation. The biodegradation property may relate to the classes biodegradable or non-biodegradable. The biodegradation property may relate to a biodegradability measure in relation to reference measurement and / or a reference habitat. The biodegradation property may relate to the biodegradability measure based on the chemical structure and / or composition of the substance.
[0036] In another embodiment the biodegradation property relates to at least one application of the substance and / or at least one application performance measure of the substance based on the chemical structure and / or composition of the substance.
[0037] BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In the following, the present disclosure is further described with reference to the enclosed figures:
[0039] Fig. 1 illustrates examples of biodegradable substances.
[0040] Fig. 2 illustrates an example method for generating numeric substance representations based on chemical structure specifications.
[0041] Fig. 3 illustrates an example system for generating at least one biodegradability property based on numeric graph representations. Fig. 4 illustrates an example of a training process for the data-driven graph representation model.
[0042] Fig. 5 illustrates an example of the training process for the data-driven property model.
[0043] Figs. 6, 7 illustrate outputs of the trained models.
[0044] Figs. 8, 9 illustrate comparative results from differently trained models.
[0045] DETAILED DESCRIPTION
[0046] Fig. 1 illustrates examples of biodegradable substances.
[0047] Biodegradable substances may be designed to degrade upon disposal by the action of living organisms. Biodegradability may relate to the environmental fate and / or behavior of the substance. Biodegradability may relate to the extent to which the substance can be decomposed by microorganisms such as such as bacteria, fungi or algae. Biodegradability may be dependent on the substance’s chemical structure, chemical weight, physical factors such as cross-linking density, branching, crystallinity or solubility, and exposure conditions such as habitat like soil, compost or aquatic system. With respect to exposure conditions the microorganisms, microbial population, nutrient concentration, temperature, pH, pO2, ionic condition, or substrate characteristics such as toxicity influence biodegradability. Biodegradability may be measured based on measured mass loss (mg / time), dissolved organic carbon (DOC, organic carbon concentration / time), oxygen consumption (e.g. though pressure measurement, e.g. Pa / time) or carbon dioxide production over time (e.g. though pressure measurement, e.g. Pa / time). Fig. 1 illustrates some biodegradable and non-bio- degradable substances including small molecules and polymers. Despite the on first sight similar structures widely different biodegradation properties are exhibited by substances.
[0048] To quantify biodegradability in the sense of a measured property of the substance is challenging and many measurement standards have been developed. Different measurement methods are defined to determine biodegradability under pre-defined laboratory conditions. For example, for wastewater OECD Test No. 301 : “Ready Biodegradability” (July 17, 1992) describes 6 methods for determination of biodegradability. Further for example, ASTM D5988-18 “standard test method for determining aerobic biodegradation of plastic materials in soil” describes the measuring of the carbon dioxide developed by microorganisms as a function of time of exposure, thus measuring the degree of biodegradability relative to a reference material. Further for example ISO 17556:2019 “plastics — determination of the ultimate aerobic biodegradability of plastic materials in soil by monitoring the oxygen demand in a respirometer or the amount of carbon dioxide evolved” yields the optimum rate of biodegradation of plastic material in a test soil by controlling the oxygen consumption or the carbon dioxide production. Further for example, ISO 14855-1 :2012 “determination of the ultimate aerobic biodegradability of plastic materials under controlled composting conditions — method by analysis of evolved carbon dioxide — Part 1 : General method” and ASTM D5338-15 “standard test method for determining aerobic biodegradation of plastic materials under controlled composting conditions, incorporating thermophilic temperatures” determine the ultimate aerobic biodegradability (means by which microorganisms entirely consume a chemical or organic substance in the presence of oxygen) of plastics based on organic compounds under controlled composting conditions by measuring the percentage conversion of the carbon into carbon dioxide and the degree of disintegration of the plastic at the end of the test. ASTM D6400-21 “standard specification for labeling of plastics designed to be aerobically composted in municipal or industrial facilities” additionally includes elemental analysis, plant germination (phytotoxicity), and mesh filtration of the resulting particles. In ISO 17088:2021 “plastics — organic recycling — specifications for compostable plastics” includes the evaluation of negative consequences on the composting process and facility and negative effects on the quality of the resulting compost, including the presence of high levels of regulated metals and other harmful components.
[0049] For aerobic biodegradation ISO 18830:2016 “plastics — determination of aerobic biodegradation of non-floating plastic materials in a seawater / sandy sediment interface — method by measuring the oxygen demand in closed respirometer”, ISO 19679:2020 “plastics — determination of aerobic biodegradation of non-floating plastic materials in a seawater / sediment interface — method by analysis of evolved carbon dioxide” were developed. The biodegradation evaluation is measured by the oxygen demand or the CO2 evolution. Further standards for example include ISO 14853:2016 “plastics — determination of the ultimate anaerobic biodegradation of plastic materials in an aqueous system — method by measurement of biogas production”, ISO 23977-1 :2020 “plastics — determination of the aerobic biodegradation of plastic materials exposed to seawater — Part 1 : method by analysis of evolved carbon dioxide” and ISO 23977-2:2020 “plastics — determination of the aerobic biodegradation of plastic materials exposed to seawater — Part 2: method by measuring the oxygen demand in closed respirometer”.
[0050] The quantified biodegradation property for a substance may depend on the measurement method and conditions used, the measurement environment and the measurement value related to the degradation process, such as mass loss, DOC, oxygen consumption or carbon dioxide production over time. The measurement method and the measured characteristics may be provided as metadata per measurement point related to biodegradability.
[0051] Fig. 2 illustrates an example method for generating numeric substance representations based on chemical structure specifications.
[0052] The chemical structure specification of the biodegradable substance may be mapped to a numerical graph representation. Such mapping may include the determination of feature vectors based on a pre-defined feature specification. Pre-defined feature specifications for chemical structures may include, but not limited to, atom types, atom arrangement like ring or chain, hybridization, number of bonds, bond types, bond arrangement like ring or chain, conjugation, stereo or the like. The features may be implemented as one-hot encoding or in other words based on a pre-defined feature specification. The pre-defined feature specification may relate to atom features corresponding to atoms of the chemical structure. The atom features may be represented as nodes or vertices of the numerical graph representation. The pre-defined feature specification may relate to bond features corresponding to bonds of the chemical structure. The bond features may be represented as edges or arcs of the numerical graph representation. In other words, the graph representation may include nodes corresponding to atoms and vertices corresponding to bonds between two atoms. The feature vector may be assigned per node and vertexthat represent atom types such as C atom or orbital hybridization, and bond type such as double bond or ring structures. Each node may encode atomic information such as the atom type, aromaticity, hybridization, number of bonds the atoms is connected with and the number of bonded hydrogen atoms to treat hydrogen implicitly. In an example, the atom type can be one-hot encoded into a pre-defined number of categorical features based on a predefined list of chemical elements. Edge features (e.g., bond type) may be explicitly included through bond type, bond being part of a ring, conjugation, and stereo. This type of graph data representation may result in a pre-defined number of categorical features per atom and per bond.
[0053] Based on the feature vectors the chemical structure may be represented in a matrix as numerical representation. For example, a feature matrix and / or adjacency matrix may be generated from the feature vectors. This way the substance may be represented as a chemical graph where nodes correspond to atoms and edges correspond to bonds between two atoms. Be assigning the feature vector to each node and each edge that includes information about atom types and bond types, the chemical structure can be mapped from graph representation e.g. via SMILES to the numerical representation e.g. via feature and / or adjacency matrix. The feature and / or adjacency matrix may represent the chemical structure specification of the surfactant in the numerical graph representation that is processable by the graph neural network (GNN).
[0054] For small substances the numeric graph representation may include categorical features per atom and per bond. For large molecules such as polymers the numeric graph representation may include categorical features per atom, per bond and macro molecule arrangement.
[0055] For polymers the monomer structures including the polymerized monomer and / or the raw monomer in unpolymerized state may be represented as chemical fingerprints, e.g. encoded as binary vectors or graph representations as described above. The chain architecture of the polymer may be encoded based on the monomer representation(s). The stoichiometry or polymer architecture may be represented by taking the sum of monomer representation^) weighted by the respective ratios. The stoichiometry or polymer architecture may be represented by architecture vectors of integer values capturing the frequency of different monomer patterns to reflect the monomers' stoichiometry e.g. including the polymerized monomer and / or the raw monomer in unpolymerized state. Graph chemical representations may further include edges to describe the average structure of repeating units weighted by their probability of occurring in polymers. This may reflect (i) the recurrent nature of polymers' repeating units, (ii) the different topologies and isomerisms of polymer chains, and (iii) their varying monomer composition and stoichiometry. The polymer graph representation may include atom and bond representations per repeating unit and / or one or more edge(s) associated with a weight reflecting the probability or frequency of the bond being present per repeating unit. By linking separate monomers e.g. including the polymerized monomer and / or the raw monomer in unpolymerized state with weighted edges, the recurrent nature of polymer chains as well as the ensemble of possible chain architectures may be represented. One possible implementation is for example described in Matteo Aldeghi and Connor W. Coley, “A graph representation of chemical ensembles for polymer property prediction”, Chem. Sci., 2022, 13, 10486-10498.
[0056] For formulations the formulation components or ingredients, their mass ratio and the interactions between formulation components may be represented by the chemical graph representation. For example, the formulation components may be represented by the small or large molecule graph representation described above. The formulation composition may be represented by edges with weights representing the relative concentrations, interactions between the formulation components and / or probabilities related to the concentrations or interactions. This way the chemical structure can be represented by numerical vectors and / or matrices that may serve as input representations to graph neural networks for generating substance specific graph representations.
[0057] Fig. 3 illustrates an example system for generating at least one biodegradability property based on numeric graph representations.
[0058] The system includes at least one data-driven graph representation model configured to map numeric graph representation (s) associated with at least atoms and bonds of the substance^) to substance specific graph representation(s) associated with atoms, bonds, and correlations related to atoms and bonds of the substance(s) and at least one data-driven property model configured to map substance specific graph representation (s) to at least one biodegradation property. The training process for the data-driven models are described in more detail below. In essence the process includes two steps: 1) the data-driven graph representation model generates substance specific graph representation(s) for one or more substances, 2) the substance specific graph representation(s) may be used in a second step to generate at least one biodegradation property. The architecture hence allows to generate substance specific graph representation(s) for multiple substances. Such substance specific graph representation(s) may be stored in relation to substance specifications. The substance specific graph representation(s) may than be used to generate the biodegradation property depending for example on measurement type such as habitat type, measurement method, measurement conditions, property type, application type or the like. This allows for more efficient setup of generation of biodegradation properties, since the substance specific graph representation^) may be generated once and based on such substance specific graph representation space provided by the GNN different data-driven property models may be trained and used.
[0059] The computing system for generating at least one biodegradability property may include a user interface configured to provide a request or instruction related to the substance the at least one biodegradation property is to be generated for. For example, at least one substance specification may be provided the formats like smiles string, graphical representation and / or numerical representation associated with substance.
[0060] The biodegradation property may relate to at least one substance structure associated with a non-biodegradable or biodegradable substance including e.g. at least one small molecule, at least one polymer, and / or at least one formulation. The biodegradation property may relate to the classes biodegradable or non-biodegradable. The biodegradation property may relate to a biodegradability measure in relation to reference measurement and / or a reference habitat, such as mass loss (mg / time), dissolved organic carbon (DOC, organic carbon concentration / time), oxygen consumption (e.g. though pressure measurement, e.g. Pa / time) or carbon dioxide production over time (e.g. though pressure measurement, e.g. Pa / time). The biodegradation property may relate to the biodegradability measure based on the chemical structure and / or composition of the substance. The biodegradation property may relate to at least one application of the substance. The biodegradation property may relate to at least one application performance measure of the substance based on the chemical structure and / or composition of the substance. Examples of application and / or application performance measures are diverse and may include polymers or formulations used for cosmetics or personal care such as surfactants and related performance measures such as foamability, foam rate, or the like.
[0061] The request or instruction may be processed mapping agent configured to map the one or more substance specification(s) to corresponding substance specific graph representation^). The mapping agent may be configured to provide substance specific graph representation^) based on the request including one or more substance specification(s). The mapping agent may be configured to provide substance specific graph representation (s) generated by providing numeric graph representation(s) to at least one data-driven graph representation model configured to map numeric graph representation(s) associated with atoms and bonds of the chemical structure(s) to substance specific graph representation(s) associated with atoms, bonds, and correlations related to atoms and bonds of the chemical structure^). The substance specific graph representation(s) may be stored in a structure store including substance specific graph representation(s) and related substance specification^). The substance specific graph representation(s) may be pre-generated and stored by the at least one data-driven graph representation model. If the substance specific graph representation(s) for the requested substance specification(s) are stored in the structure store or are pre-generated, the mapping agent may be configured to provide the substance specific graph representation(s) corresponding to the one or more substance specification^). If the substance specific graph representation(s) forthe requested substance specification^) are not stored in the structure store or are not pre-generated, the mapping agent may be configured to request model execution by the model execution engine. The mapping agent may be configured to provide the one or more substance specification(s) to the model execution engine. The model execution engine may be configured to access the trained data-driven graph representation model e.g., stored in the model store and to generate based on such access the substance specific graph representation(s).
[0062] The request may include at least one substance type relating to a small substance, a large molecule, such as a polymer or oligomer and / or a formulation. The at least one substance type may be provided with or derived from the one or more substance specification(s). The at least one substance type may be provided to the at least one data-driven graph representation model. The at least one substance type may be provided to select the at least one data-driven graph representation model. The model store may include one or more data-driven graph representation models dependent on the substance type. The model store may include one or more data-driven graph representation models trained on training data related to the substance type. The model store may include one or more data-driven graph representation models per substance type dependent on the substance sub-class. For example, in the case of small molecules as substance type the models may depend on substance classes relating to the type of molecules, functionality of the molecules, applications of the molecules or the like. For example, in the case of polymer as substance type the models may depend on substance classes relating to the type of monomers e.g. including the polymerized monomer and / orthe raw monomer in unpolymerized state, number of monomers per repeating unit such as co-polymers, terpolymers, quarterpolymers or the like, applications or the like. For example, in the case of formulation as substance type the models may depend on substance classes relating to the type of components, solutions, number of components or the like.
[0063] The request may include at least one measurement type relating to the measurement of the biodegradability property. The at least one measurement type may be provided with or derived from the one or more substance specification(s). The at least one measurement type may relate to the one or more measured quantity / ies, such as, but not limited to, mass loss (mg / time), dissolved organic carbon (DOC, organic carbon concentration / time), oxygen consumption (e.g. though pressure measurement, e.g. Pa / time) or carbon dioxide production over time (e.g. though pressure measurement, e.g. Pa / time). The at least one measurement type may relate to the one or more measurement methods, such as, but not limited to, measurement methods based on reference measurement setups as e.g. provided for in the example standards cited in the context of Fig. 1 . The at least one measurement type may relate to the one or more measurement conditions, such as but not limited to, reference measurement conditions as e.g. provided for in the example standards cited in the context of Fig. 1 . The at least one measurement type may relate to one or more habitat type(s) and / or conditions, such as but not limited to, reference measurement habitat types and / or conditions as e.g. provided for in the example standards cited in the context of Fig. 1 . The at least one measurement type may be provided to the at least one data- driven graph representation model. The at least one measurement type may be provided to select the at least one data-driven graph representation model. The model store may include one or more data-driven graph representation models dependent on the measurement type. The model store may include one or more data-driven graph representation models trained on training data related to the measurement type. One or more habitat type(s) and / or condition(s) may be provided that influence the one or more biodegradation properties. The at least one data-driven graph representation model may be trained on numeric graph representation(s) and related one or more habitat type(s) and / or condition(s). The one or more habitat type(s) and / or condition(s) may be mapped to numeric habitat condition(s) e.g. via pre-defined feature vectors as described in the context of Fig. 2. The one or more numeric habitat type(s) and / or condition(s) may be fused to a numeric graph representation (s) associated with the substance and the habitat. The fusion may represent the substance and the habitat like a formulation representation described in the context of Fig. 2. The substrate components, the substance components and their interrelation may be represented by pre-defined feature vectors and the graph representation matrix may be generated. The fused numeric graph representation(s) may be provided to the at least one data-driven graph representation model. The data-driven graph representation model may hence depend on the one or more habitat type(s) and / or condition(s) and may be selected based on one or more habitat type(s) and / or condition(s).
[0064] The mapping agent may be configured to provide the one or more substance specification^) to the model execution engine. The model execution engine may be configured to access the trained data-driven property model e.g. stored in the model store and to generated based on such access the at least one biodegradation property. Based on the provided substance specific graph representation(s) the model execution agent may be configured to generate at least one biodegradation property by providing the substance specific graph representation(s) to the at least one data-driven property model configured to map substance specific graph representation^) to at least one biodegradation property.
[0065] The at least one substance type may be provided with or derived from the one or more substance specification(s). The at least one substance type may be provided to the at least one data-driven property model. The at least one substance type may be provided to select the at least one data-driven property model. The model store may include one or more data-driven property models dependent on the substance type. The model store may include one or more data-driven property models trained on training data related to the substance type. The model store may include one or more data-driven property models per substance type dependent on the substance sub-class.
[0066] The at least one measurement type may be provided with or derived from the one or more substance specification(s). The at least one measurement type may relate to the one or more measured quantity / ies, such as, but not limited to, mass loss (mg / time), dissolved organic carbon (DOC, organic carbon concentration / time), oxygen consumption (e.g. though pressure measurement, e.g. Pa / time) or carbon dioxide production over time (e.g. though pressure measurement, e.g. Pa / time). The at least one measurement type may relate to the one or more measurement methods, such as, but not limited to, measurement methods based on reference measurement setups as e.g. provided for in the example standards cited in the context of Fig. 1 . The at least one measurement type may relate to the one or more measurement conditions, such as but not limited to, reference measurement conditions as e.g. provided for in the example standards cited in the context of Fig.
[0067] 1 . The at least one measurement type may relate to one or more habitat type(s) and / or conditions, such as but not limited to, reference measurement habitat types and / or conditions as e.g. provided for in the example standards cited in the context of Fig. 1. Habitat types may relate to soil, compost, aquatic system, sewage water or the like. Habitat conditions may relate to nutrient content, toxicity or other characteristics of the habitat affecting biodegradation processes. The at least one measurement type may be provided to the at least one data-driven property model. The at least one measurement type may be provided to select the at least one data-driven property model. The model store may include one or more data-driven property models dependent on the measurement type. The model store may include one or more data-driven property models trained on training data related to the measurement type.
[0068] One or more habitat type(s) and / or condition(s) may be provided, e.g., by the request, that influence the one or more biodegradation properties. The at least one data-driven property model may be trained on substance specific graph representation(s) and related one or more biodegradation properties dependent on one or more habitat type(s) and / or condition^). The one or more habitat condition(s) may be mapped to numeric habitat condition^). The one or more habitat type(s) and / or condition(s) may be mapped to numeric habitat condition(s) e.g. via pre-defined feature vectors as described in the context of Fig.
[0069] 2. The numeric one or more habitat condition(s) may be fused with the substance specific graph representation (s). The fusion of the numeric representations may include one or more operations such as concatenation, summing, dot products or the like. The fused substance specific graph representation (s) may be provided to the at least one property model. The property model may hence depend on the one or more habitat type(s) and / or condition^) and may be selected based on one or more habitat type(s) and / or condition(s).
[0070] The mapping agent and / or the model execution agent may be configured to provide at least one generated biodegradation property for the one or more substance specification(s) characterizing the biodegradability of the associated substance(s). From the substance specific graph representation(s) associated with the substance structure(s) having the generated at least one biodegradation property, one or more similar substance specific graph representation^) may be generated based on a representation distance measure. The one or more similar substance specific graph representation(s) and corresponding biodegradable substance structure(s) may be provided. For easy navigation and overview the one or more similar substance specific graph representation(s) and corresponding biodegradable substance structure(s) may be provided in the form of a diffusion map. This way a substance map including multiple substance specific graph representation(s) associated with biodegradable substance structures having the generated biodegradation properties may be generated based on the representation distance measure and displayed.
[0071] The at least one data-driven graph representation model and / orthe at least one data-driven property model may be trained based on training data sets relating to substance specification^) and corresponding numeric graph representation (s) for training the data-driven graph representation model and substance specific graph representation(s) and corresponding biodegradation property / ies fortraining the data-driven property model. Examples of model architectures, training process and model characteristics will be described in more detail in the context of the following figures. This shall not be considered limiting, since multiple implementations exist and the examples merely serve as illustrative examples.
[0072] Fig. 4 illustrates an example of a training process for the data-driven graph representation model.
[0073] At least one data-driven graph representation model may be trained to map numeric graph representation(s) associated with at least atoms and bonds of the substance(s) to substance specific graph representation(s) associated with at least atoms, bonds, and correlations related to atoms and bonds of the substance(s). The data-driven graph representation model may be trained to generate numeric graph representation(s) for one substance type such as small molecules, large molecules like polymers or oligomers, formulations or substance habitat combinations.
[0074] For training substance specification (s) may be provided and mapped to the numeric graph representation as for example described in the context of Fig. 2. The numeric graph representation per substance may be augmented to represent multiple augmented substance structures. The augmentation may include at least one synthetic change in atoms, bonds, atom-bond combinations, components of a macro molecule, components of small molecules or combinations of components. Multiple graph neural networks may be provided separately with augmented numeric graph representations per substance. The graph neural network may include a convolutional graph neural network or an isomorphism graph neural network. As an example a graph convolution network may be defined by
[0075] As another example a graph isomorphism network may be defined by
[0076] Thus, each graph representing the substance may have a D dimensional representation in a D dimensional space. For each layer of the graph neural network the node and / or bond states may be updated using neighboring states such as nearest neighbor atoms, next nearest neighbor atoms and so on. The updating may also be referred to as message passing. The update function for each layer may be provided for example by using the notation described above:
[0077] Wherein MLP is one example of a regression model such as a multilayer perceptron (MLP).
[0078] The MLP may include input and output layers with multiple hidden layers in between. They may utilize activation functions at each of their calculated layers. Hence, the inputs may be pushed for-ward through the MLP by taking the dot product of the input with the weights that exist between the input layer and the first hidden layer. This dot product may yield a value at the hidden layer. Then the calculated output at the current hidden layer may be transformed through one or more activation function(s) such as rectified linear units (ReLU), sigmoid function, ortanh. Once the calculated output at the hidden layer has been pushed through the activation function, it may be pushed to the next layer in the MLP by taking the dot product with the corresponding weights. These steps may be repeated until the output layer is reached. At the output layer, the calculations will either be used for a backpropagation algorithm that corresponds to the activation function that was selected for the MLP (in the case of training).
[0079] The graph neural network may provide multiple representations per augmented substance. The results per augmented substance may be pooled by one or more pooling functions such as sum, mean, max, set2set or the like to generate the substance specific graph representation per augmented substance.
[0080] The pooled representations per augmented substance may be provided to a loss function for contrastive learning. The loss function may relate to similarity measures measuring the distance of the representations computer per augmented substance per substance. An example Loss function may relate to cos similarity measure for example using the known Tanimoto similarity. For two graphs i,j from same group or per substance the loss function may defined by: with zf: RNrepresentation of i-th molecule.
[0081] This way the at least one data-driven graph representation model may be trained using a contrastive loss function that may depend on a distance measure between substance specific graph representation(s) of different substances and / or augmented substances. By training the graph neural network based on the substance representations and their augmentations, the substance specific graph representations may be learned by learning the relevance of different components of the substance representations and their correlations. This way similar substances may be mapped into a D-dimensional embedding space and the contrastive learning to enables to learn a feature space that can combine or put together points that are related and push apart points that are not related. The contrastive loss function basically tries to minimize the distance between the similar substance specific graph representations und to maximize the distance between unrelated substance specific graph representations. This may also be referred to as an unsupervised learning approach for generating substance specific graph representations. Further details of such approaches may be found for example in Wang, Y., Wang, J., Cao, Z. et al. Molecular contrastive learning of representations via graph neural networks. Nat Mach Intell 4, 279-287 (2022). https: / / doi.org / 10.1038 / s42256-022-00447-x or Improving Molecular Contrastive Learning via Faulty Negative Mitigation and Decomposed Fragment Contrast, Yuyang Wang, Rishikesh Magar, Chen Liang, and Amir Barati Farimani, Journal of Chemical Information and Modeling 2022 62 (1 1), 2713-2725, DOI: 10.1021 / acs.jcim.2c00495.
[0082] Fig. 5 illustrates an example of the training process for the data-driven property model.
[0083] The substance specific graph representations, as generated for example by the GNN trained as described in the context of Fig. 4, may be stored in relation to their substance specification. For the data driven property model, training data set including the at least one biodegradation property as measured for the substance associated with the substance specification or corresponding substance specific graph representations may be provided. The data-driven property model may be trained on the training dataset including at least one biodegradation property measured for one or more substance(s) and substance specificgraph representation(s) generated by the at least one data-driven graph representation model. The data driven model architecture may include any suitable architecture for supervised learning to generate the at least one biodegradation property based on the substance specific graph representation(s). An example of a suitable regression model may be a random forest regression model that maps the feature vectors of the substance specific graph representation(s) according to a tree structure to the at least one biodegradation property. Other possible architectures may range from simple classification models discerning biodegradable from non-biodegradable via gradient boosting methods to more elaborate e.g. transformer-based models generating the at least one biodegradation property.
[0084] Fig. 5 illustrates the training process. On the input layer of the model, the substance specific graph representation(s) potentially fused with additional feature vectors such as habitat type and / or condition, measurement type and / or condition or the like as e.g. alluded to in the context of Fig. 3, may be provided. The model weights and / or structures may be trained to generate the at least one biodegradation property. The loss function may be defined to minimize the difference between the measured at least one biodegradation property from the training data set and the generated at least one biodegradation property. Depending on property type, substance type and / or measurement type multiple models may be trained. Here transfer learning, ensemble learning or other common techniques may be used for training.
[0085] On inference or in use the thus trained data-driven property prediction model may map a given substance specific graph representation and potentially additional representations to at least one biodegradation property.
[0086] Figs. 6, 7 illustrate outputs of the trained models.
[0087] To illustrate the results of the biodegradation property generation the Fig. 6 and 7 illustrate user interfaces displaying such results. In Fig. 6 the molecular structures for small molecules are used as an illustrative basis. The molecular structure provided by the request and based on which the biodegradation property was generated is displayed together with molecular structures that are similar to the requested molecular structure. Next to the structure and the biodegradation property, in this case for example, the mass loss after 10 days, the confidence interval e.g. in terms of standard deviation of the model output, and similarity measures calculated based on numeric distance measures between respective substance specific graph representation (s).
[0088] In Fig. 7 the representations of molecular structures are used to produce a 2-d similarity or diffusion map. Each dot on the map represents one substance, e.g. small molecule. The distance between the points represents the distance measure between respective substance specific graph representation(s).
[0089] Figs. 8, 9 illustrate comparative results from differently trained models.
[0090] In table of Fig. 8 GNN models including property prediction GNN1 are compared with GNN models using the two-step approach of combining GNN model with separate property prediction model GNN2. In addition, classical fingerprint models such as Morgan fingerprints not based on GNN generated fingerprints MF were compared. More details on Morgan Fingerprints may be found in The Generation of a Unique Machine Description for Chemical Structures-A Technique Developed at Chemical Abstracts Service. H. L. Morgan Journal of Chemical Documentation 1965 5 (2), 107-1 13 DOI: 10.1021 / c160017a018.
[0091] Fig. 9 illustrates different pooling and head options and compares standard error or confidence measures. The present disclosure has been described in conjunction with preferred embodiments and examples as well. However, other variations can be understood and effected by those persons skilled in the art and practicing the claimed invention, from the studies of the drawings, this disclosure and the claims.
[0092] Any steps presented herein can be performed in any order. The methods disclosed herein are not limited to a specific order of these steps. It is also not required that the different steps are performed at a certain place or in a certain computing node of a distributed system, i.e. each of the steps may be performed at different computing nodes using different equipment / data processing.
[0093] As used herein ..determining" also includes ..initiating or causing to determine", “generating" also includes ..initiating and / or causing to generate" and “providing” also includes “initiating or causing to determine, generate, select, send and / or receive”. “Initiating or causing to perform an action” includes any processing signal that triggers a computing node or device to perform the respective action.
[0094] In the claims as well as in the description the word “comprising” or “including” or similar wording does not exclude other elements or steps and shall not be construed limiting to the elements or steps lined out. The indefinite article “a” or “an” does not exclude a plurality. A single element or other unit may fulfill the functions of several entities or items recited in the claims. The mere fact that certain measures are recited in the mutual different dependent claims does not indicate that a combination of these measures cannot be used in an advantageous implementation or further elements may be included.
[0095] Providing in the scope of this disclosure may include any interface configured to provide data. This may include an application programming interface, a human-machine interface such as a display and / or a software module interface. Providing may include communication of data or submission of data to the interface, in particular display to a user or use of the data by the receiving entity.
[0096] Any disclosure and embodiments described herein relate to methods, systems, apparatuses, devices, chemicals, materials, services, uses, computer program elements lined out above and vice versa. Advantageously, the benefits provided by any of the embodiments and examples equally apply to all other embodiments and examples and vice versa. All terms and definitions used herein are understood broadly and have their general meaning.
Claims
Claims1 . A method for generating at least one biodegradability property characterizing at least one substance, the method comprising the steps of:- providing at least one substance specific graph representation generated by providing numeric graph representation(s) to at least one data-driven graph representation model trained to map numeric graph representation(s) associated with at least with atoms and bonds of the at least one substance to at least one substance specific graph representation associated with at least atoms, bonds, and correlations related to atoms and / or bonds;- generating at least one biodegradation property by providing the at least one substance specific graph representation to the at least one data-driven property model trained to map the at least one substance specific graph representation to at least one biodegradation property;- providing the at least one generated biodegradation property for characterizing the biodegradability of the at least one substance.
2. The method of claim 1 , wherein at least one substance specification associated with the at least one substance is provided, wherein the at least one specification is mapped to corresponding numeric graph representation(s) associated with at least the atoms and bonds of the at least one substance.
3. The method of any of the preceding claims, wherein the at least one data-driven graph representation model is trained by generating per substance graph representation numeric graph representation(s) related to one or more augmented substance^), wherein the augmented substances include at least one synthetic change in the numeric graph representation associated at least with atoms, bonds, atombond combinations or substance components.
4. The method of any of the preceding claims, wherein the at least one data-driven graph representation model is trained using a contrastive loss function that depends on a distance between numeric graph representation(s) of different substances and / or augmented substance(s).
5. The method of any of the preceding claims, wherein the data-driven property model is trained on a training dataset including at least one biodegradation property measured for one or more substance(s) and substance specific graph representation(s) generated by at least one data-driven graph representation model.
6. The method of any of the preceding claims, wherein at least one substance type relating to a small substance, a large molecule, such as a polymer or oligomer and / or a formulation is provided to the at least one data-driven graph representation model and / or the at least one data-driven property model and / or is provided for selecting the at least one data-driven graph representation model and / or the at least one data- driven property model.
7. The method of any of the preceding claims, wherein at least one measurement type relating to the measurement of the biodegradability property is provided to the at least one data-driven graph representation model and / orthe at least one data-driven property model and / or is provided for selecting the at least one data-driven graph representation model and / or the at least one data-driven property model.
8. The method of any of the preceding claims, wherein, based on the substance specific graph representation associated with the at least one substance having the generated biodegradation property, one or more similar substance specific graph representation^) are generated based on a representation distance measure, wherein one or more similar substance specific graph representation(s) and corresponding substance(s) are provided.
9. The method of any of the preceding claims, wherein a substance map including multiple substance specific graph representation(s) associated with multiple substances and optionally corresponding generated biodegradation properties is generated based on a representation distance measure, wherein the substance map including a map of the substances and optionally corresponding generated biodegradation properties with respect to the representation distance measure is provided.
10. The method of any of the preceding claims, wherein one or more habitat condition(s) are provided that influence the one or more biodegradation properties, wherein the at least one data-driven property model is trained on substance specific graph representation^) and related one or more biodegradation properties dependent on one or more habitat condition(s), wherein the one or more habitat condition(s) aremapped to numeric habitat condition(s), wherein numeric one or more habitat condition^) are fused with the substance specific graph representation(s), wherein the fused substance specific graph representation(s) are provided to the at least one data-driven property model.
11. The method of any of the preceding claims, wherein one or more habitat condition(s) are provided that influence the one or more biodegradation properties, wherein the at least one data-driven graph representation model is trained on numeric graph representation(s) and related one or more habitat condition(s), wherein the one or more habitat condition(s) are mapped to numeric habitat condition(s), wherein one or more numeric habitat condition(s) are fused with the numeric graph representation^), wherein the fused numeric graph representation(s), are provided to the at least one data-driven graph representation model.
12. The method of any of the preceding claims, wherein the biodegradation property relates to at least one substance structure associated with a non-biodegradable or biodegradable substance including at least one functional chemical compound, at least one large molecule, at least one formulation, or combinations thereof, wherein the biodegradation property relates to the classes biodegradable or non-biodegrada- ble, wherein the biodegradation property relates to a biodegradability measure in relation to reference measurement and / or a reference habitat, wherein the biodegradation property relates to the biodegradability measure based on the structure and / or composition of the substance.
13. The method of any of the preceding claims, wherein the biodegradation property relates to at least one application of the substance and / or at least one application performance measure of the substance based on the chemical structure and / or composition of the substance.
14. Use of the biodegradation property and associated substance structure generated and selected according to the methods of any of claims 1 -13 to produce a biodegradable substance with the selected substance structure.