Deep neural network for biodegradability

By generating biodegradable properties through data-driven graphical representation and property models, the environmental pollution problem in the treatment of functional compound waste is solved, and the reliable and standardized generation of biodegradable substances is realized.

CN122459879APending Publication Date: 2026-07-24BASF SE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In the existing technology, the widespread use of functional compounds has led to a large amount of non-degradable waste, especially the algal bloom problem caused by phosphate accumulation, which necessitates the development of reliable and standardized methods for generating biodegradable materials.

Method used

By using data-driven graph representation models, atoms and bonds of matter are mapped to matter-specific graph representations, and data-driven property models are used to generate biodegradable properties, providing computer-implemented methods and devices for generating biodegradable substances.

Benefits of technology

It enables more reliable and standardized generation of biodegradable substances, solving the environmental pollution problem in the treatment of functional compound waste, especially the problem of phosphate accumulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122459879A_ABST
    Figure CN122459879A_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of biodegradable substances or molecules, such as small molecules or macromolecules or formulations using machine learning. Methods, devices, computer elements, biodegradable substances or uses for generating the biodegradability properties of one or more biodegradable substances are disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of biodegradable substances or molecules, such as small or large molecules or formulations using machine learning. Methods, apparatus, computer components, substances, or uses for generating one or more substances with biodegradable properties are disclosed. Background Technology

[0002] Functional compounds are generally defined as small molecules that are widely used in industrial and / or everyday products due to their broad range of applications. Uses of functional compounds include coatings, personal care products, detergents, lubricants, packaging, and foams. However, this widespread application also results in a large amount of waste containing used functional compounds. Disposing of non-degradable waste in undesignated environments is a problem. In particular, the accumulation of chemicals in the environment, such as phosphate buildup that causes algal blooms, is undesirable. Therefore, functional compounds that can be degraded are not the only option. Summary of the Invention

[0003] In one aspect, a method for generating at least one biodegradable property characterizing at least one substance is disclosed, specifically a computer-implemented method comprising the following steps:

[0004] - Provide at least one substance-specific graph representation generated by providing a numerical graph representation to at least one data-driven graph representation model, the at least one data-driven graph representation model being configured or trained to map numerical graph representations associated with at least atoms and bonds of at least one substance to at least one substance-specific graph representation associated with at least atoms, bonds and interrelationships related to atoms and / or bonds;

[0005] - Generate at least one biodegradable property by providing at least one substance-specific graph representation to at least one data-driven property model, wherein the at least one data-driven property model is configured or trained to map at least one substance-specific graph representation to at least one biodegradable property;

[0006] - Provide at least one biodegradable property generated, used to characterize the biodegradability of at least one substance.

[0007] In another aspect, an apparatus for generating at least one biodegradable property characterizing at least one substance is disclosed, the method comprising the following steps:

[0008] - An input interface configured to provide at least one substance-specific graph representation generated by providing a numerical graph representation to at least one data-driven graph representation model, the at least one data-driven graph representation model being configured or trained to map numerical graph representations associated with at least atoms and bonds of at least one substance to at least one substance-specific graph representation associated with at least atoms, bonds, and interrelationships related to atoms and / or bonds;

[0009] - A property generator configured to generate at least one biodegradable property by providing at least one substance-specific graph representation to at least one data-driven property model, wherein the at least one data-driven property model is configured or trained to map at least one substance-specific graph representation to at least one biodegradable property.

[0010] - Output interface, which is configured to provide at least one of the generated biodegradability properties for characterizing the biodegradability of at least one substance.

[0011] On the other hand, the use of generating and selecting biodegradable properties and associated material structures according to the methods or apparatus disclosed herein for producing biodegradable substances having the selected material structures is disclosed.

[0012] In another aspect, a computer element, such as a computer-readable storage medium, a computer program, or a computer program product, is disclosed, including instructions that, when executed by a computing node or computing system, direct the computing node or computing system to perform the methods disclosed herein.

[0013] Any disclosures and embodiments described herein relate to methods, apparatus, surfactants, uses, and computer components. Advantageously, the benefits provided by any embodiments and examples also apply to all other embodiments and examples.

[0014] Implementation Plan

[0015] In the following sections, embodiments and / or examples will be outlined to illustrate implementations of this disclosure. It should be understood that this disclosure is not limited to the described embodiments and / or examples.

[0016] The methods, apparatus, computer components, biodegradable substances, or uses disclosed herein allow for more reliable and standardized generation of the biodegradable properties of biodegradable substances.

[0017] Substances may contain or be functional compounds, such as small molecules, macromolecules (such as polymers or oligomers), formulations, or combinations thereof.

[0018] Functional compounds can include or be any type of functional compound. Functional compounds are typically compounds with properties used for technical purposes (i.e., fulfilling the function in a chemical product). For example, compounds that provide UV protection in sunscreens are functional compounds. Functional compounds can contain active ingredients, which are components that provide the function of the functional compound. Furthermore, active ingredients can provide the biological activity of the functional compound. For example, active ingredients can refer to antifungal agents, aromatic chemicals, UV absorbers, food additives, vitamins, nutrients, dyes, and surfactants. Functional compounds are typically characterized by their chemical structure. However, different chemical structures can exist within a single functional compound. For example, a functional compound can consist of components that are molecules undergoing tautomerism, protonation, or deprotonation. Functional compounds can consist of more than one stereoisomer. Therefore, a functional compound composed of one type of molecule can be associated with one or more chemical structures, such as a protonated structure and an uncharged structure, or two stereoisomers. Therefore, generally, a functional compound can include all molecules associated with a chemical formula through (de)protonation, isomerization (such as tautomerism and stereoisomerization). Furthermore, a functional compound can refer to any functional compound that can be described by one or more chemical structures. In embodiments, the chemical structure can be associated with a chemical formula.

[0019] Preferably, the functional compound is composed of small molecules. Preferably, the functional compound has a molecular weight of less than 10,000 g / mol. More preferably, the compound has a molecular weight of less than 800 g / mol, and even more preferably less than 400 g / mol. Furthermore, it is preferred that the functional compound exists in the environment in a molecular form that allows for a complete description of the functional compound using a simple structural formula containing relevant information. A simple molecular structure refers to a molecule that can be clearly described by covalent bonds between the atoms of the molecule. Examples of this are, for example, systems with dynamic equilibrium between several forms (such as monomers and oligomers) in the case of several inorganic acids, or ionic substances with highly localized charges that strongly interact with the solvent (e.g., via hydrogen bonding). Preferably, the functional compound has at least one of the following properties: it has an effect on living organisms, is suitable for affecting the structure of living organisms, or is suitable for affecting the function of living organisms. In one embodiment, the functional compound comprises at least one of the following functional groups: ether group, hydroxyl group, peroxy group, hydroperoxide group, carboxyl group, carboxyl group derivative, carbonyl group, amine group, imine group, hydrazine group, urea group, urethane group, thiourethane group, nitrile group, azide group, azo group, cyanate group, isocyanate group, isocyanate group, pyridine group, alkane group, olefin group, alkyne group, phenyl group, ketone group, thionyl group. Groups, aldehyde groups, thioaldehyde groups, acetal groups, ketal groups, oxime groups, hydrazone groups, nitro groups, nitroso groups, thiol groups, sulfide groups, disulfide groups, sulfonic acid derivatives, sulfinic acid derivatives, sulfate ester groups, sulfate ester group derivatives, sulfone groups, sulfoxide groups, mercapto groups, sulfide groups, phosphoalkyl groups, phosphate ester groups, phosphate derivatives, phosphonates, phosphonates, silanealkyl groups, silazanealkyl groups, organosilicon groups, borate ester groups, boranyl groups, halide groups, or combinations thereof. Preferably, the functional compound corresponds to one of the following compound categories: carboxyl derivatives, ether groups, amine groups, hydroxyl groups, carbonyl groups, alkane groups, olefin groups, benzene derivatives, pyridine derivatives, and halide groups.

[0020] Typically, macromolecules can comprise either polymers and / or oligomers. A polymer or oligomer can comprise one or more subgroups, where all subgroups together form the polymer or oligomer. For example, a subgroup can refer to a portion of a polymer or oligomer, where the subgroups are continuously linked together along a chain or network to form the polymer. Preferably, a subgroup of a polymer refers to a repeating unit describing a portion of the polymer, which, when repeated, produces a polymer chain. However, in some cases, a subgroup can also refer to a single, non-repeating portion of the polymer or oligomer. A subgroup can contain repeating portions; for example, a subgroup of a polymer can contain a repeating core also present in other subgroups and additional, non-repeating portions present in other subgroups. A subgroup can contain at least one of a polymeric monomer or an oligomer fragment. A subgroup can contain a polymeric monomer. In this context, a polymeric monomer refers to the monomer after its polymerization, sometimes also called a "monomer unit" or "monomer." In particular, a polymeric monomer does not refer to a monomer present in the reaction mixture prior to polymerization, i.e., a raw material, but rather to a repeating unit derived from a monomer that has been modified during or after polymerization.

[0021] The formulation can be any formulation containing more than two chemical and / or biological components. The formulation contains at least two components, which can comprise any chemical and / or biological entity. For example, components can comprise small molecules, polymers, etc. However, these components themselves can also be more complex chemical products.

[0022] Biodegradability can characterize the properties of at least one substance in terms of its biodegradable behavior. Biodegradability can relate to the classification of biodegradable and / or non-biodegradable substances, potentially relating to one or more habitat types and / or conditions. Habitat type can relate to the environment or substrate of interest regarding biodegradation processes or biodegradability. Habitat type can include soil, compost, wastewater, aquatic systems, or other suitable environments or substrates of interest. Habitat conditions can relate to the conditions of interest regarding biodegradation processes or biodegradability. This can include nutrient concentrations, temperature, pH, pO2, ionic conditions, toxicity, or other suitable conditions affecting biodegradation processes.

[0023] Biodegradability properties can be correlated with quantities characterizing the time evolution of a biodegradation process. Biodegradability properties can be correlated with measurements characterizing the time evolution of a biodegradation process. Biodegradability properties may include reference quantities, such as those measured under reference measurement conditions. Reference quantities and reference measurement conditions may be defined by OECD, ASTM, ISO standards, or other applicable standards (such as...). Figure 1 (The standard referenced in the context) is provided.

[0024] In contrast to numerical graph representations, at least one substance-specific graph representation is generated by at least one data-driven graph representation model based on at least one graph neural network architecture. Numerical graph representations can be generated based on predefined feature vectors that encode the substance's structure and / or composition according to specific circumstances. Therefore, numerical graph representations associated with at least at least one substance's atoms and bonds can be generated based on predefined feature vectors of at least the substance's atoms and bonds. These feature vectors can represent nodes and edges corresponding to the substance's structure and / or composition according to specific circumstances, and these feature vectors can form numerical graph representations. On the other hand, substance-specific graph representations are associated with at least atoms, bonds, and the interrelationships related to atoms and / or bonds. In other words, data-driven graph representation models can generate representations of substances that consider the interrelationships related to atoms and / or bonds, and therefore consider details related to the substance's structure and / or composition, including the interrelationships between structural or constituent elements.

[0025] In one embodiment, at least one substance specification associated with at least one substance is provided, wherein the at least one specification is mapped to a corresponding numerical diagram representation associated with at least one atom and bond of the at least one substance.

[0026] In another implementation, at least one data-driven graph representation model is trained by generating numerical graph representations associated with one or more reinforcing substances through each material graph representation, wherein the reinforcing substances include at least one synthetic variation in the numerical graph representation that is associated with at least one atom, bond, atom-bond combination, or material component.

[0027] In another implementation, a contrastive loss function is used to train at least one data-driven graph representation model, which depends on the distance between the numerical graph representations of different materials and / or reinforcing materials.

[0028] In another implementation, the data-driven property model is trained on a training dataset that includes at least one biodegradation property measured for one or more substances and a substance-specific graph representation generated by at least one data-driven graph representation model.

[0029] In another embodiment, at least one type of substance relating to small substances, macromolecules (such as polymers or oligomers), and / or formulations is provided to at least one data-driven graphical representation model and / or at least one data-driven property model, and / or is provided for selecting at least one data-driven graphical representation model and / or at least one data-driven property model.

[0030] In another embodiment, at least one measurement type related to the measurement of biodegradability properties is provided to at least one data-driven graphical representation model and / or at least one data-driven property model, and / or is provided for selecting at least one data-driven graphical representation model and / or at least one data-driven property model.

[0031] In another embodiment, one or more similar substance-specific map representations are generated based on a distance metric, based on a substance-specific map representation associated with at least one substance having the generated biodegradable properties, wherein one or more similar substance-specific map representations and corresponding substances are provided.

[0032] In another embodiment, a material map is generated based on a distance metric, the material map including multiple material-specific graphical representations associated with various materials and optionally corresponding generated biodegradability properties. A material map including mappings of materials and optionally corresponding generated biodegradability properties relative to the distance metric can be provided.

[0033] In another implementation, one or more habitat conditions influencing one or more biodegradability properties are provided, wherein at least one data-driven property model is trained on a substance-specific graph representation and the associated one or more biodegradability properties dependent on the one or more habitat conditions. The one or more habitat conditions may be mapped to numerical habitat conditions. The one or more habitat conditions in numerical form may be fused with a substance-specific graph representation. The fused substance-specific graph representation may be provided to at least one data-driven property model. Fusion in this context may involve any vector or matrix operations, such as concatenation, addition, dot product, or other arithmetic vector or matrix transformations.

[0034] In another implementation, one or more habitat conditions influencing one or more biodegradation properties are provided. At least one data-driven graph representation model can be trained on a numerical graph representation and the associated one or more habitat conditions. The one or more habitat conditions can be mapped to numerical habitat conditions. The one or more numerical habitat conditions can be fused with the numerical graph representation. The fused numerical graph representation can be provided to at least one data-driven graph representation model.

[0035] In another embodiment, biodegradability refers to the structure of at least one substance associated with a non-biodegradable or biodegradable substance, which includes at least one functional compound, at least one macromolecule, at least one formulation, or a combination thereof, or at least one small molecule, at least one polymer, and / or at least one formulation. Biodegradability may refer to a category of biodegradable or non-biodegradable. Biodegradability may refer to a biodegradability measure associated with a reference measurement and / or a reference habitat. Biodegradability may refer to a biodegradability measure based on the chemical structure and / or composition of the substance.

[0036] In another embodiment, the biodegradability relates to at least one application of the substance and / or to at least one application energy level of the substance based on its chemical structure and / or composition. Attached Figure Description

[0037] The present disclosure is further described below with reference to the accompanying drawings:

[0038] Figure 1 Examples of biodegradable substances are shown.

[0039] Figure 2 An example method for generating numerical material representations based on chemical structure specifications is illustrated.

[0040] Figure 3 An example system for generating at least one biodegradable property based on a numerical graph representation is illustrated.

[0041] Figure 4 An example illustrating the training process of a data-driven graph representation model is provided.

[0042] Figure 5 An example illustrating the training process of a data-driven model is provided.

[0043] Figure 6 , Figure 7 The output of the trained model is shown as an example.

[0044] Figure 8 , Figure 9 Examples of comparison results from models trained in different ways are shown. Detailed Implementation

[0045] Figure 1 Examples of biodegradable substances are shown.

[0046] Biodegradable substances can be programmed to degrade upon disposal through the action of living organisms. Biodegradability can be related to the environmental fate and / or behavior of a substance. Biodegradability can involve the extent to which a substance can be broken down by microorganisms such as bacteria, fungi, or algae. Biodegradability can depend on the substance's chemical structure, chemical weight, physical factors such as crosslinking density, branching, crystallinity, or solubility, and exposure conditions such as habitat, such as soil, compost, or aquatic systems. Regarding exposure conditions, microorganisms, microbial populations, nutrient concentrations, temperature, pH, pO2, ionic conditions, or substrate characteristics such as toxicity affect biodegradability. Biodegradability can be measured based on measured mass loss over time (mg / time), dissolved organic carbon (DOC, organic carbon concentration / time), oxygen consumption (e.g., by pressure measurement, such as Pa / time), or carbon dioxide production (e.g., by pressure measurement, such as Pa / time). Figure 1 Examples of biodegradable and non-biodegradable substances, including small molecules and polymers, are shown. Despite their seemingly similar structures, these substances exhibit a wide range of different biodegradable properties.

[0047] Quantifying biodegradability in the sense of measured properties of a substance is challenging, and numerous measurement standards have been developed. Different measurement methods have been defined to determine biodegradability under predefined laboratory conditions. For example, for wastewater, OECD Test No. 301, “Ready Biodegradability” (July 17, 1992), describes six methods for determining biodegradability. Furthermore, for example, ASTM D5988-18, “Standard Test Method for Determining Aerobic Biodegradation of Plastic Materials in Soil,” describes measuring the change in carbon dioxide produced by microorganisms over exposure time, thereby measuring the degree of biodegradability relative to a reference material. Furthermore, for example, ISO 17556:2019 "plastics—determination of the ultimate aerobic biodegradability of plastic materials in soil by monitoring the oxygen demand in a respirometer or the amount of carbon dioxide evolved" determines the optimal biodegradability of plastic materials in the test soil by controlling oxygen consumption or carbon dioxide production.Furthermore, for example, ISO 14855-1:2012 "determination of the ultimate aerobic biodegradability of plastic materials under controlled composting conditions—method by analysis of evolved carbondioxide—Part 1: General method" and ASTM D5338-15 "standard test method for determining aerobic biodegradation of plastic materials under controlled composting conditions, incorporating thermophilic temperatures" determine the ultimate aerobic biodegradability of organic-based plastics under controlled composting conditions (the way microorganisms completely consume chemical or organic matter in the presence of oxygen) by measuring the percentage of carbon converted to carbon dioxide and the degree of disintegration of the plastic at the end of the test. ASTM D6400-21, "Standard Specification for Labeling of Plastics Designed to Be Aerobically Composted in Municipal or Industrial Facilities," also includes elemental analysis, plant germination (phytotoxicity), and sieve filtration of the resulting granules. ISO 17088:2021, "Plastics—Organic Recycling—Specifications for Compostable Plastics," includes assessments of the negative impacts on the composting process and facilities, as well as on the quality of the resulting compost, including the presence of high levels of regulated metals and other hazardous components.

[0048] For aerobic biodegradation, ISO 18830:2016 "plastics—determination of aerobic biodegradation of non-floating plastic materials in a seawater / sandysediment interface—method by measuring the oxygen demand in a closed respirometer" and ISO 19679:2020 "plastics—determination of aerobic biodegradation of non-floating plastic materials in a seawater / sediment interface—method by analysis of evolved carbon dioxide" have been established. Biodegradation is evaluated by measuring oxygen demand or CO2 emissions.Other standards include ISO 14853:2016 "plastics—determination of the ultimate anaerobic biodegradation of plastic materials in an aqueous system—method by measurement of biogas production", ISO 23977-1:2020 "plastics—determination of the aerobic biodegradation of plastic materials exposed to seawater—Part 1: method by analysis of evolved carbondioxide", and ISO 23977-2:2020 "plastics—determination of the aerobic biodegradation of plastic materials exposed to seawater—Part 2: method by measuring the oxygen demand in closed respirometer".

[0049] The quantitative biodegradability of a substance can depend on the measurement methods and conditions used, the measurement environment, and measurements related to the degradation process, such as mass loss over time, DOC, oxygen consumption, or carbon dioxide production. Measurement methods and measured characteristics can be provided as metadata for each measurement point related to biodegradability.

[0050] Figure 2 An example method for generating numerical material representations based on chemical structure specifications is illustrated.

[0051] The chemical structure specification of a biodegradable substance can be mapped to a numerical graph representation. This mapping can include determining eigenvectors based on predefined feature specifications. Predefined feature specifications for a chemical structure can include, but are not limited to, atom type, atomic arrangement (e.g., ring or chain), hybridization, number of bonds, bond type, bond arrangement (e.g., ring or chain), conjugation, and stereochemistry. Features can be implemented as one-hot encoded features, or in other words, based on predefined feature specifications. Predefined feature specifications can involve atomic features corresponding to the atoms in the chemical structure. Atomic features can be represented as nodes or vertices in the numerical graph representation. Predefined feature specifications can also involve bond features corresponding to the bonds in the chemical structure. Bond features can be represented as edges or arcs in the numerical graph representation. In other words, the graph representation can include nodes corresponding to atoms and vertices corresponding to the bonds between two atoms. Eigenvectors can be assigned to each node and vertex, representing the atom type (e.g., C atom or orbital hybridization) and bond type (e.g., double bond or ring structure). Each node can encode atomic information such as atom type, aromaticity, hybridization, the number of bonds the atom is connected to, and the number of bonded hydrogen atoms, implicitly handling hydrogen. In one example, atom types can be uniquely encoded as a predefined number of categorical features based on a predefined list of chemical elements. Edge features (e.g., bond types) can be explicitly included by bond type, bonds as part of a ring, conjugation, and stereotype. This type of graph data representation can result in a predefined number of categorical features for each atom and each bond.

[0052] Based on eigenvectors, chemical structures can be represented numerically in matrices. For example, eigenvectors can generate eigenmatrices and / or adjacency matrices. In this way, a substance can be represented as a chemical graph, where nodes correspond to atoms and edges correspond to bonds between two atoms. By assigning eigenvectors to each node and each edge, which contain information about the types of atoms and bonds, the chemical structure can be mapped from a graph representation (e.g., via SMILES) to a numerical representation (e.g., via eigenvectors and / or adjacency matrices). The eigenvectors and / or adjacency matrices can represent the chemical structure specification of surfactants in a numerical graph representation that can be processed by graph neural networks (GNNs).

[0053] For small substances, numerical graphical representations can include categorical features of each atom and each bond. For macromolecules (such as polymers), numerical graphical representations can include categorical features of each atom, each bond, and the arrangement of the macromolecules.

[0054] For polymers, the monomer structure, including polymeric monomers and / or unpolymerized pristine monomers, can be represented as a chemical fingerprint, for example, encoded as a binary vector or graph representation as described above. The chain architecture of the polymer can be encoded based on the monomer representations. The stoichiometry or polymer architecture can be represented by summing the monomer representations weighted by appropriate ratios. The stoichiometry or polymer architecture can be represented by an architecture vector of integer values ​​that captures the frequency of different monomeric patterns to reflect the stoichiometry of the monomers, for example, including polymeric monomers and / or unpolymerized pristine monomers. The graph chemical representation can also include edges to describe the average structure of repeating units weighted by their probability of occurrence in the polymer. This can reflect (i) the recurrence nature of the polymer repeating units, (ii) the different topologies and isomerism of the polymer chains, and (iii) their different monomeric compositions and stoichiometry. The polymer graph representation can include the atomic and bond representations of each repeating unit and / or one or more edges associated with weights reflecting the probability or frequency of bond presence in each repeating unit. By connecting individual monomers (e.g., pristine monomers in a polymeric state and / or an unpolymerized state) with weighted edges, the reproducible properties of polymer chains, as well as the overall set of possible chain architectures, can be represented. A possible specific implementation is described, for example, in Matteo Aldeghi and Connor W. Coley, “A graph representation of chemical ensembles for polymer property prediction,” Chem. Sci., 2022, 13, 10486-10498.

[0055] For formulations, the formulation components or ingredients, their mass ratios, and the interactions between the formulation components can be represented by chemical graphs. For example, formulation components can be represented by the small molecule or macromolecule graphs described above. The formulation composition can be represented by weighted edges, which represent relative concentrations, interactions between formulation components, and / or probabilities related to concentration or interactions. Thus, chemical structures can be represented by numerical vectors and / or matrices, which can be used as input representations for graph neural networks to generate substance-specific graph representations.

[0056] Figure 3 An example system for generating at least one biodegradable property based on a numerical graph representation is illustrated.

[0057] The system includes at least one data-driven graph representation model configured to map numerical graph representations associated with at least the atoms and bonds of a substance to substance-specific graph representations associated with atoms, bonds, and the interrelationships related to the atoms and bonds of the substance; and at least one data-driven property model configured to map the substance-specific graph representations to at least one biodegradable property. The training process of the data-driven model is described in more detail below. Essentially, the process comprises two steps: 1) the data-driven graph representation model generates substance-specific graph representations for one or more substances; 2) the substance-specific graph representations can be used in the second step to generate at least one biodegradable property. Therefore, the architecture allows for the generation of substance-specific graph representations for multiple substances. Such substance-specific graph representations can be stored in relation to substance specifications. The substance-specific graph representations can then be used to generate biodegradable properties, depending on factors such as the measurement type, habitat type, measurement method, measurement conditions, property type, application type, etc. This allows for more efficient setup of biodegradable property generation because the substance-specific graph representations can be generated all at once, and different data-driven property models can be trained and used based on this substance-specific graph representation space provided by the GNN.

[0058] A computational system for generating at least one biodegradable property may include a user interface configured to provide requests or instructions relating to a substance for which at least one biodegradable property is to be generated. For example, at least one substance specification may be provided, in a format such as a smiley string associated with the substance, a graphical representation, and / or a numerical representation.

[0059] Biodegradability properties can relate to the structure of at least one substance associated with a non-biodegradable or biodegradable substance, including, for example, at least one small molecule, at least one polymer, and / or at least one formulation. Biodegradability properties can relate to categories of biodegradable or non-biodegradable substances. Biodegradability properties can relate to biodegradability measures associated with reference measurements and / or reference habitats, such as mass loss over time (mg / time), dissolved organic carbon (DOC, organic carbon concentration / time), oxygen consumption (e.g., by pressure measurement, such as Pa / time), or carbon dioxide production (e.g., by pressure measurement, such as Pa / time). Biodegradability properties can relate to biodegradability measures based on the chemical structure and / or composition of the substance. Biodegradability properties can relate to at least one application of the substance. Biodegradability properties can relate to at least one application energy measure of the substance based on its chemical structure and / or composition. Examples of applications and / or application energy measures are diverse and can include polymers or formulations used in cosmetics or personal care, such as surfactants, and related energy measures such as foaming properties, foaming rate, etc.

[0060] A request or instruction can be a processed mapping agent configured to map one or more material specifications to corresponding material-specific graph representations. The mapping agent can be configured to provide material-specific graph representations based on a request including one or more material specifications. The mapping agent can be configured to provide material-specific graph representations generated by providing numerical graph representations to at least one data-driven graph representation model configured to map numerical graph representations associated with atoms and bonds of a chemical structure to material-specific graph representations associated with the atoms, bonds, and their interrelationships. Material-specific graph representations can be stored in a structure repository that includes the material-specific graph representations and associated material specifications. Material-specific graph representations can be pre-generated and stored by at least one data-driven graph representation model. If a material-specific graph representation for a requested material specification is stored in the structure repository or pre-generated, the mapping agent can be configured to provide a material-specific graph representation corresponding to one or more material specifications. If a material-specific graph representation for a requested material specification is not stored in the structure repository or pre-generated, the mapping agent can be configured to request model execution by a model execution engine. The mapping agent can be configured to provide one or more material specifications to the model execution engine. The model execution engine can be configured to access, for example, trained data-driven graph representation models stored in a model repository, and generate matter-specific graph representations based on such access.

[0061] The request may include at least one substance type relating to small substances, macromolecules (such as polymers or oligomers), and / or formulations. At least one substance type may be provided with or derived from one or more substance specifications. At least one substance type may be provided to at least one data-driven graphical representation model. At least one substance type may be provided to select at least one data-driven graphical representation model. The model repository may include one or more data-driven graphical representation models depending on the substance type. The model repository may include one or more data-driven graphical representation models trained on training data related to the substance type. The model repository may include one or more data-driven graphical representation models for each substance type depending on the substance subclass. For example, in the case of small molecules as the substance type, the model may depend on the substance category related to molecule type, molecule functionality, molecule application, etc. For example, in the case of polymers as the substance type, the model may depend on the substance category related to monomer type, such as including polymerizable monomers and / or unpolymerized raw monomers, the number of monomers per repeating unit (such as copolymers, terpolymers, quaternary copolymers, etc.), application, etc. For example, in the case of formulations as the substance type, the model may depend on the substance category related to component type, solution, number of components, etc.

[0062] The request may include at least one type of measurement relating to the measurement of biodegradability properties. At least one type of measurement may be provided with or obtained from one or more material specifications. At least one type of measurement may involve one or more quantities of measurements, such as, but not limited to, mass loss over time (mg / time), dissolved organic carbon (DOC, organic carbon concentration / time), oxygen consumption (e.g., by pressure measurement, such as Pa / time), or carbon dioxide production (e.g., by pressure measurement, such as Pa / time). At least one type of measurement may involve one or more measurement methods, such as, but not limited to, measurement methods based on a reference measurement setting, such as, for example, in… Figure 1 The example standards referenced in the context are provided. At least one measurement type may involve one or more measurement conditions, such as, but not limited to, reference measurement conditions, as exemplified by, in... Figure 1 The example standards cited in the context of this study provide that at least one measurement type may involve one or more habitat types and / or conditions, such as, but not limited to, reference measurements of habitat types and / or conditions, as exemplified in [the context of] [the study]. Figure 1 The example standards referenced in the context of this document provide that at least one measurement type may be provided to at least one data-driven graph representation model. At least one measurement type may be provided to select at least one data-driven graph representation model. A model repository may include one or more data-driven graph representation models that depend on the measurement type. A model repository may include one or more data-driven graph representation models trained on training data relevant to the measurement type.

[0063] One or more habitat types and / or conditions influencing one or more biodegradation properties can be provided. At least one data-driven graph representation model can be trained on a numerical graph representation and associated one or more habitat types and / or conditions. One or more habitat types and / or conditions can be provided, for example, via... Figure 2 The predefined feature vectors described in the context are mapped to numerical habitat conditions. One or more numerical habitat types and / or conditions can be fused into a numerical graph representation associated with matter and habitat. The fusion can represent matter and habitat, such as... Figure 2 The formulation representation described in the context. Substrate components, material components, and their interrelationships can be represented by predefined eigenvectors, and a graphical representation matrix can be generated. The fused numerical graphical representation can be provided to at least one data-driven graphical representation model. Therefore, the data-driven graphical representation model can depend on one or more habitat types and / or conditions, and can be selected based on one or more habitat types and / or conditions.

[0064] A mapping agent can be configured to provide one or more material specifications to a model execution engine. The model execution engine can be configured to access, for example, a trained data-driven property model stored in a model repository, and generate at least one biodegradable property based on such access. Based on the provided material-specific graph representation, the model execution agent can be configured to generate at least one biodegradable property by providing the material-specific graph representation to at least one data-driven property model, which is configured to map the material-specific graph representation to at least one biodegradable property.

[0065] At least one material type may be provided together with or obtained from one or more material specifications. At least one material type may be provided to at least one data-driven property model. At least one material type may be provided to select at least one data-driven property model. The model repository may include one or more data-driven property models depending on the material type. The model repository may include one or more data-driven property models trained on training data related to the material type. The model repository may include one or more data-driven property models for each material type depending on the material subclass.

[0066] At least one measurement type may be provided with or obtained from one or more material specifications. At least one measurement type may involve one or more measurement quantities, such as, but not limited to, mass loss over time (mg / time), dissolved organic carbon (DOC, organic carbon concentration / time), oxygen consumption (e.g., by pressure measurement, such as Pa / time), or carbon dioxide production (e.g., by pressure measurement, such as Pa / time). At least one measurement type may involve one or more measurement methods, such as, but not limited to, measurement methods based on a reference measurement setting, such as, for example, in… Figure 1 The example standards referenced in the context are provided. At least one measurement type may involve one or more measurement conditions, such as, but not limited to, reference measurement conditions, as exemplified by, in... Figure 1 The example standards cited in the context of this study provide that at least one measurement type may involve one or more habitat types and / or conditions, such as, but not limited to, reference measurements of habitat types and / or conditions, as exemplified in [the context of] [the study]. Figure 1The example standards referenced in the context are provided. Habitat type can refer to soil, compost, aquatic systems, wastewater, etc. Habitat conditions can refer to the nutrient content, toxicity, or other characteristics of the habitat that affect the biodegradation process. At least one measurement type can be provided to at least one data-driven property model. At least one measurement type can be provided to select at least one data-driven property model. The model repository can include one or more data-driven property models that depend on the measurement type. The model repository can include one or more data-driven property models trained on training data related to the measurement type.

[0067] One or more habitat types and / or conditions influencing one or more biodegradation properties can be provided, for example, by requesting such information. At least one data-driven property model can be trained on a material-specific graphical representation and on one or more related biodegradation properties dependent on one or more habitat types and / or conditions. One or more habitat conditions can be mapped to numerical habitat conditions. One or more habitat types and / or conditions can be provided, for example, via... Figure 2 The predefined eigenvectors described in the context are mapped to numerical habitat conditions. One or more habitat conditions in numerical form can be fused with a material-specific graph representation. The fusion of numerical representations can include one or more operations, such as join, summation, dot product, etc. The fused material-specific graph representation can be provided to at least one property model. Therefore, the property model can depend on one or more habitat types and / or conditions, and can be selected based on one or more habitat types and / or conditions.

[0068] Mapping agents and / or model execution agents can be configured to provide at least one generated biodegradable property for one or more material specifications, thereby characterizing the biodegradability of the associated material. Based on the material-specific graph representation associated with a material structure having at least one generated biodegradable property, one or more similar material-specific graph representations can be generated based on a representational distance metric. One or more similar material-specific graph representations and corresponding biodegradable material structures can be provided. For ease of navigation and overview, one or more similar material-specific graph representations and corresponding biodegradable material structures can be provided in the form of a diffusion map. Thus, a material map comprising multiple material-specific graph representations associated with biodegradable material structures having the generated biodegradable property can be generated and displayed based on a representational distance metric.

[0069] At least one data-driven graph representation model and / or at least one data-driven property model can be trained based on training datasets relating to the material specifications and corresponding numerical graph representations used to train the data-driven graph representation model, and the material-specific graph representations and corresponding biodegradation properties used to train the data-driven property model. Examples of model architectures, training processes, and model characteristics will be described in more detail in the context of the following figures. This should not be considered limiting, as multiple specific implementations exist, and the examples are used for illustrative purposes only.

[0070] Figure 4 An example illustrating the training process of a data-driven graph representation model is provided.

[0071] At least one data-driven graph representation model can be trained to map a numerical graph representation associated with at least the atoms and bonds of a substance to a substance-specific graph representation associated with at least the atoms, bonds, and the interrelationships related to the atoms and bonds of the substance. The data-driven graph representation model can be trained to generate a numerical graph representation of a substance type, such as small molecules, macromolecules (e.g., polymers or oligomers), formulations, or combinations of substance habitats.

[0072] To enable training, a matter specification can be provided and mapped to a numerical graph representation, such as, for example... Figure 2 As described in the context, numerical representations of each substance can be enhanced to represent multiple enhanced substance structures. Enhancements can include at least one synthetic change in atoms, bonds, atom-bond combinations, macromolecular components, small molecule components, or combinations of components.

[0073] Multiple graph neural networks can be provided individually, each with an enhanced numerical graph representation for each substance. Graph neural networks can include convolutional graph neural networks or isomorphic graph neural networks. As an example, a graph convolutional network can be defined by the following formula:

[0074]

[0075] As another example, graph isomorphic networks can be defined by the following formula:

[0076]

[0077] Therefore, each graph representing matter can have a D-dimensional representation in D-dimensional space. For each layer of the graph neural network, node and / or bond states can be updated using neighboring states such as nearest neighbor atoms, second nearest neighbor atoms, etc. This update can also be referred to as message passing. For example, the update function for each layer can be provided, for instance, by using the notation described above:

[0078] ,

[0079] MLP is an example of a regression model, such as a Multilayer Perceptron (MLP). An MLP can include an input layer and an output layer, with multiple hidden layers between them. They can utilize activation functions at each of the layers in which they are computed. Therefore, the input is forward-pushed through the MLP by taking the dot product of the input and the weights present between the input layer and the first hidden layer. This dot product produces values ​​at the hidden layers. The computed output at the current hidden layer can then be transformed by one or more activation functions, such as a Rectified Linear Unit (ReLU), a sigmoid function, or tanh. Once the computed output at the hidden layer has been forward-pushed by the activation function, it can be forward-pushed to the next layer in the MLP by taking the dot product with the corresponding weights. These steps can be repeated until the output layer is reached. At the output layer, computation is used for a backpropagation algorithm, which corresponds to the activation function chosen for the MLP (in the case of training).

[0080] Graph neural networks can provide multiple representations for each augmentation. The results for each augmentation can be pooled using one or more pooling functions (such as summation, mean, maximum, set2set, etc.) to generate a material-specific graph representation for each augmentation.

[0081] Pooled representations of each enhancement can be fed into a loss function for contrastive learning. The loss function can involve a similarity metric measuring the distance between the computer representations of each enhancement and each material. Example loss functions could involve a cosine similarity metric, such as using the known Tanimoto similarity. For two graphs from the same group or each material... The loss function can be defined as:

[0082]

[0083] ,

[0084] in Indicates the first Each molecule.

[0085] Thus, at least one data-driven graph representation model can be trained using a contrastive loss function, which can depend on a distance metric between material-specific graph representations of different substances and / or augmented substances. By training a graph neural network based on the substance representations and their augmentations, material-specific graph representations can be learned by learning the correlations and interrelationships between the different components of the substance representations. This allows similar substances to be mapped into a D-dimensional embedding space, and contrastive learning enables the learning of a feature space that can group or combine related points while pushing unrelated points apart. The contrastive loss function essentially attempts to minimize the distance between similar material-specific graph representations and maximize the distance between unrelated material-specific graph representations. This can also be referred to as an unsupervised learning method for generating material-specific graph representations. Further details of this approach can be found, for example, Wang, Y., Wang, J., Cao, Z. et al., Molecular contrastive learning of representations via graph neural networks, Nat Mach Intell 4, 279–287 (2022). https: / / doi.org / 10.1038 / s42256-022-00447-x or Improving Molecular Contrastive Learning via FaultyNegative Mitigation and Decomposed Fragment Contrast, Yuyang Wang, Rishikesh Magar, Chen Liang and AmirBarati Farimani, Journal of Chemical Information and Modeling 2022 62 (11), 2713-2725, DOI: 10.1021 / acs.jcim.2c00495.

[0086] Figure 5 An example illustrating the training process of a data-driven model is provided.

[0087] For example, through, as Figure 4The substance-specific graph representation generated by a GNN trained as described in the context can be stored in relation to its substance specification. For data-driven property models, a training dataset can be provided comprising at least one biodegradable property or its corresponding substance-specific graph representation for a substance measurement associated with a substance specification. The data-driven property model can be trained on a training dataset comprising at least one biodegradable property for one or more substance measurements and substance-specific graph representations generated by at least one data-driven graph representation model. The data-driven model architecture can include any suitable architecture for supervised learning to generate at least one biodegradable property based on the substance-specific graph representation. An example of a suitable regression model could be a random forest regression model that maps feature vectors of the substance-specific graph representation to at least one biodegradable property according to a tree structure. Other possible architectures range from simple classification models that distinguish between biodegradable and non-biodegradable substances via gradient boosting methods to more sophisticated transformer-based models that generate at least one biodegradable property.

[0088] Figure 5 The training process is illustrated. At the input layer of the model, a material-specific graph representation, potentially fused with additional feature vectors, can be provided, such as, for example, in... Figure 3 The habitat type and / or conditions, measurement type and / or conditions, etc., are indirectly mentioned in the context. Model weights and / or structure can be trained to generate at least one biodegradability property. The loss function can be defined as minimizing the difference between the measured at least one biodegradability property from the training dataset and the generated at least one biodegradability property. Multiple models can be trained based on property type, substance type, and / or measurement type. Here, training can employ transfer learning, ensemble learning, or other commonly used techniques.

[0089] When inferred or used, such a data-driven property prediction model can map a given substance-specific graphical representation and potential additional representations to at least one biodegradability property.

[0090] Figure 6 , Figure 7 The output of the trained model is shown as an example.

[0091] To illustrate the results generated by biodegradation, Figure 6 and Figure 7 An example of a user interface for displaying such results is shown. Figure 6In this context, the molecular structure of the small molecule serves as an exemplary basis. A molecular structure, provided upon request and whose biodegradability is generated based on the molecular structure, is shown alongside molecular structures similar to the requested molecular structure. In addition to structure and biodegradability, in this case, for example, mass loss after 10 days, confidence intervals (e.g., in terms of the standard deviation of the model output), and a similarity metric calculated based on the numerical distance between the corresponding substance-specific graphical representations are also included.

[0092] exist Figure 7 In this model, molecular structure representations are used to generate 2D similarity or diffusion maps. Each point on the map represents a substance, such as a small molecule. The distance between points represents a measure of the distance between the corresponding substance-specific map representations.

[0093] Figure 8 , Figure 9 Examples of comparison results from models trained in different ways are shown.

[0094] exist Figure 8 The table compares GNN models that include property prediction GNN1 with GNN models using a two-step approach that combines the GNN model with a separate property prediction model GNN2. Additionally, classic fingerprint models, such as Morgan fingerprints (not based on GNN-generated fingerprints MF), are compared. More details on Morgan fingerprints can be found in *The Generation of a Unique Machine Description for Chemical Structures—A Technique Developed at Chemical Abstracts Service*, HL Morgan Journal of Chemical Documentation 1965 5 (2), 107-113 DOI: 10.1021 / c160017a018.

[0095] Figure 9 Different pooling and header options are illustrated, and the standard error or confidence metric is compared.

[0096] This disclosure is also described in conjunction with preferred embodiments and examples. However, by studying the accompanying drawings, this disclosure, and the claims, those skilled in the art will understand and implement other variations of the claimed invention.

[0097] Any step presented in this paper can be performed in any order. The methods disclosed herein are not limited to a specific order of these steps. Nor is it necessary to perform different steps at a specific location in a distributed system or on a specific computing node; that is, each step can be performed at different computing nodes using different equipment / data processing.

[0098] As used herein, "determine" also includes "initiate or cause determination," "generate" also includes "initiate and / or cause generation," and "provide" also includes "initiate or cause determination, generation, selection, transmission, and / or reception." "Initiate or cause execution of an action" includes any processing signal that triggers a computing node or device to perform a corresponding action.

[0099] In the claims and description, the words "comprising" or "including" or similar terms do not exclude other elements or steps and should not be construed as limiting oneself to the listed elements or steps. The indefinite articles "a" or "an" do not exclude a plurality. A single element or other unit may perform the function of several entities or items recited in the claims. The fact that certain measures are recited only in mutually different dependent claims does not mean that combinations of these measures cannot be used in advantageous embodiments or that additional elements may be included.

[0100] The provision within the scope of this disclosure may include any interface configured to provide data. This may include application programming interfaces, human-machine interfaces (such as displays), and / or software module interfaces. The provision may include communication of data or submission of data to the interface, particularly displaying data to a user or using data by a receiving entity.

[0101] Any disclosure and embodiments described herein relate to the methods, systems, apparatuses, devices, chemicals, materials, services, uses, and computer program elements listed above, and vice versa. Advantageously, the benefits provided by any embodiments and examples also apply to all other embodiments and examples, and vice versa.

[0102] All terms and definitions used in this document should be understood broadly and have their general meaning.

Claims

1. A method for generating at least one biodegradability property characterizing at least one substance, the method comprising the following steps: - Provide at least one substance-specific graph representation generated by providing a numerical graph representation to at least one data-driven graph representation model, the at least one data-driven graph representation model being trained to map numerical graph representations associated with at least atoms and bonds of the at least one substance to at least one substance-specific graph representation associated with at least atoms, bonds and interrelationships related to atoms and / or bonds; - Generate at least one biodegradable property by providing the at least one substance-specific graph representation to at least one data-driven property model, wherein the at least one data-driven property model is trained to map the at least one substance-specific graph representation to at least one biodegradable property; - Provide at least one biodegradable property generated, for characterizing the biodegradability of the at least one substance.

2. The method of claim 1, wherein at least one substance specification is provided associated with the at least one substance, wherein the at least one specification is mapped to a corresponding numerical representation associated with at least one atom and one bond of the at least one substance.

3. The method according to any one of the preceding claims, wherein the at least one data-driven graph representation model is trained by generating numerical graph representations associated with one or more reinforcing substances through each material graph representation, wherein the reinforcing substances include at least one synthetic variation in the numerical graph representation that is associated with at least one atom, bond, atom-bond combination or material component.

4. The method according to any one of the preceding claims, wherein the at least one data-driven graph representation model is trained using a contrastive loss function, the contrastive loss function depending on the distance between the numerical graph representations of different materials and / or reinforcing materials.

5. The method according to any one of the preceding claims, wherein the data-driven property model is trained on a training dataset comprising at least one biodegradation property measured for one or more substances and a substance-specific graph representation generated by at least one data-driven graph representation model.

6. The method according to any one of the preceding claims, wherein at least one type of substance relating to small substances, macromolecules such as polymers or oligomers and / or formulations is provided to the at least one data-driven graphical representation model and / or the at least one data-driven property model, and / or is provided for selecting the at least one data-driven graphical representation model and / or the at least one data-driven property model.

7. The method according to any one of the preceding claims, wherein at least one measurement type relating to the measurement of the biodegradability property is provided to the at least one data-driven graphical representation model and / or the at least one data-driven property model, and / or is provided for selecting the at least one data-driven graphical representation model and / or the at least one data-driven property model.

8. The method according to any one of the preceding claims, wherein, Based on the substance-specific graph representation associated with the at least one substance having the generated biodegradable properties, one or more similar substance-specific graph representations are generated based on a distance metric, wherein one or more similar substance-specific graph representations and corresponding substances are provided.

9. The method according to any one of the preceding claims, wherein a material mapping map is generated based on a distance representation metric, the material mapping map including multiple material-specific graph representations associated with multiple substances and optionally corresponding generated biodegradable properties, wherein the material mapping map is provided including a mapping of the substances and optionally corresponding generated biodegradable properties relative to the distance representation metric.

10. The method according to any one of the preceding claims, wherein one or more habitat conditions affecting one or more biodegradable properties are provided, wherein the at least one data-driven property model is trained on a material-specific graph representation and a related one or more biodegradable properties depending on the one or more habitat conditions, wherein the one or more habitat conditions are mapped to numerical habitat conditions, wherein the one or more habitat conditions in numerical form are fused with the material-specific graph representation, wherein the fused material-specific graph representation is provided to the at least one data-driven property model.

11. The method according to any one of the preceding claims, wherein one or more habitat conditions influencing the one or more biodegradable properties are provided, wherein the at least one data-driven graph representation model is trained on a numerical graph representation and the associated one or more habitat conditions, wherein the one or more habitat conditions are mapped to numerical habitat conditions, wherein the one or more numerical habitat conditions are fused with the numerical graph representation, wherein the fused numerical graph representation is provided to the at least one data-driven graph representation model.

12. The method according to any one of the preceding claims, wherein the biodegradability relates to at least one substance structure associated with a non-biodegradable or biodegradable substance, the non-biodegradable or biodegradable substance comprising at least one functional compound, at least one macromolecule, at least one formulation, or a combination thereof, wherein the biodegradability relates to a category of biodegradable or non-biodegradable substances, wherein the biodegradability relates to a biodegradability measure associated with a reference measurement and / or a reference habitat, and wherein the biodegradability relates to the biodegradability measure based on the structure and / or composition of the substance.

13. The method according to any one of the preceding claims, wherein the biodegradability relates to at least one application of the substance and / or to at least one application energy level of the substance based on the chemical structure and / or composition of the substance.

14. The use of the biodegradable properties and associated material structures generated and selected by the method of any one of claims 1 to 13 for producing biodegradable substances having the selected material structures.