Apparatus for determining the properties of substances
A machine learning-based property model classifies atoms into adaptive classes using atomic descriptors to predict substance properties, addressing limitations of traditional models by enhancing accuracy and reducing data requirements, facilitating efficient material selection.
Patent Information
- Application Number
- JP2025534759
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-16
- Filing Date
- 2023-12-13
- Publication Date
- 2026-01-06
AI Technical Summary
Existing thermodynamic models struggle with accurately predicting properties of substances, especially for novel functional groups or combinations, due to insufficient training data and limited applicability, requiring complex decomposition and manual tuning, and are not adaptable to different chemical environments.
A machine learning-based property model that classifies atoms into adaptive classes using atomic descriptors, allowing for accurate property determination with reduced training data, applicable to a wide variety of substances, including those not in the training set, without relying on interatomic bonds or bond orders.
Enables rapid identification of suitable substances for specific applications by predicting properties like melting point, vapor pressure, and thermal conductivity, reducing the need for extensive synthesis and testing, and improving accuracy and efficiency in material selection.
Smart Images

Figure 2026500300000001_ABST
Abstract
Description
[Technical Field]
[0001] FIELD OF THE INVENTION The present invention relates to an apparatus, a method and a computer program product for determining the properties of a pure substance or a substance in a given mixture, and further to a training apparatus, a training method and a computer program product for training a machine learning-based property model that can be used in the aforementioned apparatus, method and computer program product for determining the properties of a substance. [Background technology]
[0002] Background of the Invention In general, determining the thermodynamic properties of substances, defined as those that exhibit molecular properties or those that exhibit extended periodic 1D, 2D, or 3D structures, is a challenging problem of high industrial relevance. Many established thermodynamic models, particularly in the field of molecular materials, are so-called "group contribution" models, in which properties are determined from contributions attributable to hierarchically derived molecular functional groups. However, deriving the hierarchy to be followed to decompose a molecular structure into its individual functional groups is not only a tedious and time-consuming task for experts or complex software, but also the source of many inaccuracies in the prediction of thermodynamic properties. For example, functional groups that fit the same hierarchical pattern, when presented in different chemical environments, may exhibit significantly different behavior. However, current models cannot take such differences into account. Therefore, such predictive models for thermodynamic properties are often applicable only to a specific, very limited chemical space, and therefore, it is often not possible to screen a huge number of molecules for specific properties in an industrially relevant manner. Furthermore, even to achieve acceptable accuracy in their limited application fields, these predictive models require enormous amounts of training data to train each model, and as a result, models can only be provided for groups of molecules for which such an extensive data base already exists. However, for example, for highly novel functional groups or novel combinations of closely spaced functional groups, there is often insufficient data available to achieve the accuracy of such models that would enable practical application.
[0003] It would therefore be advantageous to provide a predictive model for determining physicochemical properties of substances that (a) allows for the determination of properties for substances belonging to classes other than those in the training set, and (b) requires less training data while allowing for greater accuracy in determining properties. Summary of the Invention [Means for solving the problem]
[0004] Summary of the Invention An object of the present invention is to provide an apparatus, method, and computer program product that enables highly accurate determination of properties of materials defined as either exhibiting molecular properties or representing extended periodic 1D, 2D, or 3D structures. A further object of the present invention is to provide a training apparatus, method, and computer program that (a) enables property determination for materials belonging to classes other than those in the training set, and (b) provides predictive models that can be used in apparatuses, methods, and computer programs for property determination, which can be trained to provide high determination accuracy using less training data. The flexibility of the present invention also enables the generation of predictive models for a wide variety of target properties, because the mathematical structure allows the model's predictor to naturally adapt to the determined properties. In contrast to classical group contribution methods, the present model does not need to be manually tuned to the target property by an expert. Furthermore, because the present model does not require interatomic bonds and bond orders as input, it is applicable to objects where these are difficult to establish, such as metal complexes and hydrogen bonds.
[0005] In a first aspect of the present invention, there is provided an apparatus for determining properties of a material defined as either exhibiting molecular properties or representing an extended periodic 1D, 2D or 3D structure, the apparatus comprising: i) an atomic descriptor providing unit for providing atomic descriptors that characterize atoms of the material with respect to a structure of a particular molecule; ii) a trained property model providing unit for providing a trained machine learning based property model, the trained property model being adapted to determine properties of the material as output when the atomic descriptors are provided as input, the property model comprising: a) an atomic classifier adapted to classify atoms of the material into one or more atomic classes based on the atomic descriptors; and b) a regression model adapted to determine properties based on the atomic classes determined for the material; iii) a property determination unit for determining properties of the material by applying the trained property model to the atomic descriptors; and iv) a property providing unit for providing the properties for further processing.
[0006] The property model includes a) an atomic classifier adapted to classify atoms of a substance into one or more atomic classes based on atomic descriptors, and b) a regression model adapted to determine properties based on the atomic classes determined for the substance. Specifically, both the atomic classifier and the regression model are part of a machine learning-based property model and can therefore be trained simultaneously, making use of atomic classes specifically determined for each training situation. Thus, the atomic classes used to determine properties do not rely on insights into the physical or chemical properties of molecules provided by experts or used as scientific conventions. Rather, the atomic classes themselves are determined by training specific to each training situation, where classifications may differ for each property considered. Additionally, these atomic classes map complex structural and chemical motifs into a low-dimensional space consisting of a limited number of categories. This allows the property model to be trained using less training data, further enabling greater accuracy in determining properties for each training situation, and finally, making it possible to make predictions for substances containing functional groups not covered by the training set. All of this is a result of the fact that the classifications thus performed are targeted to atomic features that are truly important for the property of interest.
[0007] Furthermore, due to the large number of anticipated, but often insufficiently investigated, substances potentially suitable for a particular application, today, technical product engineers faced with the technical challenge of finding a substance that is not only suitable for a particular application but also satisfies numerous additional target properties for each, must synthesize and test a vast number of anticipated substances or search through huge datasets and libraries in which potential substances are stored to find each substance that may be suitable for the application. Even when sophisticated experimental design methods are used, it is still necessary to synthesize and experimentally test a large number of possible substances. In this regard, the above-described method can assist users, such as technical product engineers, in automatically finding potentially suitable substances much more quickly. Specifically, by using the above-described method, users need only synthesize and test substances that are determined to be highly likely to satisfy the respective target properties, while substances that are far less likely to be suitable can be avoided. Therefore, the method allows users to quickly and more efficiently perform the technical task of finding a substance suitable for a technical application.
[0008] Generally, the material to be characterized is considered to exhibit either molecular properties or an extended periodic 1D, 2D, or 3D structure. Preferably, the material exhibits molecular properties, i.e., is a molecular material. Even more preferably, the material exhibits molecular properties and contains at least one of the following elements: H, C, N, O, F, Si, P, Cl, Br, and I. In general, the principles of the present invention can be applied to any type of material for which a respective training data set exists. However, the larger the intended application range to which the model is to be applied, the more training data must be provided to reach sufficient accuracy. Furthermore, it is preferred that the molecular material has a molecular weight of less than 600 g / mol, more preferably less than 300 g / mol. Furthermore, it is preferred that the specific molecule exists in the environment in a form that allows the molecule to be completely described using a simple structural formula containing the relevant information. A simple molecular structure refers to a molecule that can be unambiguously described by the covalent bonds between its atoms. Examples where this is not the case are systems that have a dynamic equilibrium between several forms, such as monomers and oligomers, as is the case for some inorganic acids, or ionic species with highly localized charges that interact strongly with the solvent, for example, via hydrogen bonding. Materials can also generally refer to periodic systems, i.e., materials in which certain molecules are structurally repeated in a defined manner in space; for example, in such cases, materials can refer to crystalline metals or semiconductors, or other systems that include periodic atomic symmetry.
[0009] The atomic descriptor providing unit is adapted to provide atomic descriptors. Specifically, the atomic descriptor providing unit may be a storage unit or may be in communicative contact with a storage unit in which the atomic descriptors are already stored. However, the atomic descriptor providing unit may refer to or be in communicative contact with the atomic descriptor determining unit to determine the atomic descriptors, for example, based on structural information of the molecular substance. Preferably, the atomic descriptors indicate atomic characteristics of the substance with respect to its 3D structure at atomic resolution. Preferably, the atomic descriptors include atomic intrinsic quantities derived from the atomic structural environment of the atoms of the substance and its electronic structure of the molecule.
[0010] The derivation of atomic descriptors from the atomic structure environment can be performed according to any formalism that can convert the three-dimensional structure and elemental composition of a molecular system into a set of per-atom descriptors. These can be end-to-end architectures such as SchNet and PhysNet, but predefined atomic descriptors, such as ANI, SOAP, and FCHL, can also be used. The atomic descriptors derived from molecular electronic structure calculations can refer to one or more of the following: partial charges, exposed surface fraction, average surface screening charge density from continuum solvation model calculations, flexibility, such as average atomic radius calculations in continuum solvation models based on the use of isodense cavities, contributions to the energy or free energy in COSMO-RS- or COSMO-SAC-like solvation models, such as misfit H-bond donating, misfit H-bond accepting, dispersion or total residual contributions, pairwise contact probabilities, COSMO-RS-like or COSMO-SAC-like solvation models, and NMR shielding constants.
[0011] The trained property model providing unit is adapted to provide a trained machine learning-based property model. Specifically, the trained property model is adapted, i.e., trained, to determine a property of a substance as an output when atomic descriptors are provided as input. The determined property may refer to any property of the substance for which the trained machine learning-based property model is trained. Specifically, in a preferred embodiment, the property is a physicochemical property referring to any of melting point, boiling point, vapor pressure, heat of vaporization, heat capacity, flash point, auto-ignition temperature, liquid density, critical temperature and / or pressure, electrical and / or thermal conductivity, glass transition temperature, viscosity, surface tension, refractive index, general mechanical properties, total energy, dipole moment, polarizability, HOMO-LUMO or band gap, ionization potential, electron affinity, activity coefficient, partition coefficient, solubility, cloud point, critical micelle concentration, acid and base dissociation constants, hydrolytic stability, corrosivity, and catalytic activity for a given chemical reaction, or spectroscopic properties, such as transition energies and intensities. Furthermore, in preferred embodiments, application properties may refer to either a) general properties other than purely physicochemical properties, such as various types of toxicological and ecotoxicological behavior, biodegradability in various habitats, odor, ozone depletion potential, global warming potential, etc., and b) performance properties as additives in mixtures, such as effect on the octane and cetane number of fuels, stability against UV radiation or oxidation, anti-corrosion action on the surface of a solution or coating, flame retardancy, effect on tactility, etc.
[0012] In a preferred embodiment, the determined property refers to the melting point of the substance. Determining the melting point of a substance using a property model makes it possible to determine very quickly whether the substance is suitable for agrochemical or pharmaceutical applications, or as a resin raw material for, for example, epoxy or polyurethane systems, i.e., without having to synthesize the substance and perform complex measurement sequences. In particular, for these applications, it is important that the substance does not have an excessively high melting point. Therefore, using a property model, potential substances can be very quickly scanned for their general suitability. Additionally or alternatively, in one embodiment, the property refers to vapor pressure. For example, if the substance refers to a potential compatible plastic, in this way, before synthesizing the plastic additive or preparing a sample for application testing, it can be checked whether the substance's volatility is sufficiently low to ensure that the additive substance remains in the desired material and maintains its effectiveness. This can, for example, increase the safety of the synthesized product. Furthermore, determining the vapor pressure of a substance, together with other properties such as activity coefficients, allows for the identification of suitable conditions for, for example, a separation process, such as distillation. Additionally or alternatively, in one embodiment, the property refers to thermal conductivity. This allows the device to be applied to search for polymeric materials that can function as thermal conductors or insulators, enabling virtual screening of new insulating materials that can be used, for example, to select appropriate materials for electrical devices that avoid overheating. Additionally or alternatively, in one embodiment, the property refers to viscosity. Using the device to determine the viscosity of a potential substance allows, for example, determining whether the viscosity falls within a given range, for example, in the context of reactive resins or polymer processing using an extrusion process followed by injection molding or film blowing. Additionally or alternatively, in one embodiment, the property refers to solubility. Using the device to screen potential substances for solubility can be advantageous in process design, for example, to find the ideal solvent and / or precipitant for a given chemical synthesis or extraction / purification workup.Furthermore, it is possible to ensure that potential active ingredients, i.e., substances, are available for biological uptake or that demixing does not occur in formulated products without having to utilize complex test sequences. Additionally or alternatively, the property refers to ionization potential. By training a property model to determine ionization potential, the device can be applied to screen potential battery materials, such as base materials, solvents, and oxidation-resistant membranes. Furthermore, ionization potential can also be used as a physical descriptor of a molecule that can be correlated with other application properties, so that the determination of ionization potential can be reused in conjunction with another trained property model to determine another property of the substance. Additionally or alternatively, in one embodiment, the property refers to one or more spectroscopic properties, such as transition energy and / or intensity. In this case, the device can also be applied in a validation process to verify that the synthesis steps actually resulted in the desired product. Additionally or alternatively, the property refers to either the flame retardancy of the substance itself as a material or the flame retardancy of a physical mixture of a polymer and a substance that functions as a flame retardant additive when mixed. Determining the flame retardancy of polymers added as additives allows virtual screening to find new flame retardants that are more effective than existing substances in plastic formulations. Furthermore, in this way, inherently flame retardant materials can be searched for very quickly and efficiently.
[0013] The trained property model includes a) an atomic classifier adapted to classify atoms of molecules into one or more atomic classes based on atomic descriptors, and b) a regression model adapted to determine properties based on the atomic classes determined for the substance. The trained property model is a machine learning-based model in which the atomic classifier and the regression model are trained simultaneously. Specifically, the atomic classifier is a machine learning-based classifier. The atomic classifier may refer to any machine learning classifier that provides one or more classes for objects and can be trained to classify objects into respective classes based on provided characteristics of the objects, in this case, based on provided atomic descriptors. Specifically, the atomic classes utilized by the atomic classifier are not predetermined, for example, by user considerations or an expert class hierarchy, but are trained simultaneously with the regression model so that the optimal atomic class for each case, for example, for each property to be determined, is identified by training alone. Thus, the atomic classes of the atomic classifier are determined purely in a data-driven manner. Furthermore, the regression model that utilizes the atomic classes determined for molecules to determine properties may refer to any machine learning-based regression model. Generally, a machine learning regression model refers to any machine learning model that assigns a continuous dependent variable to a set of input variables; in this application, the continuous dependent variable refers to a set of input variables to a property and the value of each atom class. Preferably, the regression model refers to any one of multiple linear regression, artificial neural networks, Gaussian process regression, kernel algorithms, support vector regression, and random forest algorithms. There are many advantages to using a classifier that can learn atom classes based on training data and the following regression model. Specifically, the atom classes in this case provide a description of the characteristics and influence of each atom of a substance on the property learned from the training data, and are therefore specific to each application and do not rely on known, convenient classifications. This allows for the provision of unique information already as part of the model, potentially reducing the amount of training data, for example by reducing the dimensionality of the information provided by the training data.Furthermore, it becomes possible to train models to utilize more meaningful descriptors, which also makes it possible to interpret the training results, i.e., the feature models. Furthermore, the inventors have found that the feature models described above are more robust to artifacts that may be caused by biased training data.
[0014] The property determination unit is then adapted to determine properties of the substance by applying the trained property model to the atomic descriptors. Specifically, the atomic descriptors are provided as input to the trained property model, which then provides determined properties for the substance as output. The determined properties may then be provided for further processing, for example, by the property providing unit. For example, further processing may refer to providing the determined properties to a user as output, for example, via a screen or any other suitable method. However, the determined properties may also be provided for further processing related to a specific application, for example, internally in a screening process for screening multiple substances, in which case each property is not necessarily provided to a user but is used only to determine, for example, a list of potential substances suitable for the application, and further processed, for example, by an automatic synthesis unit for automatically synthesizing the substances on the list. Furthermore, the determined properties may be further processed by being provided to a database and then stored, for example, in a database search operation, before further processing.
[0015] Preferably, processing the determined property includes determining a control signal for controlling and / or monitoring a production process based on the determined property. The production process may refer to a production process for a substance, or may refer to a production process for a product in which the substance is utilized. For example, if the determined property indicates that the substance has a particular vapor pressure, generating the control signal may include generating a control signal for controlling and / or monitoring a production process for a polymer in which the substance is utilized and taking the vapor pressure of the substance into account. In a preferred embodiment, the control signal indicates a machine-viable formulation of the substance, particularly if the comparison indicates that the determined property of the substance is within a predetermined range around the provided target property. Further, the method may include controlling and / or monitoring the production process based on the control signal.
[0016] Furthermore, the process of processing properties may also refer to selecting one or more substances based on their respective determined properties. For example, if properties have been determined for multiple potential substances, selecting may include comparing the determined properties of different substances with predetermined selection criteria and selecting substances whose determined properties satisfy these criteria. Specifically, in one embodiment, the method includes receiving target properties of substances, comparing the received target properties with the determined properties, and providing a control signal in response to the comparison. The control signal may refer to any signal that enables further control of a technical system. For example, the control signal may be adapted to control an interface to provide the results of the comparison on an interface. In a preferred embodiment, the comparison refers to verification of the target properties, and the verification is positive if the determined properties are within a predetermined range centered on the target properties. In this case, the control signal may simply be adapted to control a user interface to provide an indication of a positive or negative verification result. Preferably, however, the control signal refers to a recipe, i.e., a formulation, of one or more substances that satisfy a particular target property, i.e., that have been verified positively. A recipe is generally defined as instructions on how a substance can be synthesized or formulated. In particular, the recipe includes starting materials and respective parameters for producing the material from the starting materials. Preferably, the control signal includes the recipe in a form that directly enables automatic control of a respective industrial system or piece of work equipment to produce the material. In particular, if the result of the comparison indicates that the determined property is within a predetermined range around the target property, the control signal preferably indicates a machine-executable recipe for the material.
[0017] In one embodiment, a molecular substance is considered in more detail as an ensemble of several conformers. In this case, the atomic descriptor of a molecule may refer to a weighted average of the atomic descriptors of different conformers of a substance. In general, conformation refers to the phenomenon of conformational isomerism, in which isomers can be interconverted not only by rotation around chemical bonds, which are primarily formal single bonds, but also by other types of interconversion, such as ring inversion or pseudorotation in cyclic non-aromatic hydrocarbons. All of these potential interconversions are subject to the criterion that they are kinetically possible, i.e., that they actually occur at the temperature of interest and within a reasonable time scale. In this case, atoms can exist in a substance in different environments, which may result in different atomic descriptors for different conformers. Here, the accuracy of the determination can be improved by utilizing an atomic descriptor that refers to a weighted average of the atomic descriptors of different conformers. Specifically, the weight may indicate the probability that a particular conformer of a molecule exists in a substance. For example, if two possible conformers are expected to exist in a substance with the same probability, the weight may indicate the proportion of the conformers expected in the substance. Preferably, the weight of the weighted average refers to the Boltzmann weight. The conformational energy or free energy required to determine the Boltzmann weight can be obtained from quantum chemistry-based calculations, semi-empirical calculations, or machine learning calculations. Furthermore, the weight may also refer to a trained weight that is trained together with the property model based on the same training data used to train the property model.
[0018] In one embodiment, the atomic classifier is provided with predetermined constraints for classifying atoms of a substance based on atomic descriptors. Providing predetermined constraints on classification can further increase the efficiency of classification, for example by constraining classification to chemically sensitive classes. Preferably, in one embodiment, the atomic classifier is adapted to classify atoms of a substance into one or more atomic classes in an element-specific manner, such that atoms referring to different elements are classified into different atomic classes. Providing this constraint for training the atomic classifier makes it possible to directly take into account that different elements often have different chemical characteristics that affect properties differently. Thus, this constraint can make it possible to further reduce the amount of training data required to train a property model. However, if a more flexible property model is desired, the atomic classifier can also be trained without the element-specific restrictions on atom classes established by the atomic classifier during training.
[0019] In one embodiment, the atomic classifier is adapted to classify atoms of a substance into one or more atomic classes. When an atom is classified into two or more atomic classes, an atomic class weight is provided for each atomic class to which the atom is partially assigned based on the conformance of a given atomic property to the class definition. In some cases, an atom may contain features, i.e., atomic descriptors, that result in the atom being assigned to multiple atomic classes. In this case, the atom may then be considered to affect each property according to both atomic classes. However, in this case, only one atom has the effect associated with each atomic class, and therefore an atomic class weight is provided to prevent this one atom from having the same influence on a property as a molecule containing one atom in each of the atomic classes. Specifically, the atomic class weight is assigned based on the conformance of a given atomic property to the class definition for each property. However, in one embodiment, the classifier may also be adapted to only allow classification of atoms in one atomic class. This leads to less flexibility in the training process, but allows for further reduction in training data.
[0020] In a further aspect of the present invention, a training apparatus for training a machine learning-based property model is provided, the apparatus comprising: i) a training data providing unit for providing training data for training the property model, the training data respectively comprising atomic descriptors and corresponding known properties of a plurality of different substances; ii) a trainable property model providing unit for providing a machine learning-based property model, the property model comprising: a) an atomic classifier trainable to classify atoms of the substance into one or more atomic classes based on the atomic descriptors; and b) a regression model trainable to determine properties based on the atomic classes determined for the molecule; and iii) a training unit for training the provided property model based on the training data, such that the trained property model is adapted to determine properties of the substance as output when atomic descriptors of the substance are provided as input.
[0021] Generally, training data including atomic descriptors and corresponding known properties of multiple different substances can be provided in the form of a database in which features and properties or substances are stored. Such a database can be based, for example, on simulation data, measurement data, sequence data, etc. Preferably, the training data refers to only one property of the substances so that a property model trained based on this training data is specialized to determine each property. However, two or more properties can also be provided in the training data so that the training property model can determine two or more properties for each substance. Generally, any known training algorithm for training data-driven, particularly machine learning-based, models can be utilized. Preferably, during the training of the property model, the descriptors of the substances that have the greatest influence on the properties are also determined, and the model is then trained based on these most influential descriptors. To determine these most influential descriptors, for example, cluster analysis or PCA analysis tools can be utilized. Specifically, the descriptors can be used to determine the application space of the training data, which is then defined by the descriptors of the substances covered by the data. The determination of the most influential descriptors can then be performed as a dimensionality reduction of the application space. An algorithm can then be applied to optimize the training data in the application space, for example, to cover the application space with as little training data as possible.
[0022] Preferably, training the feature model includes iteratively optimizing model parameters of the provided feature model based on the training data until the trained feature model provides a predetermined accuracy improvement. Specifically, the model parameters include classification parameters and regression parameters, and the iterative optimization of the model parameters includes optimizing the classification parameters of the atomic parameters and the regression parameters of the regression model for each iteration step. In general, the accuracy improvement may refer to the difference between the result of a current iteration step and the result of a previous iteration step, as is typical in iterative methods. When the iterations converge, as they usually should, the difference between the results of each iteration step becomes smaller with each step, and the iterations may be stopped when this difference, i.e., accuracy improvement, falls below a predetermined accuracy improvement, i.e., a predetermined threshold. Possible variations of this procedure include maintaining certain parameters for several iteration steps and then adapting them again, or applying a convergence criterion that takes into account the evolution of accuracy over several iteration steps.
[0023] In a further aspect, a system for screening potential substances for a predetermined application is presented, the system comprising: a) the apparatus described above; b) a target property providing unit for providing target substances and target properties of potential target substances; and c) a screening unit adapted to determine a property of each of the potential target substances using the apparatus, compare the determined property of the potential target substance with the target property, and, based on the comparison, either: i) determine the potential target substance as the target substance; or ii) provide a new potential target substance and repeat the property determination using the new potential target substance. Preferably, the method further comprises generating a control signal indicative of a machine-executable formulation of the target substance for controlling and / or monitoring production of the target substance. Preferably, the substance is a polymer, and the machine-executable formulation comprises a synthetic specification for the polymer. Further, the method may comprise controlling and / or monitoring production of the target substance based on the control signal.
[0024] The target property providing unit is configured to provide a target property that indicates a characteristic of a substance. Specifically, providing may refer to receiving the target property from a user's input, for example, using the respective input units. Furthermore, providing may also refer to accessing a storage unit in which the target property is already stored. Furthermore, providing may include receiving the target property, for example, from another source via a network connection, and providing the received property. Generally, the target property may refer to a single target value, for example, a target vapor pressure, or may refer to a range of values to be met by the substance. Furthermore, the target property may also refer to any kind of target function, for example, a timed sequence of properties. Furthermore, the target property providing unit is configured to provide potential target substances. Specifically, providing may refer to receiving the potential target substances from a user's input, for example, using the respective input units. Furthermore, providing may also refer to accessing a storage unit in which the potential target substances are already stored. Furthermore, providing may include receiving the potential target substances, for example, from another source via a network connection, and providing the received potential target substances. The potential target substances may then be considered as starting materials for screening for substances that satisfy the target property. The respective descriptors for the substances can be determined based on the provided potential target substances, for example as already described above. In general, the descriptors can be determined by accessing a storage unit in which descriptors for the respective substances are already stored. Also, any of the above-described methods for determining descriptors based on the provided substances can be used.
[0025] After utilizing the property model to determine properties for the potential target material, the determined properties are compared with the target properties. Based on the comparison, it is determined whether the potential target material has been determined as a target material, and if so, the iteration may stop at this point. Furthermore, based on the comparison, it may be determined to provide a new potential target material and repeat the property determination using the new potential target material. Thus, at this point, an iteration is performed in which property determination is repeated using the property model and descriptors of the potential target material until one of the potential target materials is determined as a target material. Specifically, the comparison may include determining whether the determined properties of the potential target material are within a predetermined range around the target property; if so, the target may be considered satisfied, and the potential target material is determined as a target material. If the determined properties are outside the predetermined range around the target property, the target is determined to be unsatisfied, and a new potential target material is provided that may satisfy the target property.
[0026] In general, the iterations performed can refer to any search or directed search of the potential target substance space. For example, new potential target substances can simply be randomly selected from a vast number of potential target substances generated in silico. However, specific rules for generating new potential target substances can also be applied based on a comparison between the determined properties and the target properties, with or without consideration of simultaneous optimization of additional target properties of the substance. In general, known methods for generating new target substances can be utilized, for example, evolutionary algorithms or Bayesian optimizers can be used. For example, such algorithms can be used to change the molecular structure or periodic atomic structure of a substance to generate new potential target substances.
[0027] Then, iterations may be performed over steps of determining properties of the new potential target substance by utilizing the descriptors of the new potential target substance, as described above. Optionally, determining descriptors from the potential target substance may also be part of the iterations if the descriptors are not already provided or present in the storage unit. Furthermore, it is preferable that the same property model is used in all iteration steps for determining properties. However, in some cases, different property models may be used in different iteration steps. For example, if other descriptors are utilized for the new potential target substance, another property model may also be more appropriate. This may be the case, for example, if some descriptors are not available or not available with sufficient accuracy for the new potential target substance.
[0028] After the iterations have stopped, e.g., after a potential target substance has been determined as a target substance, or if a new potential target substance cannot be selected or generated, the results of the iterations may be provided to the user. For example, if none of the possible potential target substances meet the target properties, the user may be notified of the failure to determine a target substance. If a target substance can be determined, the target substance may be provided to the user as output. For example, the determined target substance may then be provided to an output unit or a computing unit for further processing. Preferably, the provision of the target substance leads to further processing, e.g., generating a control signal, as already described above.
[0029] In a further aspect of the present invention, a method is provided for determining properties of a material defined as either indicative of molecular nature or representing an extended periodic 1D, 2D or 3D structure, the method comprising: i) providing atomic descriptors that characterize the atoms of the material with respect to the material's structure; ii) providing a trained machine learning-based property model, the trained property model adapted to determine properties of the material as output when the atomic descriptors are provided as input, the property model comprising: a) an atomic classifier adapted to classify atoms of the molecule into one or more atomic classes based on the atomic descriptors; and b) a regression model adapted to determine properties based on the atom classes determined for the material; and iii) determining properties of the material by applying the trained property model to the atomic descriptors.
[0030] In a further aspect of the present invention, a training method for training a machine learning-based property model is provided, the method including: i) providing training data for training the property model, the training data including atomic descriptors and corresponding known properties of a plurality of different substances, respectively; ii) providing a machine learning-based property model, the property model including: a) an atomic classifier trainable to classify atoms of the substance into one or more atomic classes based on the atomic descriptors; and b) a regression model trainable to determine properties based on the atomic classes determined for the molecules; and iii) training the provided property model based on the training data, such that when atomic descriptors of a particular molecule are provided as input, the trained property model is adapted to determine as output properties of the substance defined as either indicative of molecular properties or representing extended periodic 1D, 2D, or 3D structures.
[0031] In a further aspect, a method for screening potential substances for a given application is provided, the method comprising: a) providing a target substance and a target property for a potential target substance; and b) determining a property for each of the potential target substances using the method described above, comparing the determined property of the potential target substance with the target property, and based on the comparison, either i) determining the potential target substance as the target substance, or ii) providing a new potential target substance and repeating the property determination using the new potential target substance. In a further aspect of the present invention, a computer program product for training a machine learning-based feature model is provided, the computer program product comprising program code means for causing a feature model training apparatus as described above to perform the training method as described above.
[0032] In a further aspect of the present invention there is provided a computer program product for determining a property of a substance, the computer program product comprising program code means for causing an apparatus as described above to carry out a method as described above.
[0033] It is to be understood that the above-mentioned devices, the above-mentioned methods and the above-mentioned computer program products have similar and / or identical preferred embodiments, in particular as defined in the respective dependent claims.
[0034] It is to be understood that a preferred embodiment of the invention can also be any combination of the dependent claims or the above-mentioned embodiments with the respective independent claim.
[0035] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. [Brief explanation of the drawings]
[0036] [Figure 1] 1 shows, in a schematic and exemplary manner, an embodiment of an apparatus for determining a property of a substance; [Figure 2]1 illustrates, in a schematic and exemplary manner, one embodiment of a method for determining a property of a substance. [Figure 3] 1 illustrates, schematically and exemplarily, an embodiment of an apparatus for training a feature model. [Figure 4] 1 shows, schematically and exemplarily, a flow chart of a method for training a feature model. [Figure 5] 1 illustrates a schematic diagram of an exemplary application of the present invention; [Figure 6] 1 illustrates a schematic diagram of an exemplary application of the present invention; [Figure 7] 1 shows a schematic and exemplary application of the present invention for formulation optimization. [Figure 8] 1 shows, in a schematic and exemplary manner, a method for refining the determination of the respective property model and / or for screening new substances for application. DETAILED DESCRIPTION OF THE INVENTION
[0037] Detailed Description of the Embodiments FIG. 1 schematically and exemplarily illustrates one embodiment of an apparatus 100 for determining properties of a material, defined as either exhibiting molecular properties or representing an extended periodic 1D, 2D, or 3D structure. The apparatus 100 comprises an atomic descriptor providing unit 110, a trained property model providing unit 120, and a property determination unit 130. Furthermore, the apparatus may also comprise property providing units not shown in FIG. 1 . Specifically, the apparatus 100 may be implemented in any form of software and / or hardware that can be dedicated to the task of determining properties of a material or that can provide additional functionalities. The apparatus 100 may further comprise an input unit 141, such as a keyboard, a mouse, a touchscreen, etc., and / or an output unit 142 for outputting information to a user as audio output, e.g., on a display, by utilizing a lighting unit, etc. However, the input unit 141 and the output unit 142 may also be omitted and the apparatus 100 may, for example, be provided as part of the software and / or hardware of a network of computing devices, for example of a cloud, in which case the apparatus may be connected to other computing devices which may act as input or output units or which may provide information to the apparatus, for example in the form of stored information, and / or which may receive information from the apparatus, for example the result of a decision, which may then be further processed without explicitly providing it as output to the user.
[0038] The atomic descriptor providing unit 110 is adapted to provide atomic descriptors, e.g., stored in a respective database of atomic descriptors or predetermined from respective knowledge of a particular molecule. Generally, the atomic descriptors indicate characteristics of atoms of a substance with respect to the structure of the substance. For example, in a preferred embodiment, the atomic descriptors include atomic intrinsic quantities derived from the molecular electronic structure of the substance and / or from atomic features indicative of the atomic structural environment of the atoms of the substance. The atomic descriptor providing unit 110 is then adapted to provide the atomic descriptors, e.g., to the characterization unit 130.
[0039] The apparatus 100 further includes a trained property model providing unit 120 adapted to provide a trained machine learning-based property model. The trained property model is adapted, i.e., trained, to determine a property of a substance as an output when atomic descriptors are provided as input. Specifically, the property of a substance determined by the trained property model may refer to a physicochemical property and / or an application property. Physicochemical properties refer to physical or chemical characteristics of a substance, such as melting point, viscosity, etc., defined as either indicative of molecular properties or representing an extended periodic 1D, 2D, or 3D structure. Application properties refer to application-specific properties of a substance, such as biodegradability, toxicity, flash point, etc., defined as either indicative of molecular properties or representing an extended periodic 1D, 2D, or 3D structure. In general, some properties may be considered to belong to both physicochemical properties and application-specific properties, depending, for example, on the respective intended applications. In general, the trained property model may be trained using, for example, a training apparatus and a training method as described with reference to FIGS. 3 and 4 .
[0040] Generally, the trained feature model includes an atomic classifier adapted to classify atoms of molecules into one or more atomic classes based on atomic descriptors, and a regression model adapted to determine features based on the atomic classes determined for the molecules. Specifically, the atomic classifier may refer to any machine learning classifier that can provide one or more classes for objects and be trained to classify the objects into the respective classes based on provided features of the objects, in this case, based on provided atomic descriptors. The atom classes used by the atomic classifier are not predetermined, for example, by user considerations or an expert class hierarchy, but are trained simultaneously with the regression model so that the optimal atom class for each feature to be determined is identified by training alone. Thus, the atom classes of the atomic classifier are determined purely in a data-driven manner. Furthermore, the regression model that utilizes the atomic classes determined for molecules to determine features may refer to any machine learning-based regression model. Preferably, the regression model refers to any one of linear regression, multiple linear regression, artificial neural network, Gaussian process regression, kernel algorithm, support vector regression, and random forest algorithm.
[0041] Preferred examples of classifiers and regression models are described in more detail below. Preferably, the utilized classifier assigns each atom of a material a probability of belonging to a particular atomic group, e.g., in the form of an atom class coefficient. During training, this can be achieved by first applying multiple linear regression to each atomic descriptor vector to generate non-normalized atom class coefficients. This can then be followed by normalization via a softmax function. The atomic descriptor vectors can be used as input to the classifier either directly or after additional transformations, e.g., using an artificial neural network. Alternative normalization schemes to softmax can be used, provided the optimization function of the property model is differentiable and the final atom class coefficients are non-negative and sum to 1 for each atom. One example is to square each value in the non-normalized atom class coefficient vector and divide by the sum of the squared values. The normalized atom class assignments can then be used in a regression model to determine the final properties and impose additional constraints on the training procedure. Examples include orthogonality constraints that encourage the assignment of chemically different atoms to different groups, and entropy-based constraints that limit the number of groups clustered together during training.
[0042] The trained property model providing unit 120 is then adapted to provide the trained property model, for example, to the property determination unit 130. The property determination unit 130 may then be adapted to determine properties of the substance by applying the trained property model to the atomic descriptors, i.e., by providing the atomic descriptors of the substance as input to the trained property model. The output of the trained property model then refers to or indicates the respective properties for which the trained property model was trained. In general, the property determination unit 130 may also be adapted to apply two or more property models to the atomic descriptors, e.g., apply different property models trained for different respective properties, such that the property determination unit 130 is adapted to determine two or more properties of the substance. The different trained property models may then be applied subsequently to or in parallel with the atomic descriptors. Specifically, in some embodiments, the property determination unit may be adapted to utilize one or more results, i.e., the properties of the substance determined by the trained property model, again as atomic descriptors, and / or to determine further atomic descriptors that are then provided again to the respective trained property models to yield further properties of the substance.
[0043] Optionally, the apparatus 10 may further include a characteristic providing unit for providing the characteristics for further processing. For example, the characteristic providing unit may refer to an output unit 142 that processes the characteristics and can provide the characteristics to a user, for example, via a display. However, the characteristic providing unit may also be configured as an interface for interfacing with another system, for example, a production process control system. In this case, the further processing may refer to generating control data for controlling the production process control system based on the determined characteristics. For example, the production process may be configured to produce a material if the determined characteristics meet target characteristics, or to set one or more process parameters of the production process based on the determined characteristics, for example, when the material is utilized during production and each determined characteristic affects a production parameter.
[0044] FIG. 2 schematically and exemplarily illustrates a method for determining a property of a material. Method 200 includes step 210 of providing atomic descriptors, e.g., as described above. Furthermore, method 200 includes step 220 of providing a trained machine-learning-based property model. Specifically, the trained machine-learning-based property model may refer to a property model as described above, e.g., trained using the apparatus and method as described with reference to FIGS. 3 and 4. Generally, step 210 of providing atomic descriptors and step 220 of providing a trained machine-learning-based property model can be performed in any order or even simultaneously. In the final step 230, method 200 then includes determining a property of the material by applying the provided trained property model to the atomic descriptors.
[0045] 3 shows a schematic and exemplary training apparatus 300 for training a machine learning-based feature model that can be used in the apparatus and method as described above. The training apparatus 300 includes a training data providing unit 310, a trainable feature model providing unit 320, and a training unit 330.
[0046] The training data providing unit 310 is adapted to provide training data for training the feature model. Specifically, the training data includes atomic descriptors of a plurality of different substances, each containing a different substance, and corresponding known properties. For example, the training data providing unit 310 may be connected to a database 340 that stores a plurality of such data sets including atomic descriptors of molecules and corresponding properties. Based on the desired training of the feature model, for example, based on the desired output of one or more respective properties, the training data providing unit 310 may also be adapted to select and provide respective training data from the database 340. However, data from such database 340 may also be selected by a user or selected by the training data providing unit 310 based on specific criteria provided by the user.
[0047] The trainable feature model providing unit 120 is then adapted to provide a trainable machine learning-based feature model. Specifically, the structure of the trainable machine learning-based feature model may follow the structure of the trained feature model as described above. In general, a trainable feature model can be considered as referring to an adaptable feature model, where parameters determining the processing of inputs in the feature model to generate the feature model's output are used. For example, the parameters can initially be set to predetermined values, which are then adapted based on training data during training of the feature model so that the feature model can subsequently perform its function. However, a trainable feature model may also refer to an already trained feature model, e.g., a feature model already trained using a respective training dataset. In some cases, it may be desirable to retrain such an already trained feature model, e.g., to improve accuracy or to extend the applicability of the feature model, e.g., to different materials or additional properties. In this case, retraining the feature model may refer to providing a new training dataset to the feature model to further adapt the trainable feature model's parameters.
[0048] The training unit 330 is then adapted to train the provided feature model based on the training data and the provided trainable feature model. Preferably, training the feature model refers to iterative training in which model parameters of the provided feature model are optimized for their respective tasks based on the training data until the trained feature model provides a predetermined improvement in accuracy. Specifically, thresholds can be set that allow determining when the training of the feature model has converged to a state where no further improvement beyond the respective threshold is expected with additional training. Furthermore, it is preferred that the model parameters, which can be divided into classifier parameters and regression parameters, are trained simultaneously, e.g., optimized together in each iteration step.
[0049] FIG. 4 schematically and exemplarily illustrates a flowchart of a method for training a machine learning-based feature model. The method 400 includes a step 410 of providing training data for feature training, e.g., according to the principles described above. Furthermore, the training method 400 includes a step 420 of providing a trainable machine learning-based feature model. The trainable machine learning-based feature model preferably has a structure as already described above. Generally, the step 410 of providing training data and the step 420 of providing a trainable machine learning-based feature model can be performed in any order, or even simultaneously. Furthermore, the method 400 includes a step 430 of training the provided feature model based on the training data, such that the trained feature model is adapted to determine, as output, a property of a substance consisting of a specific molecule when atomic descriptors of the specific molecule are provided as input. Generally, the training can be performed according to any of the principles already described above.
[0050] 5 and 6 illustrate an exemplary application of the above-described apparatus for reducing the energy effort associated with extractive distillation as a method for separating fluid mixtures. The quality of chemical production processes is often characterized by the purity of the product, which must be achieved by the purity of the extract and the separation of by-products generated in the reaction process. When conventional distillation processes are limited by a small separation factor between components, e.g., between the product and similar by-components, additives can be provided to increase the separation factor. Screening for such additives using the above-described apparatus can result in technically and economically viable distillation processes, particularly with extractive distillation, with significantly reduced energy effort.
[0051] To screen for optimal additives for extractive distillation, the following thermodynamic data are preferably used: a) the pure component vapor pressures of the components to be separated and the pure component vapor pressures of the additive; and b) activity coefficients, or, as a first screening approximation, the activity coefficients at infinite dilution of the components to be separated when the additive is used as the solvent. Selectivity can be defined as the ratio of the activity coefficients of the components for separation at infinite dilution in the additive solvent. Selectivity, combined with the pure component vapor pressure data, is directly related to the separation factor and the effort, i.e., investment and energy consumption of the distillation stage, of the extractive distillation using the additive. Furthermore, to avoid two-phase partitioning of the mixture in the column, the reciprocal values of the activity coefficients of the components in the additive can be applied as a first approximation on a volumetric basis.
[0052] The above-described apparatus can be utilized to determine the pure component vapor pressures and activity coefficients of different potential additives, thus allowing for the preselection of potentially advantageous additives for a wide variety of components for optimal distillation and separation processes. For example, the following equation can be utilized to optimize for the maximum relative volatility difference of the desired product relative to the other components, utilizing the determined vapor pressure and infinite dilution activity coefficient for the pure product, respectively:
[0053]
number
[0054] In this formula,
[0055]
number
[0056] and
[0057]
number
[0058] are the infinite dilution activity coefficients of the products and constituents, respectively,
[0059]
number
[0060] and
[0061]
number
[0062] is the determined vapor pressure of each pure product. The vapor pressure and infinite dilution activity coefficient of each pure product can be determined as properties of each material using the apparatus and method described above, and an iterative optimization procedure can be used to determine potential additives that maximize relative volatility.
[0063] The following further example illustrates the application of the above-described apparatus and method to determining the temperature-dependent vapor pressure of a pure substance, for example, for the aforementioned application. The model in this application preferably consists of a SchNet model to determine appropriate atomic descriptors, followed by linear regression and softmax activation as a classifier to classify atomic groups, preferably allowing up to 50 different groups. The frequency of the groups in the substance can then be used to determine the base 10 logarithm of the vapor pressure by linear regression, followed by multiplication with a coefficient f that depends on the target temperature T and the reference normal boiling point TNBP of the substance. For example, the relationship f = (T / TNBP - 1) / (T / TNBP - 0.125) can be used. A training data set can be determined based on the vapor pressure curves and normal boiling points obtained from the DIPPR data set. In this example, the substances are constrained to contain only the elements C, H, N, and O, resulting in a set of 1016 substances. For each substance, the vapor pressure can be evaluated at 20 different temperatures within the limits specified in the DIPPR database, resulting in a total of 20,580 data points. Of these, 80% can be used for training, 10% for validation, and 10% for testing. The split can be performed randomly, with the condition that data belonging to the same compound can only occur in one of the subsets. The training procedure can utilize the Huber loss to the base 10 logarithm of the determined true pure substance vapor pressure. The density function theory-optimized structure, as well as the experimental temperature and reference normal boiling point, can serve as inputs to the regression model. The loss can be minimized using mini-batch stochastic gradient descent with the ADAM algorithm. Training can be monitored by an early stopping procedure that operates on the validation set and stops when the prediction error improvement remains below a predefined threshold. The best model can then be selected based on the lowest validation error achieved during training and its prediction error evaluated using the test set.The model generated by this procedure provided a prediction error of 21.25% for the pure substance vapor pressure of previously unseen substances, which is sufficiently accurate to screen each large set of novel substances for potential target substances and can be further increased using the screening procedure described in Figure 8.
[0064] FIG. 7 illustrates a schematic and exemplary application of the present invention for optimizing a formulation. In this embodiment, a target formulation that satisfies one or more predetermined target properties is provided. The formulation includes at least two, preferably multiple, components, which may be any substance according to the present invention. In most cases, one or more of the components of the target formulation are fixed, for example, due to the respective application requirements. However, in this example, at least one variable component exists within the predetermined application limits. Specifically, a list of multiple potential substances that can be used as additional components can be provided, optionally along with the respective amounts of each potential substance as an additional component. The above-described apparatus and method can then be utilized to determine one or more properties of a formulation that includes a first potential substance as an additional component. In particular, the properties of the formulation's substances can be determined, and the relative amounts of each substance can be utilized to determine each property of the formulation. The determined one or more properties can then be compared with the respective target properties, i.e., the respective values of the properties are compared with the respective values of the target properties. Based on the comparison, it can then be determined whether the formulation satisfies the predetermined target, i.e., whether it satisfies one or more target properties. If the formulation satisfies the target, the current formulation can be determined as the target formulation. If the formulation does not meet the target, a new formulation may be provided by modifying the amounts of additional ingredients and / or current ingredients according to predetermined limits. Specifically, respective rules may be utilized to automatically generate a new formulation based on a previous formulation. For example, a rule may determine that first the amount of an additional ingredient is increased or decreased by a predetermined increment until it reaches a respective predetermined limit, and then that ingredient is replaced by a new ingredient from the list of potential ingredients starting from a predetermined starting amount. However, more sophisticated rules may also be utilized.
[0065] FIG. 8 illustrates a schematic and exemplary method for refining the determination of each property model and / or screening new substances for an application. In this example, a determination of each application, new substance, and therefore, property model, i.e., each target property, is first provided. The intended application also determines a respective chemical space for potential target substances. For example, in a human care application, the chemical space must exclude potentially harmful substances. However, the chemical space may also be defined based on other considerations, such as substance availability, interactions with other ingredients in the formulation, etc. Based on the application space, multiple potential new substances covering the application space can be determined. Then, using the above-described apparatus and method, the properties of each of these potential new substances can be determined very quickly to screen the application space. In particular, atomic descriptors are determined for each of these potential new substances, and trained classifiers and regression algorithms are applied to determine the properties accordingly. As a result, each property can be compared with the target property, and potential new substances that satisfy the target property or satisfy the target property within predetermined limits, which may initially be set broadly, are selected. Additionally or alternatively, some of the potential new substances in the application space can be selected based on other criteria, e.g., arbitrarily or to cover at least a predetermined extent of the application space. The selected substances are then synthesized. This small number of substances can then be subjected to measurements of their respective application tests or atomic descriptors. For example, the application properties and / or atomic descriptors of these substances can be measured or derived from the respective measurements. These measurement data can then be used to update, and in particular, retrain, the property models. The updated property models can then be used again to determine the properties of new substances in the application space. This results in higher accuracy of each property model, particularly for the region of the application space in which the measured substances were selected.Thus, when the above steps are applied iteratively, for example by screening in subsequent steps only potential new substances in the regions of the application space that come closest to satisfying the target property in the previous step, the property model becomes increasingly accurate in regions where new target substances that satisfy the target property are expected.
[0066] Further embodiments and principles of the present invention are described below. Generally, a property model, such as those described above, is adapted to learn the assignment of atoms to different atomic classes based on training data including material properties and atomic descriptors indicative of the atomic structural environment and / or atomic descriptors derived from electronic structure calculations. The appropriate atomic classes and atomic descriptors for a particular property determination are determined in a purely data-driven manner as part of the overall regression problem. The number and combination of atomic classes present in a material, which may be of molecular nature or may be an extended periodic material, are then used in the regression model to determine the property.
[0067] To improve the generally high predictive power of this approach, high-quality training data and appropriate structural and computational descriptors can be used. In this case, atomic classes can be further trained with respect to properties of interest. Apart from models for direct property determination, the invention described above can also be applied to delta learning using a) physical / engineered material characterization models, and b) the presented regression models for identifying and correcting systematic errors in the physical models.
[0068] For example, a preferred embodiment for training a property model using the training apparatus and method described above is described below. Preferably, the training data for setting up the property model includes a) physicochemical and / or application data for properties that depend on the chemistry of the respective substances, and b) atomic descriptors, i.e., features, derived from the atomic structural environment within these substances and / or obtained from electronic structure calculations of the respective substances.
[0069] The physicochemical or application properties on which the property model can be trained may be generally valid or may depend on external conditions, such as temperature, pressure, the composition of the mixture, etc. For example, the properties on which the property model can be trained refer to physicochemical properties including one or more of the following: melting point, boiling point, vapor pressure, heat of vaporization, heat capacity, flash point and autoignition temperature, liquid density, critical temperature and pressure, electrical and thermal conductivity, glass transition temperature, viscosity, surface tension, refractive index, mechanical properties, total energy, dipole moment, polarizability, HOMO-LUMO gap or band gap, ionization potential, electron affinity, activity coefficient, partition coefficient, e.g., solubility in all types of equilibria between solid, liquid and gas, cloud point, critical micelle concentration, acid and base dissociation constants, hydrolytic stability, corrosivity, catalytic activity for a given chemical reaction, or spectroscopic properties, e.g., transition energies and intensities in IR and UV-Vis spectra. Furthermore, properties may refer to a) general properties other than purely physicochemical properties, such as various types of toxicological and ecotoxicological behavior, biodegradability in various habitats, odor, ozone depletion potential, global warming potential, and / or b) application properties referring to behavior as an additive in a mixture, such as influence on the octane and cetane number of the fuel, stability against UV radiation or oxidation, flame retardancy, effect on the sense of touch.
[0070] The derivation of atomic descriptors from the atomic structure environment can be performed according to any format that can convert the three-dimensional structure and elemental composition of a molecular system into a set of per-atom descriptors. For example, end-to-end architectures such as SchNet and PhysNet can also be trained to learn to determine an appropriate set of atomic descriptors directly from structural data and thus can adapt to a wide range of different chemical systems, provided sufficient data is available. Alternatively, predefined atomic descriptors such as ANI, SOAP, and FCHL can be used, which will increase accuracy for small datasets. Multiple conformers of a molecule may be relevant to the property. To account for this, atomic descriptors can be derived from multiple conformers and combined using different methods, such as weighted averaging, where the weights can be Boltzmann weights of the conformers or weights learned during the property model training procedure.
[0071] Atomic descriptors, which refer to atomic intrinsic quantities, are derived from molecular electronic structure calculations and may refer to one or more of the following: partial charges, exposed surface fraction, average surface screening charge density from continuum solvation model calculations, flexibility, e.g., average atomic radius calculations in continuum solvation models based on the use of isodense cavities, contributions to the energy or free energy in COSMO-RS- or COSMO-SAC-like solvation models, e.g., misfit H-bond donating, misfit H-bond accepting, dispersion or total residual contributions, pairwise contact probabilities, COSMO-RS-like or COSMO-SAC-like solvation models, and NMR shielding constants. The same methods described above for considering conformers can also be applied to these atomic descriptors.
[0072] The trainable feature model is then trained based on the atomic descriptors of the molecules in the training data, for example as described above, to adapt the atoms of the molecules to be assigned to different atomic classes, which is then used to train the regression provided by the regression model. In general, apart from simple multiple linear regression, more complex strategies such as artificial neural networks (ANNs), kernel methods, support vector regression, or random forests can also be used by the regression model. During training, the overall regression problem provided by the training data applied to the feature model can be optimized, preferably in an iterative manner, during which the optimal parameters / coefficients and atom classes are simultaneously determined for the feature model. The selection of the atomic descriptors actually used can also be optimized during training.
[0073] The property model can be adapted to perform atom class assignment through a classification layer such as softmax. This step provides considerable flexibility; for example, the classifier can be adapted to introduce orthogonality during optimization. Furthermore, the classifier can be constrained to select atom classes to be element-specific, i.e., N and O atoms never share the same atom class. Without this restriction, the classifier is more flexible and has the advantage of working with fewer atom classes overall, further improving transferability, i.e., obtaining a model that works for elements for which little or even no data is available in the training set. Furthermore, the classifier can be adapted so that atoms can belong entirely to one atom class or can also have several assignments. In the latter case, the importance of different assignments can be mapped by atom class weights. For example, the sum of the weights can then be forced to sum to 1, although this is not required. In general, the greater flexibility in the latter case can be used to account for the different and powerful effects of potential nonlinearities of different atoms on a given property.
[0074] As previously mentioned, for the regression model that determines the target property, a simple multiple linear regression may be used, or more complex strategies such as artificial neural networks (ANNs), kernel methods, support vector regression, or random forests may be utilized. For example, the training procedure for the property model, performed by the training unit of the training device, results in both assigned atom classes and determined values of the target property. Generally, the iterative training is terminated when the prediction error and optional atom class constraints, such as orthogonality, converge within a given threshold.
[0075] The property model trained according to the above procedure can determine the target property of a substance based on any of the atomic descriptors, which can be derived from the structure of a molecule. This can be a structure obtained from an experimental technique, such as X-ray diffraction, but perhaps a more relevant use case for the presented model is the derivation of molecular properties from calculated structures and descriptors. This means that only calculations are required to determine the target property, and no experiments are required. This is a major advantage, for example, when a) virtual screening is performed when the candidate of interest is difficult to purchase or must be synthesized, or when the property of interest for a compound is easily available but involves expensive, error-prone, or dangerous measurement procedures.
[0076] To make such a determination based on molecular or periodic calculations, the first step typically involves one or more calculations to obtain atomic descriptors that refer to the structure of one conformer that is most stable under the application conditions, or to the structure of an ensemble of related conformers. In the case of a conformer ensemble, conformer weights according to the Boltzmann equation must also be calculated from the relative energies or free energies, if available. Then, depending on the trained atomic descriptors, atomic descriptors can be determined, for example, from each structure. This allows all input parameters for determining the target property using the property model described above to be available.
[0077] The determined physicochemical or application properties can generally be used to accelerate product or process development by significantly shortening the time until information is available about how a new compound will behave or how a compound will behave under different conditions. In the specific case of also learning the total energy of a substance as a target property, the resulting total energy and gradient with respect to nuclear coordinates based on atomic classes can also serve as a computational method for determining the most likely molecular or periodic structures and their ensembles as described above.
[0078] The general advantages of having predictive models for physicochemical or application properties based solely on calculations have already been presented in the previous sections. Furthermore, a specific advantage of the present invention is that it eliminates the need for, for example, meticulous, time-consuming, model-developer-dependent, and potentially biased hierarchical group assignments. Furthermore, atom classes can be assigned differently for each property, reflecting the fact that different atoms may play dominant roles for different properties. The present invention is also generally applicable to molecules with unclear / ambiguous chemical bonding situations, such as inorganic substances, metal-organic compounds, and molecules with noncovalent interactions, where molecular graphs derived by chemoinformatics tools are sensitive to small changes in bond lengths, and thus assignments to hierarchical functional groups are prone to errors. Apart from directly determining properties, the present invention is also particularly useful for delta learning, using the presented atom-class-based model to a) determine physical properties, and b) identify and correct systematic errors in the physical properties. Furthermore, the property models provide an unbiased level of insight into atomic similarities, and therefore functional group similarities, which may open new avenues for interpreting how properties are actually embodied and thus inspiring ideas on how to improve them.
[0079] Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims.
[0080] In the claims, the word "comprise" does not exclude other elements or steps and the indefinite article "a" or "an" does not exclude a plurality.
[0081] A single unit or device may fulfill the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.
[0082] Procedures such as providing atomic descriptors and properties, providing property models, applying or training property models, etc., performed by one or more units or devices may be performed by any other number of units or devices. These procedures may be implemented as program code means of a computer program and / or as dedicated hardware.
[0083] The computer program product may be distributed on a suitable medium, such as an optical storage medium or a solid-state medium, supplied together with or as part of other hardware, stored / distributed on other media, but may also be distributed in other forms, for example via the Internet or other wired or wireless telecommunications systems.
[0084] Any unit described herein may be a processing unit that is part of a classical computing system. A processing unit may include a general-purpose processor, a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or any other dedicated circuit. Any memory may be physical system memory and may be volatile, nonvolatile, or a combination of both. The term "memory" may include any computer-readable storage medium, such as a non-volatile mass storage device. If a computing system is distributed, processing and / or storage capabilities may also be distributed. A computing system may include multiple structures as "executable components." The term "executable component" is a structure well understood in the computing field, which may be software, hardware, or a combination thereof. For example, when implemented in software, those skilled in the art will understand that the executable component structure may include software objects, routines, methods, etc. that can be executed on a computing system. This may include both executable components in the computing system's heap or on a computer-readable storage medium. The executable component structure may reside on a computer-readable medium such that, when interpreted by one or more processors of the computing system, e.g., by processor threads, it causes the computing system to perform a function. Such structures may be directly computer readable by a processor, for example where the executable components are binary, or may be structured to be interpretable and / or compiled to generate such binaries, whether in a single step or multiple steps, that are directly interpretable by a processor. In other examples, the structures may be hard-coded or hard-wired logic gates that are implemented exclusively or nearly exclusively in hardware, for example in a field programmable gate array (FPGA), application specific integrated circuit (ASIC), or other dedicated circuitry.Thus, the term "executable component" is a term for a structure well understood by those skilled in the computing arts, whether implemented in software, hardware, or a combination. Any embodiments herein are described with reference to operations performed by one or more processing units of a computing system. When such operations are implemented in software, one or more processors direct the operation of the computing system in response to execution of the computer-executable instructions that make up the executable component. A computing system may also include communications channels that enable the computing system to communicate with other computing systems, for example, over a network. A "network" is defined as one or more data links that enable the transmission of electronic data between computing systems and / or modules and / or other electronic devices. When information is transferred or provided to a computing system via a network or another communications connection, e.g., either hardwired, wireless, or a combination of hardwired and wireless, the computing system properly considers the connection to be a carrier medium. A carrier medium may include a network and / or data link that can be used to carry desired program code means in the form of computer-executable instructions or data structures and that can be accessed by a general-purpose computing system or a special-purpose computing system, or a combination thereof. Although not all computing systems require a user interface, in some embodiments a computing system includes a user interface system for use in interfacing with a user. The user interface serves as an input or output mechanism to the user, for example, via a display.
[0085] Those skilled in the art will appreciate that at least portions of the present invention may be implemented in networked computing environments having many types of computing system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, cellular phones, PDAs, pagers, routers, switches, data centers, wearable devices such as eyeglasses, etc. The present invention may also be practiced in distributed system environments where tasks are performed by both local and remote computing systems that are linked, for example, through a network, by either hardwired data links, wireless data links, or a combination of hardwired and wireless data links. In a distributed system environment, program modules may be located in both local and remote memory storage devices.
[0086] Those skilled in the art will also understand that at least a portion of the present invention may be implemented in a cloud computing environment. A cloud computing environment may be distributed, but this is not required. If distributed, a cloud computing environment may be distributed internationally within an organization and / or have components held across multiple organizations. For purposes of this specification and the claims that follow, "cloud computing" is defined as a model that enables on-demand network access to a shared pool of configurable computing resources, such as networks, servers, storage, applications, and services. The definition of "cloud computing" is not limited to any of the many other advantages that may be obtained when such a model is deployed. The computing systems in the figures, as described, include various components or functional blocks that may implement various embodiments disclosed herein. The various components or functional blocks may be implemented on a local computing system or may be implemented in a distributed computing system that includes elements that reside in the cloud or that implement aspects of cloud computing. The various components or functional blocks may be implemented as software, hardware, or a combination of software and hardware. The computing systems shown in the figures may include more or fewer components than those shown in the figures, and some of the components may be combined where circumstances permit.
[0087] Any reference signs in the claims should not be construed as limiting the scope.
[0088] The present invention relates to an apparatus for determining a property of a substance. A providing unit provides atomic descriptors that indicate characteristics of atoms of the substance with respect to the structure of a particular molecule. A model providing unit provides a trained property model, which is trained to determine a property of the substance as an output when the atomic descriptors are provided. The property model includes: a) an atomic classifier adapted to classify atoms of the substance into one or more atomic classes based on the atomic descriptors; and b) a regression model adapted to determine a property based on the atomic classes determined for the substance. A determining unit determines the property of the substance by applying the trained property model to the atomic descriptors.
Claims
1. 1. An apparatus for determining the properties of a material defined as either exhibiting molecular properties or exhibiting extended periodic 1D, 2D or 3D structures, said apparatus (100) comprising: an atomic descriptor providing unit (110) for providing atomic descriptors characteristic of atoms of said substance with respect to said structure of said substance; a trained property model providing unit (120) for providing a trained machine learning based property model, the trained property model being adapted to determine properties of the substance as an output when the atomic descriptors are provided as an input, the property model including: a) an atomic classifier adapted to classify atoms of the substance into one or more atomic classes based on the atomic descriptors; and b) a regression model adapted to determine the properties based on the atomic classes determined for the substance; a property determination unit (130) for determining the property of the material by applying the trained property model to the atomic descriptors; and a characteristic providing unit for providing said characteristic for further processing.
2. The apparatus of claim 1 , wherein the atomic descriptors comprise atomic intrinsic quantities derived from the electronic structure of the material and / or from atomic features indicative of the atomic structural environment of the atoms of the material.
3. The device according to claim 1 or 2, wherein the properties of the specific molecule include physicochemical properties and / or application properties, where physicochemical properties refer to physical or chemical characteristics of a substance consisting of the specific molecule, and application properties refer to application-specific properties of the substance.
4. 10. An apparatus according to any one of the preceding claims, wherein, when the substance is of molecular nature and comprises molecules of different conformations, the atomic descriptors of the substance refer to a weighted average of the atomic descriptors of different conformers.
5. The apparatus of claim 4 , wherein the weights in the weighted average refer to Boltzmann weights or to trained weights that have been trained together with the feature model based on the same training data as that used to train the feature model.
6. 10. The apparatus of claim 1, wherein the atom classifier is adapted to classify the atoms of the substance into the one or more atom classes, such that atoms referring to different elements are classified into different atom classes.
7. 10. The apparatus of claim 1, wherein the atom classifier is adapted to classify atoms of the substance into one or more atom classes, and where the atoms are classified into more than one atom class, an atom class weight is provided for each of the atom classes to which the atoms are assigned based on the probability that the atoms belong to the respective atom class.
8. 10. The apparatus of claim 1, wherein the regression model refers to any one of a multiple linear regression, an artificial neural network, a kernel algorithm, a support vector regression, and a random forest algorithm.
9. A training apparatus for training a machine learning based feature model, the apparatus (300) comprising: a training data providing unit (310) for providing training data for training the property model, the training data including atomic descriptors and corresponding known properties of a plurality of different substances, respectively; a trainable property model providing unit (320) for providing a machine learning based property model, the property model including: a) an atom classifier trainable to classify atoms of a substance into one or more atom classes based on atomic descriptors; and b) a regression model trainable to determine the property based on the atom classes determined for a molecule; a training unit (330) for training the provided property model based on the training data so that the trained property model is adapted to determine a property of a substance as an output when an atomic descriptor of the substance is provided as an input.
10. 1. A system for screening potential substances for a given application, the system comprising: The device of claim 1; a target property providing unit for providing target properties of target substances and potential target substances; a screening unit adapted to utilize the device to determine a property of each of the potential target substances, compare the determined property of the potential target substances with the target property, and based on the comparison, either i) determine the potential target substance as the target substance, or ii) provide a new potential target substance and repeat the determination of the property using the new potential target substance.
11. 1. A computer-implemented method for determining properties of a material, defined as either exhibiting molecular properties or representing extended periodic 1D, 2D or 3D structures, said method (200) comprising: providing atomic descriptors (210) that characterize atoms of the material with respect to the structure of the material; providing (220) a trained machine learning based feature model adapted to determine as output a feature of the material when the atomic descriptors are provided as input, the feature model comprising: a) an atomic classifier adapted to classify atoms of the material into one or more atomic classes based on the atomic descriptors; and b) a regression model adapted to determine the feature based on the atomic classes determined for the material; determining 230 the properties of the material by applying the trained property model to the atomic descriptors; and providing the characteristics for further processing.
12. 1. A computer-implemented method for screening potential agents for a given application, the method comprising: Providing target characteristics of target substances and potential target substances; 12. A computer-implemented method comprising: utilizing the method of claim 11 to determine a property of each of the potential target materials; comparing the determined property of the potential target material to the target property; and based on the comparison, either: i) determining the potential target material as the target material; or ii) providing a new potential target material and repeating the determination of the property using the new potential target material.
13. 1. A computer-implemented training method for training a machine learning based feature model, the method (400) comprising: providing training data for training the property model (410), the training data including atomic descriptors and corresponding known properties of a plurality of different substances; providing (420) a trainable feature model for providing a machine learning based feature model, the feature model including: a) an atom classifier trainable to classify atoms of a material into one or more atom classes based on atomic descriptors; and b) a regression model trainable to determine the feature based on the atom classes determined for the material; and training (430) the provided property model based on the training data such that the trained property model is adapted to determine a property of a substance as output when atomic descriptors of the substance are provided as input.
14. 14. A computer program product for training a machine learning based feature model, comprising program code means for causing a feature model training device according to claim 9 to perform the training method according to claim 13.
15. A computer program product for determining a property of a substance, comprising program code means for causing an apparatus according to any one of claims 1 to 8 to carry out the method according to claim 11.