Device for determining property of substance
Through the machine learning-based property model, using atomic descriptors and classifier regression models, the problem of insufficient accuracy in the new functional group combination is solved, and the rapid and accurate screening and synthesis of material properties is achieved, which is suitable for a variety of industrial applications.
Patent Information
- Application Number
- CN202380085987.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-16
- Filing Date
- 2023-12-13
- Publication Date
- 2025-07-22
AI Technical Summary
The existing thermodynamic properties prediction models are insufficient in the face of new functional groups or close adjacent functional groups combinations, and require a large amount of training data, which cannot be effectively applied to industrial-related molecular screening.
Using machine learning-based property models, the properties of matter are determined through atomic descriptors and trained atomic classifiers and regression models, reducing dependence on expert knowledge, and suitable for extended periodic structures and molecular matters, which can improve accuracy with less training data.
Accurate prediction of the properties of matter outside the training set is achieved, the number of substances synthesized and tested is reduced, the efficiency and accuracy of industrial screening substances are improved, and it is suitable for rapid determination of various properties.
Smart Images

Figure CN120359573A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an apparatus, a method, and a computer program product for determining properties of substances in pure substances or given mixtures. Further, the present invention also relates to a training apparatus, a training method, and a computer program product for training a machine learning-based property model, which can be applied to the apparatus, the method, and the computer program product for determining the above properties of substances. Background Art
[0002] Generally, determining the thermodynamic properties of substances, which are defined, for example, as either having molecular properties or exhibiting extended periodic 1D, 2D, or 3D structures, is a challenging problem with high industrial relevance. Many established thermodynamic models, especially in the field of molecular substances, are so-called "group contribution" models, where properties are determined by contributions attributed to hierarchically derived functional groups in a molecule. However, deriving the hierarchy followed for decomposing a molecular structure into the corresponding functional groups is not only a cumbersome and time-consuming task for experts or sophisticated software but also a cause of many errors in predicting thermodynamic properties. For example, functional groups that adapt the same hierarchical pattern may exhibit significantly different behaviors in different chemical environments. However, such differences cannot be taken into account in current models. Therefore, such thermodynamic property prediction models can generally only be applied to specific and very limited chemical spaces and thus often fail to achieve industrial-relevant screening of specific properties for a large number of molecules. In addition, even to achieve acceptable accuracy within their limited application fields, these prediction models require a large amount of training data to train the corresponding models, so that models can only be provided for molecular groups for which such a large data basis already exists. However, for quite new functional groups or new combinations of closely adjacent functional groups, for example, there is usually not enough available data to achieve the accuracy that enables reasonable application of the model.
[0003] Therefore, it would be advantageous to provide a prediction model for determining the physicochemical properties of substances, which can (a) determine the properties of substances belonging to other classes outside the training set class and (b) improve the accuracy of determining properties while requiring less training data. Summary of the Invention
[0004] The object of the present invention is to provide a device, a method and a computer program product, so as to be able to very accurately determine the properties of substances defined as either having molecular properties or exhibiting extended periodic 1D, 2D or 3D structures. In addition, another object of the present invention is to provide a training device, a training method and a computer program, so as to be able to provide a prediction model that can be applied to the device, method and computer program for determining properties, and this prediction model: (a) can also determine the properties of substances belonging to other categories outside the training set category; and (b) can be trained to provide high determination accuracy with less training data. The flexibility of the present invention also allows the generation of prediction models for a variety of target properties, because the mathematical structure will naturally adapt the predictors of the model to the determined properties. Compared with the classical group contribution method, this model does not require experts to manually tune for the target property. This model also does not require the bonds and bond orders between atoms as inputs, which makes it applicable to cases where it is difficult to determine these bonds and bond orders, such as metal complexes and hydrogen bonding.
[0005] In a first aspect of the present invention, there is provided a device for determining the properties of a substance defined as either having molecular properties or exhibiting extended periodic 1D, 2D or 3D structures, wherein the device comprises: i) an atomic descriptor providing unit for providing atomic descriptors indicating the characteristics of the atoms of the substance with respect to a specific molecular structure; ii) a trained property model providing unit for providing a trained machine learning-based property model, wherein the trained property model has been trained to determine the properties of the substance as an output when these atomic descriptors are provided as inputs, and the property model comprises: a) an atom classifier adapted to classify the atoms of the substance into one or more atom categories based on these atomic descriptors; and b) a regression model adapted to determine the property based on the atom categories determined for the substance; iii) a property determination unit for determining the properties of the substance by applying the trained property model to these atomic descriptors; and iv) a property providing unit for providing the property for further processing.
[0006] Since the property model includes: a) an atomic classifier adapted to classify the atoms of a substance into one or more atomic classes based on atomic descriptors, and b) a regression model adapted to determine a property based on the atomic classes determined for the substance, in particular since both the atomic classifier and the regression model are part of a machine learning-based property model and can thus be trained simultaneously, enabling the use of atomic classes specifically determined for the corresponding training cases. Therefore, the atomic classes used to determine the property do not depend on the corresponding insights into the physical or chemical properties of the molecule provided by an expert or as a scientific convention, but are determined automatically by the training for the corresponding training cases, where the classification may vary depending on the property under consideration. In addition, these atomic classes map complex structures and chemical motifs into a low-dimensional space consisting of a finite number of classes. This enables the use of less training data to train the property model, further enables higher accuracy to be obtained when determining the property for the corresponding training cases, and ultimately enables predictions to be made for functional group-containing substances not covered by the training set. All of this is a consequence of the fact that the classification performed in this way is more targeted at those atomic features that are truly important for the property of interest.
[0007] Furthermore, due to the incredibly large number of potentially applicable substances for a specific application that are usually not yet fully explored, today, when faced with the technical task of finding a substance that is not only suitable for a specific application but also meets a multitude of additional target properties, a technical product engineer must synthesize and test a large number of possible substances, or carefully examine large datasets and libraries storing potential substances in order to find the corresponding substance that may be suitable for the application. Even when designed using sophisticated experimental methods, a large number of possible substances still need to be synthesized and experimentally tested. In this context, the above method can assist users (such as technical product engineers) in automatically and more quickly discovering potentially suitable substances. In particular, by using the above method, the user only needs to synthesize and test a much smaller number of potentially suitable substances, which have been determined to be likely to meet the corresponding target properties. Thereby, unnecessary synthesis and testing of substances can be avoided. Therefore, the method enables the user to more quickly and efficiently complete the technical task of finding substances suitable for technical applications.
[0008] Typically, the substance whose properties need to be determined is considered to either have molecular properties or exhibit an extended periodic 1D, 2D, or 3D structure. Preferably, the substance has molecular properties, i.e., is a molecular substance; more preferably, the substance has molecular properties and contains at least one of the following elements: H, C, N, O, F, Si, P, Cl, Br, and I. Generally, the principles of the present invention can be applied to any type of substance for which there is a corresponding training dataset. However, the larger the expected application range to which the model needs to be applied, the more training data needs to be provided to achieve sufficient accuracy. Additionally, preferably, the molecular substance has a molecular weight of less than 600 g / mol, more preferably less than 300 g / mol. Further, preferably, the specific molecule exists in the environment in a form that can be completely described by a simple structural formula containing relevant information. A simple molecular structure refers to a molecule that can be clearly described by the covalent bonds between the atoms of the molecule. Instances where this is not the case are, for example, systems with a dynamic equilibrium between several forms, such as monomers and oligomers, as in the case of several inorganic acids, or ionic substances with very localized charges that interact strongly with the solvent, such as through hydrogen bonds. A substance can also generally refer to a periodic system, i.e., a substance in which a specific molecule is structurally repeated in a defined manner in space. For example, in such a case, the substance can refer to a crystalline metal, semiconductor, or other system containing periodic atomic symmetry.
[0009] The atom descriptor providing unit is adapted to provide atom descriptors. In particular, the atom descriptor providing unit can be a storage unit or can communicate with a storage unit in which atom descriptors are already stored. However, the atom descriptor providing unit can also query or communicate with an atom descriptor determining unit that is used to determine atom descriptors, for example, based on the structural information of the molecular substance. Preferably, the atom descriptor indicates the characteristics of the atoms of the substance with respect to the atomic-level resolution 3D structure. Preferably, the atom descriptor includes atom-specific quantities derived from the atomic structural environment of the atoms of the substance and the electronic structure of the molecule.
[0010] Deriving atomic descriptors from the atomic structure environment can be carried out according to any formalism that can transform the three-dimensional structure and elemental composition of a molecular system into a set of atomic-level descriptors. These can be end-to-end architectures such as SchNet, PhysNet, but predefined atomic descriptors such as ANI, SOAP, FCHL can also be used. Atomic descriptors derived from molecular electronic structure calculations can refer to, for example, one or more of the following: partial charges, exposed surface fraction, average surface screened charge density calculated by a continuous solvation model, average atomic radius calculated based on the use of a flexible (e.g., isodensity cavity) continuous solvation model, contributions to energy or free energy in a COSMO-RS type or COSMO-SAC type solvation model, such as mismatch contributions, hydrogen bond donor contributions, hydrogen bond acceptor contributions, dispersion force contributions or total residual contributions, pairwise contact probabilities from a COSMO-RS type or COSMO-SAC type solvation model, NMR shielding constants.
[0011] The trained property model providing unit is adapted to provide a trained machine learning-based property model. In particular, the trained property model is configured (i.e., has been trained) to determine the property of a substance as output when provided with atomic descriptors as input. The determined property can refer to any property of the substance for which the trained machine learning-based property model has been trained. In particular, in a preferred embodiment, the property is a physicochemical property, which can refer to any of the following: melting point, boiling point, vapor pressure, heat of vaporization, heat capacity, flash point, autoignition temperature, liquid density, critical temperature and / or pressure, electrical conductivity and / or thermal conductivity, glass transition temperature, viscosity, surface tension, refractive index, general mechanical properties, total energy, dipole moment, polarizability, HOMO-LUMO or band gap, ionization potential, electron affinity, activity coefficient, partition coefficient, solubility, cloud point, critical micelle concentration, acid and base dissociation constants, hydrolysis stability, corrosivity, and catalytic activity towards a given chemical reaction, or spectroscopic properties such as transition energies and intensities. Further, in a preferred embodiment, the application property can refer to any of the following: a) general properties beyond pure physicochemistry, such as various types of toxicity and ecotoxicity behavior, biodegradability in different habitats, odor, ozone depletion potential, global warming potential, etc.; and b) performance characteristics as an additive in a mixture, such as the effect on fuel octane and cetane numbers, stability against UV radiation or oxidation, corrosion protection of surfaces in solutions or coatings, flame retardancy, effect on touch, etc.
[0012] In a preferred embodiment, the determined property refers to the melting point of a substance. Determining the melting point of a substance using a property model enables very quickly (i.e., without synthesizing the substance and performing a complex sequence of measurements) determining whether the substance in question is suitable for agrochemical or pharmaceutical applications, or for example as a resin raw material for epoxy resin or polyurethane systems. In particular, for these applications, it is important that the substance in question does not have an overly high melting point. Thus, using this property model, potential substances can be scanned very quickly for their general applicability. Additionally or alternatively, in one embodiment, the property refers to the vapor pressure. For example, if the substance refers to a potential plastic-compatible substance, in this way it is possible to check whether the volatility of the substance is low enough to ensure that the additive substance remains in the desired material and maintains its efficacy before synthesizing the plastic additive or preparing a sample for application testing. For example, this can enhance the safety of the composition. Furthermore, the determination of the vapor pressure of a substance together with other properties such as the activity coefficient enables the determination of suitable conditions for separation processes such as distillation. Additionally or alternatively, in one embodiment, the property refers to the thermal conductivity. This enables the device to be applied to finding polymeric materials that can act as heat conductors or insulators and to perform virtual screening of new insulating materials that can be used, for example, to select suitable materials for electric devices to avoid overheating. Additionally or alternatively, in one embodiment, the property refers to the viscosity. Determining the viscosity of potential substances using this device enables determining whether the viscosity is within a given range, for example, in the context of polymer processing such as using reactive resins or injection molding or film blowing after an extrusion process. Additionally or alternatively, in one embodiment, the property refers to the solubility. Screening potential substances for solubility using this device is advantageous for designing processes for finding, for example, ideal solvents and / or precipitants for a given chemical synthesis or extraction / purification scheme. Furthermore, without using a complex test sequence, it can be ensured that potential active ingredients (i.e., substances) are bioavailable or that they do not stratify in a formulated product. Additionally or alternatively, the property refers to the ionization potential. Training a property model to determine the ionization potential enables applying the device to screening potential battery materials such as antioxidant matrices, solvents, films. Additionally, the ionization potential can also be used as a physical descriptor of a molecule that can be correlated with other application properties, such that the determination of the ionization potential itself can again be applied in the case of another trained property model for determining another property of a substance. Additionally or alternatively, in one embodiment, the property refers to one or more spectroscopic properties such as transition energy and / or intensity. In this case, the device can also be applied to verification processes where it is necessary to verify that the synthesis step has indeed yielded the desired product. Additionally or alternatively, the property refers to the flame retardancy, which can be the flame retardancy of the substance itself as a material or the flame retardancy of a physical mixture of a polymer and the substance, where the substance acts as a flame retardant additive.Determining the flame retardancy of additive-modified polymers enables virtual screening to find novel flame retardants that are more efficient than the existing substances in the plastic formulation. Additionally, materials with inherent flame retardancy can be screened very quickly and effectively in this way.
[0013] The trained property model includes: a) an atom classifier adapted to classify the atoms of a molecule into one or more atom classes based on atom descriptors; and b) a regression model adapted to determine a property based on the atom classes determined for a substance. The trained property model is a machine learning-based model, where the atom classifier and the regression model are trained simultaneously. In particular, the atom classifier is a machine learning-based classifier. The atom classifier can refer to any machine learning classifier that can be trained to provide one or more classes of an object and classify the object into the corresponding class based on the characteristics of the provided object (in this case, based on the provided atom descriptors). In particular, the atom classes utilized by the atom classifier are not predetermined, for example, by user consideration or an expert class hierarchy, but are trained simultaneously with the regression model such that for each corresponding case, for example, for each corresponding property to be determined, the optimal atom classes can be determined solely through training. Thus, the atom classes of the atom classifier are determined in a purely data-driven manner. Additionally, the regression model that determines a property using the atom classes determined for a molecule can refer to any machine learning-based regression model. Generally, a machine learning regression model refers to any machine learning model that assigns a continuous dependent variable to a set of input variables, where, in this application, the continuous dependent variable refers to a property and the set of input variables refers to the values of the corresponding atom classes. Preferably, the regression model refers to any one of multiple linear regression, artificial neural network, Gaussian process regression, kernel algorithms, support vector regression, and random forest algorithms. Utilizing a classifier that can learn atom classes based on training data and the subsequent regression model has many advantages. In particular, in this case, the atom classes provide a description of the characteristics of the corresponding atoms of a substance and their impact on the property, which is learned from the training data and thus specific to each application and independent of known convenient classifications. This enables the provision of inherent information that is already part of the model, thereby leading to the possibility of reducing the amount of training data, for example, by reducing the dimensionality of the information provided by the training data. Additionally, this enables the training of the model to utilize more meaningful descriptors, which also enable the interpretation of the results of the training, i.e., the property model. Further, the inventors have found that the above property model is more robust to artifacts that may be caused by biased training data.
[0014] Then, the property determination unit is adapted to determine the property of a substance by applying a trained property model to atomic descriptors. In particular, the atomic descriptors are provided as input to the trained property model, which then provides the determined property of the substance as output. The determined property can then be provided, for example, by the property provision unit for further processing. For example, further processing can refer to providing the determined property as output to a user, for example, via a screen or any other suitable means. However, the determined property can also be provided for further processing for a specific application, for example, it can be used internally in a screening process to screen a variety of substances, where, in this case, the corresponding property is not necessarily provided to the user, but is used, for example, only to determine a list of potentially suitable substances for the application and for further processing, such as automatically synthesizing the substances on the list by an automatic synthesis unit. In addition, the determined property can also be provided to a database and stored, and then further processed, for example, in a search operation of the database.
[0015] Preferably, the processing of the determined property includes determining a control signal for controlling and / or monitoring a production process based on the determined property. The production process can refer to the production process of the substance, or can refer to the production process of a product utilizing the substance. For example, if the determined property indicates that the substance has a specific vapor pressure, the generation of the control signal can include: generating a control signal for controlling and / or monitoring the production process of a polymer utilizing the substance, taking into account the vapor pressure of the substance. In a preferred embodiment, the control signal indicates a machine-executable formulation of the substance, particularly when comparing indicates that the determined property of the substance is within a predetermined range around the provided target property. In addition, the method can include controlling and / or monitoring the production process based on the control signal.
[0016] In addition, a process of a processing nature may also refer to the step of selecting one or more substances based on separately determined properties. For example, if corresponding properties have been determined for a plurality of potential substances, the selection may include comparing the determined properties of the different substances with a predetermined selection criterion and selecting the substances whose determined properties meet these criteria. In particular, in one embodiment, the method includes: receiving a target property of a substance, comparing the received target property with the determined property, and providing a control signal based on the comparison result. The control signal may refer to any signal capable of further controlling a technical system. For example, the control signal may be adapted to control an interface to provide the comparison result on the interface. In a preferred embodiment, the comparison refers to the verification of the target property, wherein if the determined property falls within a predetermined range around the target property, the verification is affirmative. In this case, the control signal may be adapted to simply control the user interface to provide an indication of an affirmative or negative verification result. However, preferably, the control signal refers to a formulation, i.e., a preparation method, of one or more substances that meet the specified target property, i.e., are affirmatively verified. A formulation is generally defined as an instruction on how to synthesize or prepare a substance. In particular, the formulation includes starting substances and the corresponding parameters required to generate the substance from the starting substances. Preferably, the control signal includes a formulation in a form that directly allows automatic control of the corresponding industrial system or labor equipment used to produce the substance. In particular, preferably, when the result of the comparison refers to the determined property being within a predetermined range around the target property, the control signal indicates the machine-executable formulation of the substance.
[0017] In one embodiment, the molecular substance is considered in more detail as a collection of several conformational isomers. In this case, the atomic descriptors of the molecule can refer to the weighted average of the atomic descriptors of the different conformational isomers of the substance. Generally, conformation refers to conformational isomerism, in which the isomers can be interconverted by rotation around chemical bonds (mainly formal single bonds), or can also be achieved by other types of interconversions (such as ring flipping or pseudorotation of cyclic non-aromatic hydrocarbons). For all these potential interconversions, the standard considers them to be kinetically possible, that is, at the temperature of interest, they will actually occur on a reasonable time scale. In this case, a substance in which an atom may exist in different environments may result in different atomic descriptors for different conformational isomers. Here, by utilizing atomic descriptors, which are the weighted average of the atomic descriptors of different conformational isomers, the accuracy of determination can be improved. In particular, the weights can indicate the probability that a particular conformational isomer of the molecule exists in the substance. For example, if it is expected that two possible conformational isomers have the same probability of existing in the substance, the weights can indicate the expected ratio of the conformational isomers in the substance. Preferably, the weights for the weighted average refer to Boltzmann weights. The conformational energy or free energy required to determine the Boltzmann weights can be obtained through quantum chemical calculations, semi-empirical calculations, or machine learning calculations. In addition, the weights can also refer to the trained weights that are co-trained with the property model based on the same training data as the training data used to train the property model.
[0018] In one embodiment, the atomic classifier is provided with predetermined constraints for classifying the atoms of a substance based on the atomic descriptors. Providing predetermined constraints for the classification enables further improvement of the classification efficiency by restricting the classification, for example, to chemically reasonable categories. Preferably, in one embodiment, the atomic classifier is adapted to classify the atoms of a substance into one or more element-specific atomic categories such that atoms of different elements are classified into different atomic categories. Providing constraints for the training of the atomic classifier enables direct consideration of the fact that in many cases, different elements will have different chemical properties that have different effects on the properties. Therefore, this constraint allows further reduction of the amount of training data required to train the property model. However, if a more flexible property model is desired, the atomic classifier can also be trained without imposing any element-specific restrictions on the atomic categories constructed by the atomic classifier during training.
[0019] In one embodiment, an atom classifier is adapted to classify atoms of a substance into one or more atom classes, where, if an atom is classified into more than one atom class, atom class weights are provided to each atom class to which the atom is partially assigned, based on the degree of match between a given atomic property and the class definition. In some cases, an atom may include features that result in the atom being assigned to more than one atom class, i.e., atom descriptors. In such cases, the atom can be considered to affect the corresponding properties according to two atom classes. However, since in such cases there is only one atom with an impact associated with the corresponding atom class, atom class weights are provided to prevent the potential impact of this one atom from being the same as that of a molecule in which each corresponding atom class contains one atom. In particular, the atom class weights are assigned based on the degree of match between the given atomic property and the class definition regarding the corresponding property. However, in one implementation, the classifier can also be adapted to only allow an atom to be classified into one atom class. This results in a less flexible training process but enables further reduction of the training data.
[0020] In another aspect of the present invention, a training apparatus for training a machine learning-based property model is proposed, where the apparatus includes: i) a training data providing unit for providing training data for training the property model, where the training data respectively include atom descriptors of a plurality of different substances and corresponding known properties; ii) a trainable property model providing unit for providing a machine learning-based property model, where the property model includes: a) an atom classifier trainable to classify atoms of a substance into one or more atom classes based on the atom descriptors; and b) a regression model trainable to determine the property based on these atom classes determined for the molecule, and iii) a training unit for training the provided property model based on the training data such that the trained property model is adapted to determine the property of the substance as an output when provided with the atom descriptors of the substance as an input.
[0021] Typically, training data including atomic descriptors of various different substances and corresponding known properties can be provided in the form of a database that stores the characteristics and properties of substances. Such a database can be based on, for example, simulation data, measurement data, sequencing data, etc. Preferably, the training data relates to only one property of the substance, such that the property model trained based on the training data is dedicated to determining the corresponding property. However, more than one property can also be provided in the training data, such that the trained property model can determine more than one property for each substance. Generally, any known training algorithm can be used to train a data-driven, particularly machine learning-based model. Preferably, during the training of the property model, the substance descriptors that have the greatest impact on the property are also determined, and the model is trained based on these most impactful descriptors. To determine these most impactful descriptors, for example, clustering analysis or PCA analysis tools can be utilized. In particular, the descriptors can be used to determine the application space of the training data, where the application space is defined by the substance descriptors covered by the data. Then, the determination of the most impactful descriptors can be performed as a dimensionality reduction of the application space. Then, an algorithm for optimizing the training data in the application space can be applied, for example, in order to cover the application space with as little training data as possible.
[0022] Preferably, the training of the property model includes iteratively optimizing the model parameters of the provided property model based on the training data until the trained property model achieves a predetermined accuracy improvement. In particular, the model parameters include classification parameters and regression parameters, and the iterative optimization of the model parameters includes optimizing the classification parameters of the atomic parameters and the regression parameters of the regression model in each iteration step. Generally, the accuracy improvement can refer to, for example, typically for an iterative method, the difference between the result of the current iteration step and the result of the previous iteration step. When the iteration converges, as typically occurs, the difference between the results of each iteration step becomes smaller and smaller with each step, where if the difference (i.e., the accuracy improvement) is below a predetermined accuracy improvement (i.e., a predetermined threshold), the iteration can be stopped. Possible variations of this process include keeping certain parameters constant for several iteration steps and then adjusting these parameters, or using a convergence criterion that takes into account the development of the accuracy over several iteration steps.
[0023] On the other hand, a system for screening potential substances in a predefined application is proposed, wherein the system comprises: a) the device as described above, b) a target property providing unit configured to provide a target property and potential target substances of a target substance, and c) a screening unit, wherein the screening unit is adapted to determine the corresponding properties of the potential target substances using the device, compare the determined properties of the potential target substances with the target property, and perform one of the following operations based on the comparison: i) determine the potential target substance as the target substance, or ii) provide new potential target substances and repeat the determination of the properties using the new potential target substances. Preferably, the method further comprises generating a control signal indicating a machine-executable recipe of the target substance for controlling and / or monitoring the production of the target substance. Preferably, the substance is a polymer and the machine-executable recipe includes the synthesis specifications of the polymer. Additionally, the method may include controlling and / or monitoring the production of the target substance based on the control signal.
[0024] The target property providing unit is configured to provide a target property indicating the characteristics of a substance. In particular, the providing may refer to, for example, receiving the target property from a user's input using a corresponding input unit. Additionally, the providing may also refer to accessing a storage unit in which the target property is stored. Further, the providing may also include, for example, receiving the target property from other sources via a network connection and providing the received target property. Generally, the target property may refer to a target value, such as a target vapor pressure, or a numerical range that the substance needs to satisfy. Additionally, the target property may also refer to any type of target function, such as the temporal sequence of the property. Further, the target property providing unit is configured to provide potential target substances. In particular, the providing may refer to, for example, receiving the potential target substances from a user's input using a corresponding input unit. Additionally, the providing may also refer to accessing a storage unit in which the potential target substances are stored. Further, the providing may also include, for example, receiving the potential target substances from other sources via a network connection and providing the received potential target substances. Then, the potential target substances may be regarded as starting substances for screening substances that satisfy the target property. The corresponding descriptors of the substances may be determined based on the provided potential target substances, for example, as described above. Generally, the descriptors may be determined by accessing a storage unit in which the descriptors of the corresponding substances are stored. Additionally, any of the above methods for determining descriptors based on the provided substances may be utilized.
[0025] After determining the properties of the potential target substance using the property model, the determined properties are compared with the target properties. Based on this comparison, a decision is made whether to identify the potential target substance as the target substance, where, in this case, the iteration can be stopped at this point. Additionally, based on this comparison, it is also possible to identify new potential target substances and repeat the determination of properties using the new potential target substances. Thus, an iteration will be performed at this time, in which the determination of properties using the property model and the descriptors of the potential target substance is repeated until one of the potential target substances is identified as the target substance. In particular, this comparison can include determining whether the determined properties of the potential target substance are within a predetermined range around the target properties, where, in this case, the target can be considered satisfied and the potential target substance can be identified as the target substance. If the determined properties are outside the predetermined range around the target properties, it is determined that the target is not satisfied and new potential target substances that may satisfy the target properties are provided.
[0026] Generally, the iteration performed can refer to any search of the potential target substance space or can be a directed search. For example, new potential target substances can be simply randomly selected from a large number of computer-simulated potential target substances. However, it is also possible to apply specific rules for generating new potential target substances based on the comparison between the determined properties and the target properties, and to consider or not consider the simultaneous optimization of more target properties of the substance. Generally, known methods can be used to generate new target substances. For example, evolutionary algorithms or Bayesian optimizers can be used. For example, such algorithms can be used to change the molecular structure or periodic atomic structure of the substance to generate new potential target substances.
[0027] Then, the iteration can be performed within the step of determining the properties of the new potential target substance by using the descriptors of the new potential target substance as described above. Optionally, if the descriptors have not been provided or do not exist on the storage unit, determining the descriptors from the potential target substance can also be part of the iteration. Additionally, it is preferred to use the same property model to determine the properties in all iteration steps. However, in some cases, different property models can also be used in different iteration steps. For example, if other descriptors of the new potential target substance are used, then another property model may be more appropriate. For example, this may be the case if some descriptors for the new potential target substance are not available or cannot be obtained with suitable accuracy.
[0028] After the iteration has stopped, for example, after a potential target substance has been determined as the target substance, or if no new potential target substance can be selected or generated, the result of the iteration can be provided to the user. For example, if all possible potential target substances do not meet the target properties, the user can be notified that the target substance has not been determined. If the target substance can be determined, the target substance can be provided to the user as an output. For example, the determined target substance can then be provided to an output unit or a computing unit for further processing. Preferably, the provision of the target substance will result in further processing to generate a control signal, as described above, for example.
[0029] In another aspect of the present invention, a method for determining the properties of a substance is proposed, where the substance is defined as either having molecular properties or exhibiting an extended periodic 1D, 2D, or 3D structure. The method includes: i) providing atomic descriptors indicating the characteristics of the atoms of the substance with respect to the structure of the substance; ii) providing a trained machine learning-based property model, where the trained property model is adapted to determine the properties of the substance as an output when these atomic descriptors are provided as inputs. The property model includes: a) an atomic classifier adapted to classify the atoms of the molecule into one or more atomic classes based on these atomic descriptors; and b) a regression model adapted to determine the properties based on the atomic classes determined for the substance; and iii) determining the properties of the substance by applying the trained property model to these atomic descriptors.
[0030] In another aspect of the present invention, a training method for training a machine learning-based property model is proposed. The method includes: i) providing training data for training the property model, where the training data respectively includes atomic descriptors of a variety of different substances and their corresponding known properties; ii) providing a machine learning-based property model, where the property model includes: a) an atomic classifier that can be trained to classify the atoms of a substance into one or more atomic classes based on atomic descriptors; and b) a regression model that can be trained to determine the properties based on the atomic classes determined for the substance; and iii) training the provided property model based on the training data such that the trained property model is adapted to determine the properties of a substance defined as either having molecular properties or exhibiting an extended periodic 1D, 2D, or 3D structure as an output when the atomic descriptors of a specific molecule are provided as inputs.
[0031] On the other hand, a method for screening potential substances in a predefined application is proposed, wherein the method comprises: a) providing the target properties of the target substance and potential target substances; and b) determining the corresponding properties of the potential target substances by using the method as described above, comparing the determined properties of the potential target substances with the target properties, and performing one of the following operations based on the comparison: i) determining the potential target substance as the target substance; or ii) providing new potential target substances and repeating the determination of the properties by using the new potential target substances. On the other hand of the present invention, a computer program product for training a property model based on machine learning is proposed, wherein the computer program product comprises program code means for causing the property model training means as described above to execute the training method as described above.
[0032] On the other hand of the present invention, a computer program product for determining the properties of a substance is proposed, wherein the computer program product comprises program code means for causing the means as described above to execute the method as described above.
[0033] It should be understood that the means as described above, the method as described above, and the computer program product as described above have similar and / or exactly the same preferred embodiments, in particular as defined in the corresponding dependent claims.
[0034] It should be understood that the preferred embodiments of the present invention may also be any combination of the dependent claims or the above embodiments and the corresponding independent claims.
[0035] These and other aspects of the present invention will become apparent and be elucidated with reference to the embodiments described hereinafter. Description of the Drawings
[0036] In the drawings:
[0037] Figure 1 An embodiment of a device for determining the properties of a substance is schematically and exemplarily shown.
[0038] Figure 2 An embodiment of a method for determining the properties of a substance is schematically and exemplarily shown.
[0039] Figure 3 An embodiment of a device for training a property model is schematically and exemplarily shown.
[0040] Figure 4 A flowchart of a method for training a property model is schematically and exemplarily shown.
[0041] Figure 5 and Figure 6 An exemplary application of the present invention is schematically shown.
[0042] Figure 7 Schematically and exemplarily shows the application of the present invention for optimizing formulations, and
[0043] Figure 8 schematically and exemplarily shows a method for improving the accuracy of determination of corresponding property models and / or for screening new substances for applications. Detailed implementation manners
[0044] Figure 1 Schematically and exemplarily shows an embodiment of a device 100 for determining the properties of a substance, which substance is defined as either having molecular properties or exhibiting an extended periodic 1D, 2D or 3D structure. The device 100 includes an atomic descriptor providing unit 110, a trained property model providing unit 120 and a property determining unit 130. In addition, the device may further include Figure 1 a property providing unit not shown in the figure. In particular, the device 100 can be implemented in any form of software and / or hardware, which software and / or hardware can be dedicated to the task of determining the properties of a substance or can provide additional more functions. The device 100 may also include an input unit 141 (such as a keyboard, a mouse, a touch screen, etc.) and / or an output unit 142 for outputting information to the user (such as by means of a display, an audio output, using a lighting unit, etc.). However, the input unit 141 and the output unit 142 may also be omitted, and the device 100 may also be provided, for example, as part of the software and / or hardware of a network of computing devices (such as the cloud), where in this case, the device can be connected to other computing devices, which may act as input or output units, or allow information to be provided to the device and / or received from the device (such as determination results) (for example, in the form of stored information), and this information can then be further processed without explicitly providing it as an output to the user.
[0045] The atomic descriptor providing unit 110 is adapted to provide atomic descriptors, such as atomic descriptors stored in a corresponding atomic descriptor database, or atomic descriptors pre-determined according to the corresponding knowledge of a specific molecule. Generally, the atomic descriptor indicates the characteristics of the atoms of the substance regarding the substance structure. For example, in a preferred embodiment, the atomic descriptor includes atom-specific quantities derived from the molecular electronic structure of the substance and / or atomic features indicating the atomic structure environment of the atoms of the substance. Then, the atomic descriptor providing unit 110 is adapted to provide, for example, the atomic descriptor to the property determining unit 130.
[0046] Further, the apparatus 100 includes a trained property model providing unit 120 adapted to provide a trained machine learning-based property model. The trained property model is adapted to (i.e., has been trained to) determine a property of a substance as an output when provided with an atomic descriptor as an input. In particular, the property of the substance determined by the trained property model may refer to physicochemical properties and / or application properties. Physicochemical properties refer to the physical or chemical characteristics of a substance that are defined as either having molecular properties or exhibiting extended periodic 1D, 2D, or 3D structures, such as melting point, viscosity, etc. Application properties refer to the application-specific characteristics of a substance that are defined as either having molecular properties or exhibiting extended periodic 1D, 2D, or 3D structures, such as biodegradability, toxicity, flash point, etc. Generally, for example, depending on the corresponding intended application, certain properties may be considered to belong to both physicochemical properties and application-specific properties. Generally, the trained property model may utilize, for example, the training apparatus and training method described with respect to Figure 3 and Figure 4 for training.
[0047] Generally, the trained property model includes: an atom classifier adapted to classify atoms of a molecule into one or more atom classes based on an atomic descriptor; and a regression model adapted to determine a property based on the atom classes determined for the molecule. In particular, the atom classifier may refer to any machine learning classifier that can be trained to provide one or more classes of an object and classify the object into the corresponding classes based on the characteristics of the provided object (in this case, based on the provided atomic descriptor). The atom classes utilized by the atom classifier are not, for example, predetermined by user consideration or an expert class hierarchy, but are trained simultaneously with the regression model such that for each corresponding case, for example, for each corresponding property to be determined, the optimal atom classes can be determined only through training. Thus, the atom classes of the atom classifier are determined in a purely data-driven manner. In addition, the regression model that determines a property using the atom classes determined for the molecule may refer to any machine learning-based regression model. Preferably, the regression model refers to any one of linear regression, multiple linear regression, artificial neural network, Gaussian process regression, kernel algorithms, support vector regression, and random forest algorithms.
[0048] Preferred examples of the classifier and regression model will be described in more detail below. Preferably, the classifier utilized assigns a probability (e.g., in the form of an atomic class coefficient) of belonging to a certain atomic group to each atom of the substance. This can be achieved during the training process by first applying multiple linear regression to each atomic descriptor vector, resulting in unnormalized atomic class coefficients. Subsequently, normalization can be performed through the softmax function. The atomic descriptor vector can be directly used as the input to the classifier or, after additional transformation using, for example, an artificial neural network, used as the input to the classifier. An alternative normalization scheme for the softmax can be employed as long as the optimization function of the property model is differentiable and the final atomic class coefficients are non-negative and the sum of the coefficients for each atom is one. An example is to square each value in the vector of unnormalized atomic class coefficients and divide by the sum of the squared values. Then, the normalized atomic assignment can be used in the regression model to determine the final property and to impose additional constraints on the training process. Examples include orthogonality constraints to encourage atoms with different chemical properties to be assigned to different groups and entropy-based constraints to limit the number of populated groups during training.
[0049] Then, the trained property model of unit 120 is adapted to provide, for example, the trained property model to the property determination unit 130. Then, the property determination unit 130 is adapted to determine the property of the substance by applying the trained property model to the atomic descriptor (i.e., by providing the atomic descriptor of the substance as input to the trained property model). Then, the output of the trained property model refers to or indicates the corresponding property for which the trained property model was trained. Typically, the property determination unit 130 may also be adapted to apply more than one property model to the atomic descriptor, for example, apply different property models trained for different corresponding properties, such that the property determination unit 130 is adapted to determine more than one property of the substance. Then, these different trained property models can be applied to the atomic descriptor subsequently or in parallel. In particular, in some embodiments, the property determination unit may be adapted to use one or more results determined by the trained property model (i.e., the properties of the substance) again as atomic descriptors and / or determine more atomic descriptors that can be provided again to the corresponding trained property model, thereby generating more properties of the substance.
[0050] Optionally, the device 10 may further include a property providing unit for providing properties for further processing. For example, the property providing unit may refer to the output unit 142, which processes the property so that it can be provided to the user, for example, via a display. However, the property providing unit may also be configured as an interface for docking with other systems (such as a production process control system). In this case, the further processing may refer to generating control data for controlling the production process control system based on the determined property. For example, if the determined property meets the target property, the production process may be controlled to produce the substance; or, for example, if the substance is utilized in the production process and the corresponding determined property affects the production parameters, the production process may be configured to set one or more process parameters of the production process based on the determined property.
[0051] Figure 2 A method for determining the properties of a substance is schematically and exemplarily shown. The method 200 includes step 210: providing an atomic descriptor, as described above. In addition, the method 200 includes step 220: providing a trained machine learning-based property model. In particular, the trained machine learning-based property model may refer to the property model as described above, which has been trained using, for example, the Figure 3 and Figure 4 devices and methods described. Generally, step 210 of providing the atomic descriptor and step 220 of providing the trained machine learning-based property model may be performed in any order, or even simultaneously. Then, in the final step 230, the method 200 includes determining the properties of the substance by applying the provided trained property model to the atomic descriptor.
[0052] Figure 3 A training device 300 for training a machine learning-based property model, which can be applied to the devices and methods as described above, is schematically and exemplarily shown. The training device 300 includes a training data providing unit 310, a trainable property model providing unit 320, and a training unit 330.
[0053] The training data providing unit 310 is adapted to provide training data for training a property model. In particular, the training data respectively includes atomic descriptors of a plurality of different substances composed of different substances and corresponding known properties. For example, the training data providing unit 310 may be connected to a database 340 that stores a plurality of such data sets including atomic descriptors of molecules and corresponding properties. Based on the desired training of the property model, for example, based on the desired output of one or more corresponding properties, the training data providing unit 310 may also be adapted to select corresponding training data from the database 340 to provide the training data. However, the data in such a database 340 may also be selected by the user, or selected by the training data providing unit 310 based on specific criteria provided by the user.
[0054] Then, the trainable property model providing unit 120 is adapted to provide a trainable machine learning-based property model. In particular, the structure of the trainable machine learning-based property model may be consistent with the structure of the trained property model as described above. Generally, a trainable property model may be regarded as a property model in which the parameters for determining the processing of the input in the property model to generate the output of the property model are adjustable. For example, these parameters may be set to predetermined values at the beginning and then adjusted during the training process of the property model based on the training data, so that the property model can subsequently perform its function. However, the trainable property model may also refer to a trained property model, such as a property model that has been trained with a corresponding training data set. In some cases, it may be desirable to retrain such a trained property model, for example, in order to improve accuracy or expand the applicability of the property model, such as expanding it to different substances or more properties. In this case, retraining of the property model may refer to providing a new training data set to the property model to further adjust the parameters of the trainable property model.
[0055] Then, the training unit 330 is adapted to train the provided property model based on the training data and the provided trainable property model. Preferably, the training of the property model refers to iterative training, in which the model parameters of the provided property model are optimized based on the training data for its corresponding task until the trained property model achieves a predetermined accuracy improvement. In particular, a threshold may be set, and through this threshold, it can be determined when the training of the property model has converged to a state where it is expected that no further improvement beyond the corresponding threshold can be obtained through more training. In addition, preferably, the model parameters, which can be divided into classifier parameters and regression parameters, are trained simultaneously, for example, jointly optimized in each iteration step.
[0056] Figure 4A flow chart of a method for training a machine learning-based property model is schematically and exemplarily shown. Method 400 includes step 410: providing training data for training properties, for example according to the above principles. Further, the training method 400 includes step 420: providing a trainable machine learning-based property model. The trainable machine learning-based property model preferably has a structure as described above. Generally, step 410 of providing training data and step 420 of providing a trainable machine learning-based property model can be performed in any order, or even simultaneously. Further, method 400 includes step 430: training the provided property model based on the training data, so that the trained property model is suitable for determining the properties of a substance composed of a specific molecule as an output when the atomic descriptors of the specific molecule are provided as input. Generally, training can be performed according to any principle as described above.
[0057] Figure 5 and Figure 6 An exemplary application of the above-mentioned device for reducing energy input in the context of extractive distillation as a method for separating fluid mixtures is shown. The quality of chemical production processes is usually characterized by the purity of the product, which needs to be achieved by the purity of the educts and the separation of the by-products generated during the reaction. If conventional distillation processes are limited to a small separation factor of components (such as product and similar by-components), additives can be provided to increase the separation factor. Using the above-mentioned device to screen such additives, a technically and economically more optimal distillation process can be obtained, in particular, a significant reduction in energy input can be achieved, for example when extractive distillation is used.
[0058] In order to screen the best additive for extractive distillation, the following thermodynamic data are preferably used: a) the pure component vapor pressure of the component to be separated and the pure component vapor pressure of the additive; and b) the activity coefficient or the activity coefficient of the component to be separated at infinite dilution when the additive is used as a solvent (as a preliminary screening approximation). The selectivity can be defined as the ratio of the activity coefficients of the component separation at infinite dilution in the additive solvent. This selectivity, combined with the pure component vapor pressure data, is directly related to the separation factor and investment of the extractive distillation using the additive, that is, the investment and energy consumption of the distillation stage. In addition, the reciprocal value of the activity coefficient of the component in the additive can be used as a preliminary approximation of the capacity standard to avoid two-phase separation of the mixture in the tower.
[0059] The apparatus described above can be used to determine the pure component vapor pressures and activity coefficients of different potential additives, thereby allowing preselection of potentially beneficial additives for a wide variety of components to achieve optimal distillation and separation processes. For example, the following equation can be used to optimize for the maximum relative volatility difference of the desired product with respect to other components, using the respectively determined vapor pressures and infinite dilution activity coefficients of the pure products:
[0060]
[0061] In this equation, and are the infinite dilution activity coefficients of the product and the component, respectively, and and are the determined vapor pressures of the corresponding pure products. The corresponding vapor pressures and infinite dilution activity coefficients of the pure products can be determined as the properties of the corresponding substances using the apparatus and method described above, and an iterative optimization scheme can be employed to determine potential additives that maximize the relative volatility.
[0062] In the following further example, the application of the above apparatus and method in the determination of the temperature-dependent vapor pressures of pure substances, such as for the above applications, is described. The model preferably consists of a SchNet model in this application for determining suitable atomic descriptors, followed by linear regression and Softmax activation, as a classifier for classifying atomic groups, preferably allowing up to 50 different groups. The frequency of groups in a substance can be used to determine the decimal logarithm of the vapor pressure through linear regression, and then multiplied by a factor f related to the target temperature T and the reference normal boiling point TNBP of the substance. For example, the following relationship can be utilized: f = (T / TNBP - 1) / (T / TNBP - 0.125). The training dataset can be determined based on the vapor pressure curves and normal boiling points taken from the DIPPR dataset. In this example, the substances are limited to contain only C, H, N, and O elements, resulting in a set of 1016 substances. For each substance, the vapor pressure can be evaluated at 20 different temperatures within the range specified in the DIPPR database, ultimately obtaining a total of 20580 data points. Among them, 80% can be used for training, 10% for validation, and 10% for testing. This split can be performed randomly, but subject to the condition that data belonging to the same compound can only appear in one of the subsets. The training process can utilize the Huber loss on the determined and true decimal logarithms of the vapor pressures of pure substances. The structures optimized using density functional theory, as well as the experimental temperature and reference normal boiling point, can be used as inputs to the regression model. The loss can be minimized by using mini-batch stochastic gradient descent with the ADAM algorithm. An early stopping process can be used to monitor the training, which runs on the validation set and stops training when the improvement in the prediction error remains below a predefined threshold. The best model can be selected based on the lowest validation error achieved during training and its prediction error evaluated using the test set. The model generated through this process has a prediction error of 21.25% for the vapor pressures of pure substances of substances not seen before, which is accurate enough for screening a large set of corresponding new substances to find potential target substances, and can even be further improved using Figure 8 the screening procedure described.
[0063] Figure 7 Schematically and by way of example, the application of the present invention for optimizing a formulation is shown. In this embodiment, a target formulation that meets one or more predetermined target properties needs to be provided. The formulation comprises at least two (preferably a plurality of) components, and these components can be any substances according to the present invention. In most cases, for example due to the corresponding application requirements, one or more of the components in the target formulation are fixed. However, in this example, there is at least one variable component within the scope of the predetermined application. In particular, a list of a plurality of potential substances that can be used as additional components can be provided, and optionally the corresponding amounts of the potential substances as additional components. The devices and methods described above can be used to determine one or more properties of a formulation comprising a first potential substance as an additional component. In particular, the properties of the substances in the formulation can be determined, and the corresponding relative amounts of the substances can be used to determine the corresponding properties of the formulation. Then, the one or more determined properties can be compared with the corresponding target properties, that is, the corresponding values of the properties are compared with the corresponding values of the target properties. Then, based on the comparison, it can be determined whether the formulation meets the predetermined target, that is, whether it meets one or more target properties. If the formulation meets the target, the current formulation can be determined as the target formulation. If the formulation does not meet the target, a new formulation can be provided by changing the amount of the additional component and / or the current component according to a predetermined limit. In particular, corresponding rules can be used to automatically generate a new formulation based on the previous formulation. For example, the rule can determine: first increase or decrease the amount of the additional component in a predetermined increment until the corresponding predetermined limit is reached; then, starting from a predetermined starting amount, replace this component with a new component in the list of potential components. However, more complex rules can also be used.
[0064] Figure 8A method for improving the accuracy of corresponding property model determination and / or for screening new substances for applications is schematically and exemplarily shown. In this example, first, a corresponding application for the new substance and thus for the property model (i.e., the determination of the corresponding target property) is provided. This intended application also determines the corresponding chemical space of the potential target substances. For example, in human care applications, the chemical space must exclude potentially harmful substances. However, the chemical space can also be defined based on other considerations, such as the availability of the substance, its interaction with other components in the formulation, etc. Based on the application space, a variety of potential new substances covering the application space can be determined. Then, the above-mentioned device and method can be used to very quickly determine the corresponding properties of these potential new substances to screen the application space. In particular, for these potential new substances, atomic descriptors are determined, the corresponding trained classifier and regression algorithms are applied, and the properties are determined accordingly. Thus, the corresponding properties can be compared with the target property, and potential new substances that meet the target property or meet the target property within a predetermined limit can be selected, and these predetermined limits can be set relatively wide at the beginning. Additionally or alternatively, a part of the potential new substances in the application space can be selected based on other criteria (e.g., arbitrarily) or in order to cover the application space to a certain predetermined extent. Then, the selected substances are synthesized. The corresponding application tests or measurements of atomic descriptors can be performed on this small amount of substances. For example, the application properties and / or atomic descriptors of these substances can be measured or derived from the corresponding measured values. Then, these measurement data can be used to update, especially to retrain, the property model. Then, the updated property model can be used again to determine the properties of new substances in the application space. This results in an increase in the corresponding accuracy of the property model, especially for the region of the application space from which the tested substances have been selected. Therefore, when the above steps are iteratively applied, for example, by only screening in subsequent steps the region of the application space where the potential new substances are closest to meeting the target property in the previous step, the property model becomes increasingly accurate in the region where new target substances that meet the target property are expected to exist.
[0065] Further examples and principles of the present invention will be described below. Generally, a property model (e.g., the property model as described above) is adapted to learn the assignment of atoms to different atom classes based on training data that includes the properties of substances and atomic descriptors indicating the atomic structure environment and / or atomic descriptors derived from electronic structure calculations. The suitable atom classes and atomic descriptors for determining a specific property are determined in a purely data-driven manner as part of an overall regression problem. Then, the number and combination of atom classes present in a substance (which can have molecular properties or be an extended periodic material) are used in the regression model to determine the property.
[0066] To improve the generally high predictive power of this method, high-quality training data and suitable structural and computational descriptors can be used. In this case, the atomic classes can be trained more accurately for the property of interest. In addition to the model that directly determines the property, the present invention as described above can also be applied to incremental learning, using a) a physical / engineering material property determination model and b) the proposed regression model to identify and correct the systematic errors of the physical model.
[0067] Preferred embodiments of training a property model using, for example, the above-described training apparatus and method will be described below. Preferably, the training data for establishing a property model includes: a) physicochemical property data and / or application property data that depend on the chemical properties of the corresponding substances; and b) atomic descriptors, i.e., features, derived from the atomic structure environment within these substances and / or obtained from the electronic structure calculations of the corresponding substances.
[0068] The physicochemical properties or application properties for which the property model can be trained can be generally valid or depend on external conditions such as temperature, pressure, composition of the mixture, etc. For example, the properties for which the property model can be trained can refer to physicochemical properties, including one or more of the following: melting point, boiling point, vapor pressure, heat of vaporization, heat capacity, flash point and autoignition temperature, liquid density, critical temperature and pressure, electrical conductivity and thermal conductivity, glass transition temperature, viscosity, surface tension, refractive index, mechanical properties, total energy, dipole moment, polarizability, HOMO-LUMO or band gap, ionization potential, electron affinity, activity coefficient, partition coefficient, solubility (e.g., in various equilibria between solids, liquids, and gases), cloud point, critical micelle concentration, acid and base dissociation constants, hydrolysis stability, corrosivity, catalytic activity for a given chemical reaction, or spectral properties (e.g., transition energies and intensities in IR and UV-Vis spectra). In addition, the properties can refer to application properties, which can refer to: a) general properties beyond pure physicochemistry, such as various types of toxicity and ecotoxicity behavior, biodegradability in different habitats, odor, ozone depletion potential, global warming potential; and / or b) the role as an additive in a mixture, such as the effect on the octane number and cetane number of fuels, stability against UV radiation or oxidation, flame retardancy, the effect on the feel.
[0069] Deriving atomic descriptors from the atomic structural environment can be carried out according to any formalism that can transform the three-dimensional structure and elemental composition of a molecular system into a set of atomic-level descriptors. For example, end-to-end architectures (such as SchNet, PhysNet) can also be trainable and capable of directly learning from the structural data to determine an appropriate set of atomic descriptors, thus being able to adapt to many different chemical systems as long as there is sufficient data. Alternatively, predefined atomic descriptors such as ANI, SOAP, FCHL can be used, which can provide higher accuracy for small data sets. Multiple conformational isomers of a molecule may be relevant for a certain property. To take this into account, atomic descriptors can be derived from multiple conformational isomers and combined through different schemes (such as weighted averaging), where the weights can be the Boltzmann weights of the conformational isomers or weights learned during the training of the property model.
[0070] Atomic descriptors related to atom-specific quantities can be derived from molecular electronic structure calculations and can refer to, for example, one or more of the following: partial charge, exposed surface fraction, average surface shielding charge density calculated by a continuous solvation model, average atomic radius calculated based on the use of a flexible (e.g., isodensity cavity) continuous solvation model, contributions to energy or free energy in a COSMO-RS-type or COSMO-SAC-type solvation model, such as mismatch contributions, hydrogen bond donor contributions, hydrogen bond acceptor contributions, dispersion force contributions or total residual contributions, pairwise contact probabilities from a COSMO-RS-type or COSMO-SAC-type solvation model, NMR shielding constants. For these atomic descriptors, the methods for handling conformational isomers described above can also be applied.
[0071] Then, the trainable property model is adapted to be trained based on the atomic descriptors of the molecules in the training data (such as described above) to assign the atoms of the molecules to different atomic classes, which are then used to train the regression provided by the regression model. Usually, in addition to simple multiple linear regression, the regression model can also adopt more complex strategies such as artificial neural networks (ANNs), kernel methods, support vector regression or random forests. During training, the overall regression problem provided by the training data applied to the property model can preferably be optimized in an iterative manner, during which the optimal parameters / coefficients of the property model and the atomic classes are determined simultaneously. During training, the selection of the atomic descriptors actually used can also be optimized.
[0072] The property model can be adapted to perform atomic category assignment via a classification layer (e.g., softmax). This step provides a significant amount of flexibility. For example, the classifier can be adapted to introduce orthogonality during optimization. Additionally, the classifier can be restricted to select element-specific atomic categories, i.e., an N atom and an O atom will never share the same atomic category. Without this restriction, the classifier is more flexible, has the advantage of using fewer atomic categories, and can also improve transferability, i.e., obtaining a model applicable to elements with only a small amount of data or even no data in the training set. Moreover, the classifier can also be adapted such that an atom can fully belong to one atomic category or have several assignments. In the latter case, the importance of different assignments can be mapped by atomic category weights. For example, the sum of the weights can be forced to add up to one, but this is not necessary. Generally, the higher flexibility in the latter case can be used to describe the potential non-linear effects of different atoms on the given property with different intensities.
[0073] As previously mentioned, for a regression model to determine a target property, simple multiple linear regression can be employed, or more complex strategies such as artificial neural networks (ANNs), kernel methods, support vector regression, or random forests can be used. The training process of the property model, for example, executed by the training unit of the training device, will yield both the assigned atomic categories and the determined values of the target property. Generally, the iterative training ends if the prediction error and optional atomic category constraints (e.g., orthogonality) converge within a given threshold.
[0074] The property model trained according to the above process is capable of determining the target property of a substance based on atomic descriptors, where the atomic descriptors can be derived from the structure of the molecule. This can be a structure obtained through experimental techniques (e.g., X-ray diffraction), but for the proposed model, a more relevant use case is to derive molecular properties from computational structures and descriptors. This means that only calculations are required to determine the target property without experiments, which is a great advantage in the following situations: for example, a) performing virtual screening where the candidates of interest are difficult to purchase or even need to be synthesized; or b) for compounds, although easily obtainable, the properties of interest involve expensive, error-prone, or dangerous measurement processes.
[0075] To perform such determination based on molecular or periodic calculations, the first step is to perform one or several calculations to obtain atomic descriptors, which refer to, for example, the structure of a conformational isomer (usually the most stable conformational isomer under the application conditions) or the structure of a set of related conformational isomers. In the case of a set of conformational isomers, it is also necessary to calculate the conformational weights according to the Boltzmann equation based on the relative energy or free energy (if available). According to the trained atomic descriptors, the atomic descriptors can then be derived, for example, from the corresponding structures. Thus, all input parameters are available for determining the target property through the property model as described above.
[0076] The determined physicochemical properties or application properties can generally be used to accelerate the development of products or processes by significantly shortening the time required to obtain information about how a new compound behaves or how a compound behaves under different conditions. In certain cases, when the total energy of the substance is also learned as a target property, the obtained total energy and its gradient with respect to the nuclear coordinates based on atomic classes can also be used as a calculation method to determine the most probable molecular or periodic structures and their sets mentioned above.
[0077] The general benefits of prediction models that obtain physicochemical properties or application properties solely based on calculations have been given in the above section. Further, specific advantages of the proposed invention are, for example, that there is no need to provide a cumbersome, time-consuming, model-developer-dependent and potentially biased hierarchical assignment of groups. In addition, the atomic classes can be assigned in different ways for each property, which reflects that for different properties, different atoms can play a dominant role. The present invention is generally also applicable to molecules with unclear / ambiguous chemical bonding situations, such as inorganic substances, metal-organic compounds, molecules with non-covalent interactions, where the molecular graph derived by cheminformatics tools is sensitive to small changes in bond lengths and the assignment to hierarchical functional groups is error-prone. In addition to directly determining properties, the present invention is also particularly useful for incremental learning, using a) a physically determined model and b) the proposed atom-class-based model to identify and correct systematic errors in the physical model. In addition, the property model can provide unbiased insights into atomic and functional group similarities, opening up new avenues for explaining how properties are actually realized, thus inspiring ideas for improving these properties.
[0078] By studying the drawings, the present disclosure and the appended claims, those skilled in the art can understand and implement other variations of the disclosed embodiments when practicing the claimed invention.
[0079] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality.
[0080] A single unit or device can implement the functions of several items recited in the claims. The fact that certain measures are recited in mutually different dependent claims does not mean that a combination of these measures cannot be used advantageously.
[0081] Processes such as providing atomic descriptors and properties, providing property models, applying or training property models, etc., which are performed by one or more units or devices, can be performed by any other number of units or devices. These processes can be implemented as program code means of a computer program and / or dedicated hardware.
[0082] A computer program product can be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium provided together with or as part of other hardware, but can also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems.
[0083] Any unit described herein can be a processing unit as part of a classical computing system. The processing unit can include a general-purpose processor and can also include a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or any other specialized circuit. Any memory can be a physical system memory, which can be volatile, non-volatile, or some combination of the two. The term "memory" can include any computer-readable storage medium, such as a non-volatile mass storage device. If the computing system is distributed, the processing and / or storage capabilities can also be distributed. The computing system can include multiple structures as "executable components". The term "executable component" is a structure that is well understood in the computing field to be software, hardware, or a combination thereof. For example, when implemented in software, those of ordinary skill in the art will understand that the structure of an executable component can include software objects, routines, methods, etc. that can be executed on a computing system. This can include executable components in the computing system heap or on a computer-readable storage medium. The structure of an executable component can exist on a computer-readable medium such that when interpreted by one or more processors of the computing system (e.g., by a processor thread), it causes the computing system to perform a function. Such a structure can be directly computer-readable by a processor, for example, as is the case when the executable component is binary, or it can be constructed to be interpretable and / or compilable, for example, whether in a single stage or multiple stages, to generate such binary that can be directly interpreted by the processor. In other cases, the structure can be hard-coded or hard-wired logic gates that are implemented specifically or almost exclusively in hardware, such as within a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or any other specialized circuit. Thus, the term "executable component" is a term for a structure well known to those of ordinary skill in the computing field, whether implemented in software, hardware, or a combination. Any embodiment herein is described with reference to actions performed by one or more processing units of a computing system. If such actions are implemented in software, one or more processors, in response to having executed computer-executable instructions that make up an executable component, direct the operation of the computing system. The computing system can also include a communication channel that allows the computing system to communicate with other computing systems, for example, via a network. A "network" is defined as one or more data links that enable the transfer of electronic data between computing systems and / or modules and / or other electronic devices. When information is transmitted or provided to a computing system via a network or another communication connection (e.g., hard-wired, wireless, or a combination of hard-wired and wireless), the computing system properly treats the connection as a transmission medium. The transmission medium can include a network and / or a data link, which can be used to carry desired program code means in the form of computer-executable instructions or data structures and can be accessed by a general-purpose computing system or a specialized computing system or a combination.While not all computing systems require a user interface, in some embodiments, a computing system includes a user interface system for interacting with a user. The user interface acts as an input or output mechanism for the user, for example, via a display.
[0084] Those skilled in the art will understand that at least part of the present invention can be practiced in a network computing environment having a variety of types of computing system configurations, including personal computers, desktop computers, laptop computers, messaging processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, pagers, routers, switches, data centers, wearable devices (such as glasses), and the like. The present invention can also be practiced in a distributed system environment, where local and remote computing systems linked, for example, by a hardwired data link, a wireless data link, or a combination of hardwired and wireless data links via a network, jointly perform tasks. In a distributed system environment, program modules can be located in both local and remote memory storage devices.
[0085] Those skilled in the art will also understand that at least part of the present invention can be practiced in a cloud computing environment. A cloud computing environment can be distributed, but this is not required. When a cloud computing environment is distributed, it can be distributed across multiple countries within an organization and / or have components across multiple organizations. In this specification and the appended claims, "cloud computing" is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources (such as networks, servers, storage devices, applications, and services). The definition of "cloud computing" is not limited to any one of the many other advantages that can be obtained from such a model. The computing systems in the figures include various components or functional blocks that can implement the various embodiments disclosed herein as explained. The various components or functional blocks can be implemented on a local computing system or on a distributed computing system that includes elements residing in the cloud or implementing aspects of cloud computing. The various components or functional blocks can be implemented as software, hardware, or a combination of software and hardware. The computing systems shown in the figures can include more or fewer components than those shown in the figures, and some of the components can be combined as needed.
[0086] Any reference signs in the claims shall not be construed as limiting the scope.
[0087] The present invention relates to a device for determining the properties of a substance. A providing unit provides atomic descriptors indicative of characteristics of atoms of the substance with respect to a specific molecular structure. A model providing unit provides a trained property model, wherein the trained property model has been trained to determine the properties of the substance as an output when provided with the atomic descriptors. The property model includes: a) an atom classifier adapted to classify atoms of the substance into one or more atom categories based on the atomic descriptors; and b) a regression model adapted to determine the properties based on the atom categories determined for the substance. A determining unit determines the properties of the substance by applying the trained property model to the atomic descriptors.
Claims
1. An apparatus for determining the properties of a substance, the substance being defined as either having molecular properties or exhibiting an extended periodic 1D, 2D or 3D structure, wherein, The device (100) includes: an atomic descriptor providing unit (110) configured to provide atomic descriptors indicative of characteristics of atoms of the substance regarding the structure of the substance, a trained property model providing unit (120) configured to provide a trained machine learning-based property model, wherein the trained property model is adapted to determine a property of the substance as an output when provided with these atomic descriptors as inputs, and wherein the property model includes: a) an atomic classifier adapted to classify atoms of the substance into one or more atomic classes based on these atomic descriptors; and b) a regression model adapted to determine the property based on the atomic classes determined for the substance, a property determining unit (130) configured to determine the property of the substance by applying the trained property model to these atomic descriptors, and a property providing unit configured to provide the property for further processing.
2. The device according to claim 1, wherein, These atomic descriptors include atom-specific quantities derived from the electronic structure of the substance and / or atomic features indicative of the atomic structural environment of atoms of the substance.
3. The device according to any one of claims 1 and 2, wherein The properties of the specific molecule include physicochemical properties and / or application properties, where the physicochemical properties refer to physical or chemical characteristics of the substance composed of the specific molecule, and where the application properties refer to application-specific characteristics of the substance.
4. The apparatus according to any one of the preceding claims, wherein, If the substance has molecular properties and contains molecules of different conformations, these atomic descriptors of the substance refer to the weighted average of atomic descriptors of different conformational isomers.
5. The apparatus according to claim 4, wherein, The weights of the weighted average refer to Boltzmann weights or to trained weights trained together with the property model based on the same training data as the training data used to train the property model.
6. The device according to any one of the preceding claims, wherein, The atomic classifier is adapted to classify atoms of the substance into the one or more atomic classes such that atoms of different elements are classified into different atomic classes.
7. The device according to any one of the preceding claims, wherein, The atomic classifier is adapted to classify atoms of the substance into one or more atomic classes, wherein if an atom is classified into more than one atomic class, atomic class weights are provided to each atomic class to which the atom is assigned based on the probability that the atom belongs to the corresponding atomic class.
8. The apparatus according to any one of the preceding claims, wherein, The regression model refers to any one of multiple linear regression, artificial neural network, kernel algorithm, support vector regression, and random forest algorithm.
9. A training device for training a machine learning-based property model, wherein, The device (300) includes: a training data providing unit (310) configured to provide training data for training the property model, wherein the training data respectively includes atomic descriptors of multiple different substances and corresponding known properties, a trainable property model providing unit (320) configured to provide a machine learning-based property model, wherein the property model includes: a) an atomic classifier trainable to classify atoms of a substance into one or more atomic classes based on atomic descriptors; and b) a regression model trainable to determine the property based on the atomic classes determined for the molecule, and A training unit (330) for training a provided property model based on the training data such that the trained property model is adapted to determine the property of a substance as an output when atomic descriptors of the substance are provided as input.
10. A system for screening potential substances in predefined applications, wherein, The system includes: The apparatus according to claim 1, A target property providing unit for providing a target property of a target substance and potential target substances, and A screening unit, wherein the screening unit is adapted to determine the corresponding properties of the potential target substances using the apparatus, compare the determined properties of the potential target substances with the target property, and perform one of the following operations based on the comparison: i) determine the potential target substance as the target substance, or ii) provide new potential target substances and repeat the determination of the property using the new potential target substances.
11. A computer-implemented method for determining the properties of a substance, the substance being defined as either having molecular properties or exhibiting an extended periodic 1D, 2D or 3D structure, wherein, The method (200) includes: Providing (210) atomic descriptors indicative of characteristics of the atoms of the substance with respect to the structure of the substance, Providing (220) a trained machine learning-based property model, wherein the trained property model is adapted to determine the property of the substance as an output when these atomic descriptors are provided as input, and wherein the property model includes: a) an atom classifier adapted to classify the atoms of the substance into one or more atom classes based on these atomic descriptors; and b) a regression model adapted to determine the property based on the atom classes determined for the substance, Determining (230) the property of the substance by applying the trained property model to these atomic descriptors, and Providing the property for further processing.
12. A computer-implemented method for screening for potential substances in a predefined application, wherein, The method includes: Providing a target property of a target substance and potential target substances, and Determining the corresponding properties of the potential target substances using the method according to claim 11, comparing the determined properties of the potential target substances with the target property, and performing one of the following operations based on the comparison: i) determine the potential target substance as the target substance, or ii) provide new potential target substances and repeat the determination of the property using the new potential target substances.
13. A computer-implemented training method for training a machine learning-based property model, wherein, The method (400) includes: Providing (410) training data for training the property model, wherein the training data respectively include atomic descriptors of a plurality of different substances and corresponding known properties, Providing (420) a machine learning-based property model, wherein the property model includes: a) an atom classifier trainable to classify the atoms of a substance into one or more atom classes based on atomic descriptors; and b) a regression model trainable to determine the property based on the atom classes determined for the substance, and Training (430) the provided property model based on the training data such that the trained property model is adapted to determine the property of the substance as an output when atomic descriptors of the substance are provided as input.
14. A computer program product for training a machine learning-based property model, wherein the computer program product includes program code means for causing a property model training apparatus as claimed in claim 9 to perform the training method according to claim 13.
15. A computer program product for determining the properties of a substance, wherein the computer program product comprises program code means for causing a device as claimed in any one of claims 1 to 8 to perform the method as claimed in claim 11.