Method for screening substances having a required physical property, chemical property or physiological effect from a group of substances using electron densities

EP4681209A1Pending Publication Date: 2026-01-21FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024709779
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-03-13
Filing Date
2024-03-12
Publication Date
2026-01-21

AI Technical Summary

Technical Problem

Current methods for identifying substances with specific physical, chemical, or physiological properties are time-consuming and inefficient, relying on trial-and-error due to unknown structure-property relationships, especially in developing aroma substances or odour-active compounds, as 2D representations fail to capture the spatial domain of chemical substances.

Method used

A method using electron density cube files calculated by semiempirical molecular calculations and a trained artificial neural network to classify and assign substances to desired properties, leveraging electronegativity, electron affinity, molecular softness, or hardness weights, enabling the selection of substances with specific physical, chemical, or physiological effects without exhaustive experimental investigation.

Benefits of technology

This approach significantly reduces the need for experimental verification by precisely capturing the spatial domain of substances, saving time and resources, and enabling the identification of substances with desired properties, such as taste or odor, in a more economical and precise manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000012_0001
    Figure IMGF000012_0001
  • Figure IMGF000013_0001
    Figure IMGF000013_0001
  • Figure IMGF000013_0002
    Figure IMGF000013_0002
Patent Text Reader

Abstract

The present invention provides a method for screening substances having a requested physical property, chemical property or physiological effect from a group of substances comprising the steps • providing a group of k substances by a user, wherein k e N, wherein the molecular weight of each substance is < 400 Da; • providing classification classes or continuous property label for a physical property, chemical property or physiological effect comprising C i classes or V i values, wherein i E N; • calculating electron density cube files for every substance by semiempirical molecular calculation and providing weights Wk, wherein Wk are selected from a group comprising electronegativity cube files, electron affinity cube files, molecular softness cube files and molecular hardness cube files for every substance k; • providing a trained artificial neural network architecture for assigning the k substances to the classification classes or continuous property label, wherein the artificial neural network architecture uses the electron density cube files and weights Wk as input to calculate a tensor of weights, comprising electronegativity, electron affinity, molecular softness or molecular hardness; • assigning each substance to a class C i or a value V i by utilizing the tensor of weights of substance k as weighting for the electronic density cube files in the trained artificial neural network architecture; • displaying and / or outputting of the k substances assigned to the classes C, of the classification or assigned to the values V, of the property label; • selecting the substances assigned to the class C i with the required physical property, chemical property or physiological effect or assigned to the required value Vi of a physical property, chemical property or physiological effect; • experimental verification by a user of the physical property, chemical property or physiological effect of at least part of the selected substances.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method for screening substances having a required physical property, chemical property or physiological effect from a group of substances using electron densities

[0002] The present invention provides a method for screening substances having a requested physical property, chemical property or physiological effect from a group of substances comprising the steps

[0003] • providing a group of k substances by a user, wherein k e N, and wherein the molecular weight of each substance is < 400 Da;

[0004] • providing classification classes or continuous property label for a physical property, chemical property or physiological effect comprising C, classes or V, values, wherein i e N;

[0005] • calculating electron density cube files for every substance by semiempirical molecular calculation and providing weights Wk, wherein Wk are selected from a group comprising electronegativity cube files, electron affinity cube files, molecular softness cube files and molecular hardness cube files for every substance k;

[0006] • providing a trained artificial neural network architecture for assigning the k substances to the classification classes or continuous property label, wherein the artificial neural network architecture uses the electron density cube files and weights Wk as input to calculate a tensor of weights, comprising electronegativity, electron affinity, molecular softness or molecular hardness;

[0007] • assigning each substance to a class C, or a value V, by utilizing the tensor of weights of substance k as weighting for the electronic density cube files in the trained artificial neural network architecture; • displaying and / or outputting of the k substances assigned to the classes C, of the classification or assigned to the values V, of the property label;

[0008] • selecting the substances assigned to the class C, with the required physical property, chemical property or physiological effect or assigned to the required value Vi of a physical property, chemical property or physiological effect;

[0009] • experimental verification by a user of the physical property, chemical property or physiological effect of at least part of the selected substances.

[0010] Chemical substances have chemical properties, physical properties and physiological effects. While physical properties can be quantified by measuring underlying physical quantities, chemical properties can be quantified by measuring an underlying chemical quantity when a substance interacts with another substance. The physical properties of a substance include, for example, boiling point, melting point, volatility and the colour of the substance. The solubility in water, on the other hand, is counted among the chemical properties of a substance.

[0011] Furthermore, chemical substances can have effects on the physiology of organisms (physiological effects). This includes physical and chemical substance properties under the aspect of perceptibility or the effect on an organism. Examples are the smell and the taste of a substance.

[0012] Chemical properties, physical properties and physiological effects are of great interest for a variety of applications. According to the invention, this includes properties such as taste or odour of substances. Physiological effects are induced by properties of substances that have effects on the organism of living beings. Furthermore, according to the invention, this also includes the classification of substances into permitted and non-permitted chemicals in cosmetics and personal care. This is regulated by the use authorization according to Articles Regulation, Annex II - Restricted Substances the Annex II of the European Chemicals Agency (ECHA). The taste of substances for example directly appeals to the human sense of taste and thus decisively influences the eating behavior of humans and in particular which foods are perceived as pleasant or also as unpleasant. The taste produced by substances is therefore of great importance, especially in the food industry.

[0013] Smell is one of the five human senses and plays an important role in daily life. For example, the smell of food influences our eating behavior [1] and smells influence human memory in various evolutionary situations [2], In addition to the importance of odours for humans, they also play an important role in the economy, especially in the food and cosmetics industries, where the development of new flavours and the identification of odour-active substances are essential. In the development of new odourants, a predictive approach during molecular design is required to reduce the space of candidate substances from virtually all to a promising set of structures.

[0014] In the state of the art, sensory-trained experts in industry and science have to smell substances to determine their odour. Due to largely unknown structure-odour relationships, the principle of trial-and-error prevails in the development of aroma substances or the identification of odour-active substances. This is very time-consuming, personnel-intensive, and therefore uneconomical.

[0015] As a result, it is desirable to be able to deduce physical properties, chemical properties or physiological effects of a substance from its structure.

[0016] Machine learning models have already been used to predict certain properties. Traditionally, substances have been represented using InChi [3] (International Chemical Identifier) or SMILES [4] (Simplified Molecular Input Line Entry Specification) notation when used in machine learning tasks to predict properties of substances or to even generate new structures. Another method of encoding molecular structure and features is the SMARTS [5] (SMILES ARbitrary Target Specification) notation. These rule-based methods have the benefit of being rather straightforward to generate and being easy to understand by chemists. Such representations have been used extensively with machine learning in previous works such as those to generate a new molecular structure [6-16], Additionally, other encoding schemes such as molecular graphs have been widely combined with machine learning [8, 17-21] to predict molecular properties [22, 23] such as toxicity [24, 25] or to predict odour of substances [26-30] and even in drugs discovery [31-34] among other applications.

[0017] However, such a 2D representation of chemical substances that are essentially three- dimensional signals does not capture their spatial domain. The SMILES notation, for example, does not capture the similarity of molecules, so similar chemical substances may be encoded differently using very different SMILES strings and there’s no standard method to generate a canonical representation [35, 36], Additionally, there are molecules that cannot be defined by a graph model like those having delocalized bonds such as metal carbonyl complexes

[0035] ,

[0018] Based on the prior art, it is therefore the object of the invention to provide a method by which substances with a required physical property, chemical property or physiological effect can be selected and or generated from a predetermined set of substances without having to examine all the substances with respect to the desired property or effect by experimental methods. This purpose is addressed by the present invention by a method according to claim 1 .

[0019] Detailed Description

[0020] The invention provides a method for screening substances having a requested physical property, chemical property or physiological effect from a group of substances.

[0021] According to the invention a group of k chemical substances is provided by a user, wherein k E N. In one embodiment a group of k molecules is provided. In one embodiment of the present invention, between 20 and 5000 substances are provided, preferably between 100 and 4000 substances are provided, most preferably between 500 and 3000 substances are provided. In this case, provided means first of all that the structural formulae of the chemical substances are available and thus provided. This is possible, for example, by providing the chemical substances in the structural code SMILES, which encodes structural patterns as SMARTS [37-39], In addition, however, it is possible to have each of the chemical substances available physically present at a later time for experimental confirmation. In a preferred embodiment each of the k substances has a molecular weight < 400 Da.

[0022] Further classification classes or a continuous property label is provided according to a chemical property, physical property and / or physiological effect of a chemical substance comprising C, classes or V, values, where i E N.

[0023] If classification classes are given, in one embodiment of the invention, the classification is selected from structure-based properties of substances, in particular from the group comprising odour, taste, colour, toxicity, volatility, water solubility and permitted and nonpermitted chemicals in cosmetics and personal care.

[0024] A classification comprises several classes, for example, classification of water solubility comprises classes hydrophilic and hydrophobic or the classification of the volatility of a compound in high, medium or low volatility classes. The classification of permitted and nonpermitted chemicals in cosmetics and personal care comprises these very classes, namely permitted chemicals in cosmetics and personal care and non-permitted chemicals in cosmetics and personal care. The classification of toxicity comprises classes toxic and non-toxic. The classification of colour may include different colours as classes, for example blue, red, yellow, green. The classification of taste corresponds to different tastes, such as bitter, sour, sweet, salty and umami. The odour classification preferably comprises odour directions such as, but not limited to, 'woody, resinous', 'floral', 'fruity, not citric', 'medicinal', 'perfumed', 'light', 'heavy', 'sweet', 'aromatic', 'fragrant', 'disgusting', ‘odorless’ as classes.

[0025] If a continuous property label is given the property label comprises V; continuous values. Property labels can be for example boiling point, melting point, solubility, aroma thresholds, partition coefficient.

[0026] Further, electron density cube files for every substance are calculated by semiempirical molecular calculation and weights Wk, are provided for every substance. Weights Wk are selected from a group comprising electronegativity cube files, electron affinity cube files, molecular softness cube files and molecular hardness cube files. In one embodiment weights Wk, i.e. , electronegativity cube files, electron affinity cube files, molecular softness cube files or molecular hardness cube files are calculated as well by semiempirical molecular calculation for each of the k substances.

[0027] In a preferred embodiment of the invention, the semiempirical molecular calculation is done in the following way. First of all, the canonical SMILES representation is used as a starting point for each of the k substances. These SMILES are converted to a mol2 format and optimized by using Merck Molecular Force Field (MMFF)

[0040] method in rdkit

[0041] and saved to disk as inp files. These files are then passed into EMPIRE

[0042] software for semi-empirical molecular calculations. For this purpose, the AM1S

[0043] Hamiltonian is chosen and a full geometric optimization in the Cartesian coordinates is performed after centering the molecular structure and orienting it according to its principal axis. As the output, the wavefunction of the selected molecular structure at its electronic ground state is received as a HDF5 wavefunction file, which is then used to generate electron density cube files, electronegativity cube files, electronaffinity cube files, molecular softness cube files and molecular hardness cube files using the eh5cube software from Cepos

[0042] ,

[0028] In a further embodiment of the invention, alternative software for example GAUSSIAN

[0044] is used instead of EMPIRE. Further, alternative approximations instead of DFT or different types of geometric optimizations can be used. Such alternatives are well known in the state of the art.

[0029] The final dataset consists of cube files comprising the electronic density distribution of the k substances along with weights Wk, i.e., electronegativity or electron-affinity cube files, molecular softness cube files or molecular hardness cube files for substance k. In general, the cube files are N dimension tensors consisting of information related to the substances such as but not limited to their electron densities, electron affinity, electronegativities in 3D space.

[0030] A trained artificial neural network architecture for assigning of each of the k substances to a given classification class or continuous property label is provided, wherein the artificial neural network architecture uses the electron density cube files and electronegativity cube files, electron affinity cube files, molecular softness cube files or molecular hardness cube files to assign each of the k substances to their given classes or continuous property label. The classification or property label assignment is independent of the structure description using an artificial neural network but can be performed in one joint training.

[0031] In one embodiment, the artificial neural network architecture is a segmentation network or a Generative Adversarial Network (GAN). Both are open-source networks and known to the person skilled in the art.

[0032] In one embodiment a Binary Cross Entropy Generative Adversarial Network (BCEGAN) or a Wasserstein GAN (WGAN) or a conditional GAN (cGAN) are used to in-paint masked electron density regions. In a preferred embodiment, the Generator is a modified VNet or 3D-UNet and the Discriminator is a 3D Fully Convolutional Neural Network.

[0033] In a preferred embodiment the segmentation network is a modified 3D-UNet, modified VNet or modified Inception architecture, or similar networks.

[0034] The segmentation networks 3D-UNet, VNet and Inception architecture are known per se. Their modification will be described further below.

[0035] According to the invention each substance is assigned to a class C, of the classification or a value Vi of a property label. Therefore, in one embodiment a local map of electronegativity is calculated for each of the k substances by assigning voxels with a certain electronegativity value to a certain electronegativity class. In a further embodiment of the invention a local map of electron affinity is calculated for each of the k substances by assigning voxels with a certain electron affinity value to a certain electron affinity class. The following steps are further described for electronegativity but apply as well for electron affinity, molecular softness and molecular hardness.

[0036] Accordingly, in one embodiment a local map of electronegativity is calculated by assigning voxels with electronegativity values larger than the upper 90 percentile as electronegativity class 2 denoting high reactive sites, whereas voxels with electronegativity values less than the lower 10 percentile are assigned to the electronegativity class 1 , denoting low reactive sites and all other voxels are allocated as electronegativity class 0, denoting a medium reactive region. Hence, these labels denote regions of high, medium, and low reactivity for each of the k substances. In a preferred embodiment, an electronegativity file is created in the same way as an electron density file. While creating the electron density files, EMPIRE creates some more files such as hardness, softness, electron affinity, electronegativity among others.

[0037] The local maps of electronegativity are used as labels which are input in the neural network architecture. Additionally, since for noninert atoms, high electronegativity is correlated to high reactivity

[0045] , in one embodiment a ternary reactivity mask is derived using electronegativity cube files to create labels for this task.

[0038] In a further embodiment electronegativity is not encoded in three electronegativity classes but in two electronegativity classes.

[0039] Further, in one embodiment the electronic density cube file and their corresponding local electronegativity map is illustrated as a substance structure with an overlaid electronegativity map.

[0040] The electron densities included in the electron density cube files and the local map of electronegativity from selective sites of the substance can be used to infer properties of substance such as the classification of substances into classes C, or assigning a substance to a label value V, of physical properties, chemical properties or physiological effects. For this purpose, an adaptive max pooling layer and two fully connected layers are included in the artificial neural network architecture used for the previously described tasks. If a segmentation network such as a 3D-UNet, a VNet or an Inception architecture is basically used the adaptive max pooling layer and two fully connected layers are the modification or at least part of the modification of these segmentation networks. Accordingly, the method of the invention is based on the Density Functional Theory that states that knowing the electron density of a substance allows direct derivation of various properties such as electrostatic potentials, energies, or dipoles of a substance [46, 47],

[0041] The input electronic density is passed through the artificial neural network architecture comprising the adaptive max pooling layer and two fully connected layers to get the structure of the substance as described in a tensor of shape (N, C, B_x, B_y, B_z) along with the tensor of electronegativity which consists of three channels consisting of logits for each of the three classes of electronegativity, which are the raw predicted values from the final layers of the artificial neural network. The tensor of electronegativity is also called electronegativity logit mask in the following. Here, N is the batch size which is similar to the number of substances ; C is the number of elemental classes, corresponding to the different element types in the dataset; (B_x, B_y, B_z) corresponds to the shape of the batch cube, i.e. shape of the overall tensor made by padding all the cubes in a given batch.

[0042] Accordingly, the tensor of electronegativity comprises the spatial distribution of the electronegativity encoded in discrete electronegativity classes, preferably encoded in two or three electronegativity classes.

[0043] To classify a substance into a class C, or to assign a substance to a label value V, belonging to requested physical property, chemical property or physiological effect, the electronegativity logit mask is multiplied with the input electron densities, namely the electron density cube file, and pass the resulting tensor through the adaptive max pooling, batch normalization and fully connected layers. The idea behind this is that the logits of the electronegativity maps, which are comprised in the electronegativity logit mask, serve as weights for the electronic densities and following this with max pooling allows selection of the highest of those weighted densities from within the kernels, allowing to filter the reactive regions in the substance.

[0044] Assignment to class C, is done by using the output produced by the artificial neural network. If Ci belongs to physical properties, the output is of the shape (N, C, B_x, B_y, B_z) and the assignment is performed by using a softmax activation with 0.5 threshold, where the assigned classes C, are present as value 1 and rest as zeros. For other properties, the output is of the shape (N, C) and a sigmoid activation with threshold of 0.5 is used for class assignment. The assigned class C, is present as 1 and the rest as zeros. The assignment is optimised by minimising a loss function between the actual assignment and the assignment from the artificial neural network, e.g. by minimising the Cross Entropy Loss, mean square error loss (MSE) or L1 error.

[0045] Assignment to value V, of a property label is also done by using the output produced by the artificial neural network. If V, belongs to physical properties, the output is of the shape (N, C, B_x, B_y, B_z) and the assignment is performed by using a softmax activation with 0.5 threshold, where the assigned values V, are present as value 1 and rest as zeros. For other properties, the output is of the shape (N, C) and a linear activation with threshold of 0.5 is used for assignment. The assigned values V, are present as continuous values. The assignment is optimised by minimising a loss function between the actual assignment and the assignment from the artificial neural network, e.g. by minimising the Cross Entropy Loss, mean square error loss (MSE) or L1 error.

[0046] According to the invention substances which are assigned to the classes C, or a value VI of the classification are displaying and / or outputting. Displaying can happen on a suitable device such as a display of a PC, tablet or such a like. Outputting can also be a print of the results.

[0047] Further, the substances assigned to the class C, or a value V, with the requested physical property, chemical property or physiological effect are selected and an experimental verification by a user of the physical property, chemical property or physiological effect of at least a part of the selected substances is performed.

[0048] The experimental verification simultaneously verifies the classification of the substance or the assignment to a value V, by a user. The type of experimental verification depends on the classification or property label used. The following table provides a non-exhaustive overview of common experimental methods that can be used to verify physical properties, chemical properties and physiological effects of substances. All other common experimental methods known to those skilled in the art are equally applicable. The present invention thus enables a significant saving in cost, manpower and technical effort, since it is no longer necessary to experimentally investigate all k substances of a provided group in order to select at least one substance of a particular class C, or with a particular value Vi and thus with a particular physical property, chemical property or physiological effect. By applying the inventive method, a selection of substances is made and the subsequent experimental verification can be carried out specifically with this selection of substances. This saves time and costs compared to state-of-the-art methods. Moreover, it is not necessary to have all substances available for experimental investigations, which saves additional costs.

[0049] Advantageously, the inventive method enables to take into account the spatial domain of the substances under investigation. Thus, the calculation becomes more precise compared to state of the art methods.

[0050] To in-paint missing regions, training data consisting of original electron density are created using semi-emperical molecular calculation and additionally electron density masks are created that consists of selecting a region of (j, i, i) voxels where i e N around one atomic center and replacing the electron density in this region to a fixed value such as ones or zeroes or electron density values selected from the neighborhood of another atomic center.

[0051] In one embodiment, the neural network architecture is used to describe the structure of each of the k substances. The generated electron density cube files and electronegativity cube files, electron affinity cube files, molecular softness cube files or molecular hardness cube files for each substance are considered as a stack of images over an arbitrary dimension and calculating the structure of the substance by localizing atomic centers corresponds to the problem of assigning different atomic classes to the pixels of these stacked images, where most of the pixels belong to a so called “background class” creating sparsity, and contributing to a clear imbalance in the atomic classes is the abundance of elements in nature and in the datasets. The pixels of the stacked images are also called voxels in the following. Mostly the distribution of elements shows a clear imbalance towards elements that occur more frequently in nature such as hydrogen and carbon. These elements are then associated with an atomic class number n, where n e N and with atomic class 0 being assigned to the background voxels.

[0052] Further, for each of the k substances a label for calculating its structure is created. Such a label is preferably a tensor of the same shape as the corresponding electronic density cube file. Then, the voxels corresponding to the atomic center of the element are assigned the value corresponding to its atomic class. All other voxels are set to zero. For example, voxel locations corresponding to element hydrogen were set to value 1 , those corresponding to carbon were set to value 6 and likewise for all other relevant elements. The tensor generated in such an embodiment at the output consists of these atomic numbers at voxel locations and thus describes the structure of each of the k substances.

[0053] Training

[0054] According to the invention the artificial neural network architecture assignment of substances to classification classes is trained for a defined classification by a training dataset, wherein the training dataset consists of substances with known assignment to the classes C, of the classification or with known assignment to a value V, of a property label. In one embodiment of the invention the substances of the training dataset have a molecular weight < 400 Da.

[0055] The training of the artificial neural network is performed in two steps. First step is to investigate if the chosen artificial neural network architecture and the labeling approach can be used to learn meaningful information about the structure of a substance. As the second step, the second set of labels is input, namely, the electronegativity maps or electron affinity maps to jointly train the artificial neural network architecture to learn and predict the structure and the corresponding electronegativity maps or electron affinity maps simultaneously.

[0056] In a preferred embodiment the artificial neural network architecture is trained using PyTorch

[0050] library for Python. The hyper-parameters for both trainings are selected by performing a grid search over a common set of hyper-parameters.

[0057] The loss function for the classes of the structure is chosen as cross-entropy loss, mean square error loss (MSE) or L1 error between the actual voxels for a given class and the predicted locations for the same class and class weights can be used to counter the class imbalance. For the prediction of electronegativity maps or electron affinity maps, dice loss

[0051] is calculated between the ground truth maps and the predictions from the network and the loss function for joint training of the two tasks is the sum of the cross-entropy and dice loss. To evaluate the performance of the artificial neural network architecture, three different metrics can be used such as, L1 error, Ratio Of Prediction and dice coefficient between the ground truth and predictions.

[0058] Consider some substance M with n atomic centers defined by cartesian coordinates (x,y,z) corresponding to some class, say hydrogen as shown: substance, and this class, let there be m predictions shown The structure prediction task is to predict the n positive spikes in a sparse 3D space consisting of mostly zeros.

[0059] It can be observed that in the calculation n m and (X y^Zj) #= (x'i,y'i,z'i) along one or more dimensions indicating over or under prediction along with errors in the exact atomic center voxel locations. Concretely, often not only the actual location for the atomic center is predicted by the calculation but their immediate neighbors are often falsely calculated as well.

[0060] Thus, quantifying the performance of the artificial neural network architecture using metrics such as simple accuracies by comparing exact locations does not give the full picture. For this reason, preferably three metrics that can capture the performance of the artificial neural network architecture are calculated.

[0061] Firstly, a metric called the Ratio Of Prediction (ROP) is calculated. ROP is defined as the ratio of the number of predicted voxels for a given atomic class to the actual number of voxels for that atomic class. This helps establishing a relationship between n and m from the example above. A value of ROP > 1 indicates overprediction of the number of atomic centers, i.e. , m > n, whereas ROP < 1 indicates underprediction of the number of atomic centers, i.e., m < n.

[0062] For example, if atomic class 1 (hydrogen) has 100 atoms in a batch of 20 substances, and the model predicts 90 voxels for that atomic class in the same batch, the ROP value is then 0.9 for the batch. The ROP values are summed up and then averaged across the batch size. For an ideal prediction, the ROP is expected to approach 1 , indicating that the artificial neural network architecture calculates the correct number of atoms for a given atomic class. Additionally, the dice-coefficient between the predicted tensor and the ground truth tensor for each element class number n is calculated. This quantifies the ability of the artificial network to segment the different atomic classes. Since it’s possible that some atomic classes n occur only in a small subset of a dataset, these dice values are calculated when there exists at least one voxel corresponding to that atomic class n in a batch before being averaged over the batches that the atomic class n occurs in.

[0063] Finally, the L1 error per atomic class n between the predicted voxel locations against the actual voxel locations is calculated as well. This helps to quantify the error in the predicted locations of the atomic centers. The L1 error, however, will not allow differentiation between errors in the counts due to erroneously calculated voxels and those due to mislocated atomic center voxels, hence we use the three metrics in conjunction to evaluate the overall performance.

[0064] To evaluate the performance of the artificial neural network architecture in calculating the electronegativity maps or electron affinity maps, dice coefficients for each of the three or two electronegativity classes of the electronegativity maps or dice coefficients for each of the three or two electron affinity classes of the electron affinity maps can be calculated.

[0065] In one embodiment in the training data set there occurs a clear imbalance between substances belonging to different classes of a classification. For example, there are overall 855 substances from the ‘Allowed / Non-Hazardous’ class, and 501 substances from the ‘Prohibited / Health Hazardous’ class that can be also considered as toxic for cosmetic usage. In a preferred embodiment this class imbalance is combated by introducing a class weight in the binary cross entropy loss equal to the inverse ratio of occurrence of the two classes.

[0066] According to the invention 3D electronic densities are used as training data for an artificial neural network architecture to allow capturing of the spatial nature of substances. Additionally, the electron densities included in the electron density cube files and local electronegativity maps or electron affinity maps from selective sites of the substance are used to infer properties of substances such as the classification of substances into classes C, or assignment of a substance to a value V, of a property label of physical properties, chemical properties or physiological effects.

[0067] Thus, it is demonstrated that electronic densities can be used as training data for artificial neural network architectures to infer the structure of the underlying chemical substances. Additionally, the local electronegativity maps or local electron affinity maps of chemical compounds are segmented and can be used to identify sites of high and low electronegativity or electron affinity. These are commonly considered as active sites where reactions can take place and together with electron densities can be used in classification C, of compounds e.g. into chemical substances that are allowed in cosmetics i.e. , are not hazardous and those that are health hazards and hence prohibited.

[0068] The overall problem of using molecular electronic densities to classify a substance as hazardous for cosmetic use can be divided into three subtasks. Firstly, the underlying substance itself is inferred, i.e., the artificial neural network predicts the type of element and localizes the atomic centers of the substance based on the cartesian coordinates of its atomic center extracted from the cube file. The invention serves as a proof of concept that an artificial neural network architecture can learn meaningful information using the electron densities. Secondly, the artificial neural network architecture can simultaneously also calculate the electronegativity maps or electron affinity maps for the given compound to predict regions of high, medium, and low reactivity and finally, this information is suitable to classify substances into classes C, of requested physical properties, chemical properties or physiological effects. However, the tasks of structure prediction and classification into C, classes can be performed independently.

[0069] In one embodiment for describing the structure of each of the k substances, where the artificial neural network is a GAN, such as a BCE-GAN

[0052] or a WGAN

[0053] , the masked density cubes, i.e., those where the atomic centers were removed and replaced with a mask are passed through the Generator architecture to generate realistic electron density regions. While the discriminator is a 3D FCN that predicts if the final output is real or fake. The training is performed using the entire ECHA dataset, where 1686 substances belonged to the training class, while 183 belonged to the test class. These were optimized using Adam optimizer and a regression loss

[0054] , The regression loss defined in

[0054] was modified to consider the local differences of electron density regions for our use-case.

[0070] The evaluation of the in-painted region is done in two methods, firstly, the generated region is passed through the previously trained 3D-UNet to identify the tensor describing the spatial structure of the molecule and numerically by using the Dice score and Frechet Inception Distance (FID)

[0055] between the original electron densities and the generated densities. The initial approach allows identification of atomic center types and locations to verify if the new generated density regions have same or different atom types predicted. While FID is a numerical approach to quantify the similarity in the ground truth and predicted regions.

[0071] FID evaluation, however, requires a feature extractor. For this purpose, a 3D-ResNet10

[0056] was selected, written in Pytorch and trained to optimize for classification of electron densities into allowed or prohibited classes on the ECHA dataset without any electronegativity cube files.

[0072] In the following the method of the invention is further described by 6 figures and 4 examples.

[0073] Figure 1 illustrates electronic density (A) and the corresponding electronegativity map (B) for the molecule CC(=O)OC1CC2CC1C3C2CCC3;

[0074] Figure 2 illustrates electronegativity regions for 3 molecules (A) from a test set along with their corresponding calculated segmentation result (B);

[0075] Figure 3 (A) and (B) illustrate Confusion matrix for classification; Figure 4 (A) illustrates a native 3D Linet and (B) a modified 3D Linet;

[0076] Figure 5 (A) illustrates the electron density region ground truth for molecule

[0077] CC(C)CCCCCCCCCCCCCCCOC(=O)C[C@H](C(=O)OCCCCCCCCCCCCC CCC(C)C)O and (B) shows the in-painted region overlayed on the original electron density;

[0078] Figure 6 (A) illustrates the electron density region ground truth for molecule

[0079] C[C@H](COC(=O)c1ccccc1)OC(=O)c1ccccc1 and (B) shows the in-painted region overlayed on the original electron density.

[0080] Figure 1 to 3, as well as 5 and 6 are further explained in connection with the examples.

[0081] Figure 4 (A) illustrates a native 3D Linet as it is known to persons skilled in the art. Figure 4 (B) illustrates a modified 3D Linet according to the invention. The additional MaxPool and the two additional fully connected layers are shown.

[0082] As shown in Figure 4 (B), the 3D Linet can be used for classification into C, classes or for describing the structure of each of the k substances. These are independent tasks, but they can be performed at the same time as well by jointly optimising their loss functions. In case of property prediction, the output is the class or continuous property label, in case of structure description, the output is a tensor with each voxel marked as the predicted atomic class, such as 1 for Hydrogen, 8 for Oxygen, 0 for background and so on.

[0083] Figure 5 and 6 illustrate the results from a GAN network as is known to persons skilled in the art. Here, a region of the electron density cube is masked and the artificial neural network is trained to in-paint the missing region. The mask can be set as a fixed value such as ones, zeros or randomly selected electron densities from a different substance. This is shown in 5 (A) and 6 (A) as inputs and 5(B) and 6(B) as results.

[0084] Example 1 - calculation of structure and electron density

[0085] The structure and electron density for the molecule CC(=O)OC1CC2CC1C3C2CCC3 have been calculated by the method of the present invention.

[0086] Table 1: Shows the distribution of elements in a set of k molecules which included the molecule CC(=O)OC1CC2CC1C3C2CCC3. Structure and electronegativity map were calculated according to the invention.

[0087] Figure 1 illustrates the calculated electronic density (A) and the corresponding electronegativity map (B) for the molecule CC(=O)OC1CC2CC1C3C2CCC3. The electronegativity values have been overlaid on the molecular structure and then divided into 3 electronegativity classes. The shaded region in figure 1 (B) shows regions of high electronegativity and hence these voxels are marked as electronegativity class 2. Black regions with white stripes show regions of low electronegativity and these voxels are marked as electronegativity class 1. All other voxels are marked as electronegativity class 0 and shown as white.

[0088] Example 2 - Training of artificial neural network architecture

[0089] Table 2 shows the metrics described for training the artificial neural network architecture to calculate the structure of a substance averaged over a test set. It is observed that for almost all atomic classes n, except for Sulphur (class 8), the artificial neural network architecture calculates more voxels than present in the ground truth. This is especially the case for class 6, silicon. Table 2: Three metrics were calculated for evaluating the performance of the artificial neural network architecture for structure prediction using electronic densities as input and centers of atoms as labels. The L1 error is considerably high for atomic class 6 (silicon). This is also seen in its low Dice coefficient score.

[0090] For this purpose, the average molecular weights and the average fraction of sp3 (fsp3) carbons [48,49] for all molecules belonging to all the different element types were compared to quantify the complexity of these elemental classes n. This is summarized in Table 3. It can be observed that molecules containing Si on average have a higher fsp3 count than other molecules and a comparatively high mean molecular weight. This could be indicative of higher complexity of these molecules causing the artificial neural network architecture to make false predictions for this particular class.

[0091] Table 3: The mean fraction of sp3 carbons and the mean molecular weight calculated for each element type present in the dataset. The high fsp3 and mean molecular weight suggests that the dataset samples containing silicon consists of more ‘complicated’ molecular structures than other classes.

[0092] Table 4 shows the metrics described for the joint prediction of structure and electronegativity maps, namely, Ratio Of Prediction, L1 error, dice coefficient for structure prediction and the dice coefficient for segmentation of the electronegativity maps. Overall, there is a similar performance compared to the single structure learning task. For this machine learning task, again, silicon, atomic class 6 performs worse than the other atomic classes and this is consistent with that was observed previously. Additionally, Table 5 also shows the segmentation results for the 3 types of electronegativity maps as average class wise dice coefficient values over the test set.

[0093] Table 4: Metrics for the joint learning of molecular structure and the prediction of electronegativity maps. The average metrics obtained are similar to the singular structure prediction task. However, see atomic class 6, silicon being the least accurately predicted class again.

[0094] Table 5: Dice coefficients values were used to evaluate the overlap of the electronegativity map ground truth to the prediction. Atomic class 1 consisted of the low reactive sites, atomic class 2 consisted of the high reactive sites and atomic class 0 are all the other voxels. The atomic class imbalance between the three atomic classes can be also seen in the results with low dice coefficient for the minority classes.

[0095] To visually inspect the results of the electronegativity map calculation, three molecules were chosen randomly from the test set and their predicted electronegativity maps were overlaid on their 3D electron density structures. These are shown in figure 2 with the ground truths shown on the left (A) and the calculated regions shown on the right (B). The SMILES strings of the illustrated molecules are CCOC1(CCC2C(=C)C1CCC2(C)C)C, CC(CC(C)(C)O)O, and [C@@H]([C@H](C(=O)O)O)(C(=O)O)O respectively. The shaded, black region with white stripes and white colored regions are those marked as centers of high, low and medium electronegativity respectively, with the same interpretation as in figure 1.

[0096] Visually, the segmentation results also show what was observed in the average dice coefficients shown in Table 5. While there seems to be an overlap between the ground truth voxels and their predicted segmentation results, however, the electronegativity class imbalance between the medium reactive sites and the other two electronegativity classes can also be seen especially around the ‘edges’ of those high and low reactive regions.

[0097] Example 3 - Classification

[0098] A dataset of overall 1869 substances consisting of ‘Allowed / Non-Hazardous’ class, with 1063 samples, and 632 molecules from the ‘Prohibited / Health Hazardous’ class that can be also considered as toxic for cosmetic usage was given. After training for 38 epochs with a learning rate of 8.44e-4 and rate decay after 35 epochs and a weight decay of 2.57e-7 on Adam optimizer, an accuracy of 78.1% on the test set was achieved denoting that this strategy can be used to classify substances into health hazardous and non-hazardous classes.

[0099] The corresponding confusion matrix is shown in figure 3. According to the purpose of the invention one is more interested in not only the overall accuracy but also in the false negatives for class 0, i.e., substances originally marked as not allowed or health hazardous in the data but predicted as allowed.

[0100] To combat the class imbalance, in a further step a class weight in the binary cross entropy loss equal to the inverse ratio of occurrence of the two classes was introduced and generalized Dice loss was used instead of regular Dice loss.

[0101] Example 4 - GAN training

[0102] The ECHA dataset consisting of 1869 data points, from which 1686 electron density files were used for training and 183 for testing. The training data consisted of 1063 as allowed and 632 as prohibited samples. For the test set, 115 files were classified as allowed and 68 as prohibited.

[0103] For training of the feature extractor, that would be used for calculating the FID score, a 3D- ResnetlO was trained using the training data and then tested on the test set. It achieved a test accuracy of 68%. To generate the occluded / masked labels, random atomic centers were first selected and electron density values around the atomic centers of the size (9,9,9) and (17,17,17) voxels was set to electron density values from a random atomic center from a second randomly chosen sample. For the training set, the random sample was chosen from train set and for the test set, the random sample was chosen from the test set. Training was performed for 70 epochs and a smaller mask region outperformed the bigger mask, which makes sense as the network has more information to extrapolate and learn from. To evaluate the performance, FID scores were calculated using the feature extractor trained on ECHA dataset and additionally Dice scores were calculated for the test set samples. Two examples from the ground truth and generated data are shown in Figure 5 and 6 overlayed with the original electron density. Table 6: FID score and Dice score for the generated and original ECHA test set. Lower FID score is desired while higher Dice score is better

[0104] References

[0105] [1] a) L. G. Fine, C. E. Riera, Frontiers in physiology 2019, 10, 1151 ; b) P. Morquecho- Campos, K. de Graaf, S. Boesveldt, Food quality and preference 2020, 85, 103959.

[0106] [2] J. E. Taylor, H. Lau, B. Seymour, A. Nakae, H. Sumioka, M. Kawato, A. Koizumi, Frontiers in Neuroscience 2020, 14, 255.

[0107] [3] Heller, S. R.; McNaught, A.; Pletnev, I.; Stein, S.; Tchekhovskoi, D. InChi, the IIIPAC International Chemical Identifier. J. Cheminform.

[0108] [4] Anderson, E.; Veith, G. D.; Weininger, D. SMILES: A Line Notation and Computerized Interpreter for Chemical Structures.; 1987.

[0109] [5] Daylight Theory: SMARTS - A Language for Describing Molecular Patterns. 2012, pp 1-7.

[0110] [6] Jin, W.; Barzilay, R.; Jaakkola, T. Chapter 11 : Junction Tree Variational Autoencoder for Molecular Graph Generation. RSC Drug Discov. Ser. 2021 , 2021-Janua (75), 228- 249. https: / / doi.Org / 10.1039 / 9781788016841-00228.

[0111] [7] Takeda, S.; Hama, T.; Hsu, H.-H.; Yamane, T.; Masuda, K.; Piunova, V. A.; Zubarev, D.; Pitera, J.; Sanders, D. P.; Nakano, D. Al-Driven Inverse Design System for Organic Molecules. 2020.

[0112] [8] De Cao, N.; Kipf, T. MolGAN: An Implicit Generative Model for Small Molecular Graphs. 2018.

[0113] [9] Gomez-Bombarelli, R.; Wei, J. N.; Duvenaud, D.; Hernandez-Lobato, J. M.; Sanchez- Lengeling, B.; Sheberla, D.; Aguilera-lparraguirre, J.; Hirzel, T. D.; Adams, R. P.; Aspuru-Guzik, A. Automatic Chemical Design Using a Data-Driven Continuous Representation of Molecules. ACS Cent. Sci. 2018, 4 (2), 268-276. https : / / doi . org / 10.1021 / acscentsci .7b00572.

[0114]

[0010] Mendez-Lucio, O.; Baillif, B.; Clevert, D.-A.; Rouquie, D.; Wichard, J. De Novo Generation of Hit-like Molecules from Gene Expression Signatures Using Artificial Intelligence. Nat. Commun. 2020, 11 (1), 10. https: / / doi.org / 10.1038 / s41467-019- 13807-w.

[0115]

[0011] Skalic, M.; Jimenez, J.; Sabbadin, D.; De Fabritiis, G. Shape-Based Generative Modeling for de Novo Drug Design. J. Chem. Inf. Model. 2019, 59 (3), 1205-1214. https: / / doi.org / 10.1021 / acs.jcim.8b00706.

[0116]

[0012] Bjerrum, E. J.; Threlfall, R. Molecular Generation with Recurrent Neural Networks (RNNs). 2017.

[0117]

[0013] Jastrz^bski, S.; Lesniak, D.; Czarnecki, W. M. Learning to SMILE(S). 2016, 1-5.

[0118]

[0014] Arus-Pous, J.; Patronov, A.; Bjerrum, E. J.; Tyrchan, C.; Reymond, J. L.; Chen, H.; Engkvist, O. SMILES-Based Deep Generative Scaffold Decorator for de-Novo Drug Design. J. Cheminform. 2020, 12 (1), 1-32. https: / / doi.org / 10.1186 / s13321-020-00441- 8.

[0119]

[0015] Zhavoronkov, A.; Ivanenkov, Y. A.; Aliper, A.; Veselov, M. S.; Aladinskiy, V. A.; Aladinskaya, A. V; Terentiev, V. A.; Polykovskiy, D. A.; Kuznetsov, M. D.; Asadulaev, A.; Volkov, Y.; Zholus, A.; Shayakhmetov, R. R.; Zhebrak, A.; Minaeva, L. I.; Zagribelnyy, B. A.; Lee, L. H.; Soil, R.; Madge, D.; Xing, L.; Guo, T.; Aspuru-Guzik, A. Deep Learning Enables Rapid Identification of Potent DDR1 Kinase Inhibitors. Nat. Biotechnol. 2019, 37 (9), 1038-1040. https: / / doi.org / 10.1038 / s41587-019-0224-x.

[0120]

[0016] Popova, M.; Isayev, O.; Tropsha, A. Deep Reinforcement Learning for de Novo Drug Design. Sci. Adv. 2022, 4 (7),

[0121]

[0017] Liu, Q.; Allamanis, M.; Brockschmidt, M.; Gaunt, A. L. Constrained Graph Variational Autoencoders for Molecule Design. Adv. Neural Inf. Process. Syst. 2018, 2018-Decem (NeurlPS), 7795-7804.

[0122]

[0018] Li, Y.; Vinyals, O.; Dyer, C.; Pascanu, R.; Battaglia, P. Learning Deep Generative Models of Graphs. 2018.

[0123]

[0019] You, J.; Ying, R.; Ren, X.; Hamilton, W. L.; Leskovec, J. GraphRNN: Generating Realistic Graphs with Deep Auto-Regressive Models. 35th Int. Conf. Mach. Learn. ICML 20182018, 13, 9072-9081.

[0124]

[0020] Ma, T.; Chen, J.; Xiao, C. Constrained Generation of Semantically Valid Graphs via Regularizing Variational Autoencoders. Adv. Neural Inf. Process. Syst. 2018, 2018- Decem (Nips), 7113-7124.

[0125]

[0021] Shervani-Tabar, N.; Zabaras, N. Physics-Constrained Predictive Molecular Latent Space Discovery with Graph Scattering Variational Autoencoder. arXiv 2020, 1-42.

[0126]

[0022] Walters, W. P.; Barzilay, R. Applications of Deep Learning in Molecule Generation and Molecular Property Prediction. Acc. Chem. Res. 2021, 54 (2), 263-270. https: / / doi.org / 10.1021 / acs.accounts.0c00699.

[0127]

[0023] Hirohara, M.; Saito, Y.; Koda, Y.; Sato, K.; Sakakibara, Y. Convolutional Neural Network Based on SMILES Representation of Compounds for Detecting Chemical Motif. BMC Bioinformatics 2018, 19 (19), 526. https: / / doi.org / 10.1186 / s12859-018- 2523-5.

[0128]

[0024] Mayr, A.; Klambauer, G.; Unterthiner, T.; Hochreiter, S. DeepTox: Toxicity Prediction

[0129] Using Deep Learning. Front. Environ. Sci. 2016, 3 (FEB). https: / / doi.org / 10.3389 / fenvs.2015.00080.

[0130]

[0025] Suzuki, T.; Katouda, M. Predicting Toxicity by Quantum Machine Learning. J. Phys. Commun. 2020, 4 (12), 1-30. https: / / doi.org / 10.1088 / 2399-6528 / abd3d8.

[0026] Sanchez-Lengeling, B.; Wei, J. N.; Lee, B. K.; Gerkin, R. C.; Aspuru-Guzik, A.; Wiltschko, A. B. Machine Learning for Scent: Learning Generalizable Perceptual Representations of Small Molecules. 2019.

[0131]

[0027] Keller, A.; Gerkin, R. C.; Guan, Y.; Dhurandhar, A.; Turu, G.; Szalai, B.; Mainland, J.

[0132] D.; Ihara, Y.; Yu, C. W.; Wolfinger, R.; Vens, C.; Schietgat, L.; De Grave, K.; Norel, R.; Stolovitzky, G.; Cecchi, G. A.; Vosshall, L. B.; Meyer, P.; Bhondekar, A. P.; Boutros, P. C.; Chang, Y. C.; Chen, C. Y.; Cherng, B. W.; Dimitriev, A.; Dolenc, A.; Falcao, A. O.; Golihska, A. K.; Hong, M. Y.; Hsieh, P. H.; Huang, B. F.; Hunyady, L.; Kaur, R.; Kazanov, M. D.; Kumar, R.; Lesinski, W.; Lin, X.; Matteson, A.; Oyang, Y. J.; Panwar, B.; Piliszek, R.; Polewko-Klim, A.; Raghava, G. P. S.; Rudnicki, W. R.; Saiz, L.; Sun, R. X.; Toplak, M.; Tung, Y. A.; Us, P.; Varnai, P.; Vilar, J.; Xie, M.; Yao, D.; Zitnik, M.; Zupan, B. Predicting Human Olfactory Perception from Chemical Features of Odor Molecules. Science (80). 2017, 355 (6327), 820-826. https: / / doi.org / 10.1126 / science.aal2014.

[0133]

[0028] Genva, M.; Kemene, T. K.; Deleu, M.; Lins, L.; Fauconnier, M. L. Is It Possible to Predict the Odor of a Molecule on the Basis of Its Structure? Int. J. Mol. Sci. 2019, 20 (12). https: / / doi.org / 10.3390 / ijms20123018.

[0134]

[0029] Ldtsch, J.; Kringel, D.; Hummel, T. Machine Learning in Human Olfactory Research. Chem. Senses 2019, 44 (1), 11-22. https: / / doi.org / 10.1093 / chemse / bjy067.

[0135]

[0030] Shang, L.; Liu, C.; Tomiura, Y.; Hayashi, K. Machine-Learning-Based Olfactometer:

[0136] Prediction of Odor Perception from Physicochemical Features of Odorant Molecules. Anal. Chem. 2017, 89 (22), 11999-12005. https: / / doi.Org / 10.1021 / acs.analchem.7b02389.

[0137]

[0031] Cangea, C.; Grauslys, A.; Lid, P.; Falciani, F. Structure- Based Networks for Drug Validation. 2018, 1-5.

[0138]

[0032] Sakai, M.; Nagayasu, K.; Shibui, N.; Andoh, C.; Takayama, K.; Shirakawa, H.; Kaneko,

[0139] S. Prediction of Pharmacological Activities from Chemical Structures with Graph Convolutional Neural Networks. Sci. Rep. 2021 , 11 (1), 525. https: / / doi.Org / 10.1038 / S41598-020-80113-7.

[0140]

[0033] Gaudelet, T.; Day, B.; Jamasb, A. R.; Soman, J.; Regep, C.; Liu, G.; Hayter, J. B. R.; Vickers, R.; Roberts, C.; Tang, J.; Roblin, D.; Blundell, T. L.; Bronstein, M. M.; Taylor- King, J. P. Utilizing Graph Machine Learning within Drug Discovery and Development. Brief. Bioinform. 2021 , 22 (6), bbab159. https: / / doi.org / 10.1093 / bib / bbab159.

[0141]

[0034] Jiang, D.; Wu, Z.; Hsieh, C.-Y.; Chen, G.; Liao, B.; Wang, Z.; Shen, C.; Cao, D.; Wu, J.; Hou, T. Could Graph Neural Networks Learn Better Molecular Representation for Drug Discovery? A Comparison Study of Descriptor-Based and Graph-Based Models. J. Cheminform. 2021 , 13 (1), 12. https: / / doi.org / 10.1186 / s13321-020-00479-8.

[0035] David, L.; Thakkar, A.; Mercado, R.; Engkvist, O. Molecular Representations in Al- Driven Drug Discovery: A Review and Practical Guide. J. Cheminform. 2020, 12 (1), 1- 22. https: / / doi.Org / 10.1186 / S13321 -020-00460-5.

[0142]

[0036] O’Boyle, N. M. Towards a Universal SMILES Representation - A Standard Method to Generate Canonical SMILES Based on the InChi. J. Cheminform. 2012, 4 (1), 22. https: / / doi.Org / 10.1186 / 1758-2946-4-22.

[0143]

[0037] D. Weininger, Journal of chemical information and computer sciences 1988, 28, 31.

[0144]

[0038] Daylight Chemical Information Systems, Inc., "3. SMILES - A Simplified Chemical

[0145] Language", can be found under https: / / www.daylight.com / dayhtml / doc / theory / theory.smiles.html, 2019.

[0146]

[0039] Daylight Chemical Information Systems, Inc., "4. SMARTS - A Language for Describing

[0147] Molecular Patterns", can be found under https: / / www.daylight.com / dayhtml / doc / theory / theory.smarts.html, 2019.

[0148]

[0040] Halgren, T. A. Performance of MMFF94*. Scope, Parameterization, J. Comput. Chem. 1996, 17, 490-519.

[0149]

[0041] RDKit: Open-Source Cheminformatics, http: / / www.rdkit.org / .

[0150]

[0042] Empire & EH5cube. https: / / www.ceposinsilico.de / products / empire.htm.

[0151]

[0043] Dewar, M. J. S.; Zoebisch, E. G.; Healy, E. F.; Stewart, J. J. P. Development and Use of Quantum Mechanical Molecular Models. 76. AM 1 : A New General Purpose Quantum Mechanical Molecular Model. J. Am. Chem. Soc. 1985, 107 (13), 3902-3909. https: / / doi.org / 10.1021 / ja00299a024.

[0152]

[0044] Gaussian 16: https: / / gaussian.com / gaussian16 / ; https: / / gaussian.com / dft /

[0153]

[0045] Nordholm, S. From Electronegativity towards Reactivity-Searching for a Measure of

[0154] Atomic Reactivity. Molecules 2021 , 26 (12). https: / / doi.org / 10.3390 / molecules26123680.

[0155]

[0046] Lewis, A. M.; Grisafi, A.; Ceriotti, M.; Rossi, M. Learning Electron Densities in the Condensed Phase. J. Chem. Theory Comput. 2021 , 17 (11), 7203-7214. https: / / d0i.0rg / l 0.1021 / acs.jctc.1 C00576.

[0156]

[0047] Parr, R. G.; Weitao, Y. Density-Functional Theory of Atoms and Molecules. Oxford

[0157] University Press January 5, 1995. https: / / doi.Org / 10.1093 / OSO / 9780195092769.001.0001 .

[0158]

[0048] Lovering, F.; Bikker, J.; Humblet, C. Escape from Flatland: Increasing Saturation as an Approach to Improving Clinical Success. J. Med. Chem. 2009, 52 (21), 6752-6756. https: / / doi.org / 10.1021 / jm901241e.

[0159]

[0049] Mendez-Lucio, O.; Medina-Franco, J. L. The Many Roles of Molecular Complexity in

[0160] Drug Discovery. Drug Discov. Today 2017, 22 (1), 120-126. https: / / doi.Org / 10.1016 / j.drudis.2016.08.009.

[0050] Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; Desmaison, A.; Kbpf, A.; Yang, E.; DeVito, Z.; Raison, M.; Tejani, A.; Chilamkurthy, S.; Steiner, B.; Fang, L.; Bai, J.; Chintala, S. PyTorch: An Imperative Style, High-Performance Deep Learning Library. Adv. Neural Inf. Process. Syst. 2019, 32 (NeurlPS).

[0161]

[0051] Sudre, C. H.; Li, W.; Vercauteren, T.; Ourselin, S.; Jorge Cardoso, M. Generalised Dice Overlap as a Deep Learning Loss Function for Highly Unbalanced Segmentations. Leet. Notes Comput. Sci. (including Subser. Leet. Notes Artif. Intell. Leet. Notes Bioinformatics) 2017, 10553 LNCS, 240-248. https: / / doi.org / 10.1007 / 978-3-319- 67558-9_28.

[0162]

[0052] Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative Adversarial Networks. 2014. arXiv: 1406.2661.

[0163]

[0053] Martin Arjovsky, Soumith Chintala, and Leon Bottou. Wasserstein GAN. 2017. arXiv: 1701.07875.

[0164]

[0054] Lvwei Wang, Rong Bai, Xiaoxuan Shi, Wei Zhang, Yinuo Cui, Xiaoman Wang, Cheng Wang, Haoyu Chang, Yingsheng Zhang, Jielong Zhou, Wei Peng, Wenbiao Zhou, and Bo Huang. “A pocket-based 3D molecule generative model fueled by experimental electron density”. In: Scientific Reports 12.1 (2022), p. 15100. issn: 2045-2322. doi: 10.1038 / S41598-022-19363-6. url: https: / / doi.org / 10.1038 / s41598-022-19363-6.

[0165]

[0055] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. “GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium”. In: (2017). doi: 10.48550 / ARXIV.1706.08500. url: https: / / arxiv.org / abs / 1706.08500.

[0166]

[0056] Jiaming Gong, Wei Liu, Mengjie Pei, Chengchao Wu, and Liufei Guo. “ResNetIO: A lightweight residual network for remote sensing image classification”. In: 2022 14th Inter- national Conference on Measuring Technology and Mechatronics Automation (ICMTMA). 2022, pp. 975-978. doi: 10.1109 / ICMTMA54903.2022.00197.

Claims

Claims1 . Method for screening substances having a required physical property, chemical property or physiological effect from a group of substances comprising the steps• providing a group of k substances by a user, wherein k e N, and wherein the molecular weight of each substance is < 400 Da;• providing classification classes or continuous property label for a physical property, chemical property or physiological effect comprising C, classes or V, values, wherein i e N;• calculating electron density cube files for every substance by semiempirical molecular calculation and providing weights Wk, wherein Wk are selected from a group comprising electronegativity cube files, electron affinity cube files, molecular softness cube files and molecular hardness cube files for every substance k;• providing a trained artificial neural network architecture for assigning the k substances to the classification classes or continuous property label, wherein the artificial neural network architecture uses the electron density cube files and weights Wk as input to calculate a tensor of weights, comprising electronegativity, electron affinity, molecular softness or molecular hardness;• assigning each substance to a class C, or a value V, by utilizing the tensor of weights of substance k as weighting for the electronic density cube files in the trained artificial neural network architecture;• displaying and / or outputting of the k substances assigned to the classes C, of the classification or assigned to the values V, of the property label;• selecting the substances assigned to the class C, with the required physical property, chemical property or physiological effect or assigned to the required value Vi of a physical property, chemical property or physiological effect;• experimental verification by a user of the physical property, chemical property or physiological effect of at least part of the selected substances.

2. Method according to claim 1 , characterized in that further a tensor of shape of at least one substance is calculated by the trained artificial neural network architecture, describing the spatial structure of the at least one substance.

3. Method according to one of the preceding claims, characterized in that the trained artificial neural network architecture for the classification or property label is trained for a defined classification or property label by a training dataset, wherein the training dataset consistsof substances with known assignment to the classes C, of the classification or with known assignments to the value V, of the property label.

4. Method according to one of the preceding claims, characterized in that artificial neural network architecture is a segmentation network or a Generative Adversarial Network (GAN).

5. Method according to claim 4, characterized in that the segmentation network is modified 3D-UNet, VNet or Inception architecture, or similar networks.

6. Method according to one of the preceding claims, characterized in that the tensor of electronegativity comprises the spatial distribution of the electronegativity encoded in discrete electronegativity classes.

7. Method according to claim 6, characterized in that the electronegativity is encoded in two electronegativity classes or in three electronegativity classes.

8. Method according to one of the preceding claims, characterized in that the electronic density cube file and their corresponding local electronegativity map is illustrated as a substance structure with an overlaid electronegativity map.

9. Method according to one of the preceding claims, characterized in that the classification or property label is selected from the group of structure-based properties of molecules, in particular from the group containing odour, taste, colour, water solubility, volatility, toxicity, permitted and non-permitted chemicals in cosmetics and personal care, boiling point, melting point, solubility, aroma thresholds, partition coefficient.

10. Use of the method according to claim 1 to 9 to select at least one substance with a requested chemical property, physical property or physiological effect from a group of substances.