Method and system for predicting at least one physicochemical and / or odor characteristic value of chemical structure or composition

By integrating neural networks or multi-branch neural network models end-to-end, the problems of excessive model hyperparameters and insufficient data in the fields of chemistry and biology are solved, the accuracy and stability of predictions are improved, the demand for computing resources is reduced, and higher training speed and performance are achieved.

CN120641987APending Publication Date: 2025-09-12FIRMENICH SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380092921.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-08
Filing Date
2023-12-08
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In the fields of chemistry and biology, existing technologies have excessive numbers of hyperparameters in machine learning models and insufficient dataset size, leading to model overfitting and unstable predictions. There is also a lack of effective methods to optimize model performance and reduce data bias.

Method used

An end-to-end integrated neural network or multi-branch neural network model is used to define a digital representation of the chemical structure or composition, use multiple neural network sub-devices to make independent predictions, and use a sampling device to output random values ​​according to the probability distribution for backpropagation training, thereby reducing model parameters and improving prediction stability.

Benefits of technology

It improves the accuracy and stability of the prediction of physicochemical and odor properties of chemical structures or compositions, reduces computing resource requirements, provides a variance measure of model uncertainty, and achieves higher training speed and overall performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120641987A_ABST
    Figure CN120641987A_ABST
Patent Text Reader

Abstract

A method (100) for predicting physicochemical and / or odour characteristic values comprises the following steps:-a defining step (105) defining a representation of a chemical structure or composition,-an executing step (110) executing an integrated neural network or multi-branch neural network model of end-to-end training according to the defined representation, thereby predicting physicochemical and / or odour characteristic values,-a step (105) for predicting the physicochemical and / or odour characteristic values,-a step (105) for predicting the physicochemical and / or odour characteristic values, and-a step (110) for predicting the physicochemical and / or odour characteristic values. -a step (115) of providing physicochemical and / or odor characteristic values, the method further comprising:-a step (120) of providing example data to an end-to-end integrated neural network or multi-branch neural network device, comprising:-several neural network sub-devices configured for independent prediction,-a layer outputting at least one independently predicted distribution value,-a step (115) of providing physical and / or odor characteristic values,-a step (120) of providing example data to the end-to-end integrated neural network or multi-branch neural network device, the step (120) of providing example data to the end-to-end integrated neural network or multi-branch neural network device, and-said layer comprises sampling means configured to output random values,-an operating step (125) operating the end-to-end integrated neural network or multi-branch neural network means, and-an obtaining step (130) obtaining a trained integrated neural network or multi-branch neural network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention aims to provide a method for predicting at least one physicochemical and / or odor property value of a chemical structure or composition, a system for predicting at least one physicochemical and / or odor property value of a chemical structure or composition, and a method for effectively assembling a chemical structure or composition.

[0002] Especially suitable for the flavor and fragrance industry. Background Art

[0003] In a scientific experiment, the measured values ​​305 stored in the database (such as Figure 3 ) can vary depending on the experimental environment. Ideal environments, such as the International Space Station, are often required to achieve stable conditions. Even under near-perfect environmental conditions, technicians may use slightly different instruments and sample preparation. Because experimental variables vary depending on the measurement conditions used, combining experimental data from different sources can be difficult. Statistical methods have been developed to homogenize experimental data to reduce these variations, with moderate to good results. In machine learning, this variation can exist between training and test sets, as well as between known and future data.

[0004] Another well-known issue with machine learning models is the number of hyperparameters in the model, which can significantly affect the model’s ability to overfit the training data310, as Figure 3 As shown in Figure 2, the amount of available data may not be sufficient to train the required number of hyperparameters in a neural network. Therefore, it is best to consider the minimum required to extract a meaningful numerical representation of the data. This can be achieved by choosing layers with a smaller number of parameters. An example of such an effort is to replace dense layers with compact convolutional layers. Convolutional layers are specialized for extracting local features from data, such as images. However, even with compact and efficient convolutional layers, networks with billions of parameters can still be generated. Popular giant models indicate that such models can only be trained using very large datasets, such as those available for images, literature, and music. In contrast, to avoid expensive experiments and animal testing, datasets in scientific fields such as chemistry and biology are often limited to hundreds to thousands of data points. Therefore, in such fields, efforts should be made to reduce the number of parameters to match the size of the data.

[0005] One way to compensate for an oversized network is data augmentation. In fact, as the augmentation rate increases, performance also improves, which means that the selected network size requires more data, which means that the network size can be reduced (Tetko, IV, Karpov, P., Van Deursen, R. et al. State-of-the-art augmented NLP transformer models for direct and single-step retrosynthesis. Nat Commun 11, 5575 (2020)). At the same time, data augmentation can be used to identify whether the network has reached critical parameterization, that is, the point where augmentation has no or little effect on model performance. Not all models support data augmentation. For example, graph neural networks (GNNs) are invariant to representation shuffling. Therefore, GNNs are incompatible with existing data augmentation methods for natural language processing or images.

[0006] A third issue is that the model training process 315 defines the important dynamics of modeling. Often, there are numerous variables that could influence the model's decisions. This may explain why hyperparameterized optimization strategies are needed to improve model performance or efficiency. Beyond the chosen model, the data split between the training and test sets also plays a significant role. A variety of methods, ranging from leave-one-out-all to random splits to K-fold cross-validation, can be used to simulate and evaluate the quality of a model on unseen data. Ultimately, the model's predictions are only an educated guess, depending on the training conditions, model size, optimization parameters, and data split used. Upon completion, it is impossible to know whether the optimal model was truly trained. However, one assumes that the calculated model is the optimal model for the test points used. This is a general limitation of data modeling methods, as one should not expect the same performance results on all future unseen data. It is important to note that future performance can also vary significantly, depending on the sample size used for evaluation on unseen data and any sample bias that may be introduced in the unseen data. One approach to partially address these shortcomings is to predict an accurate theoretical endpoint as a standard metric for model evaluation. The molecular weight of a molecule in chemistry is one such endpoint.

[0007] There are three main branches of using neural networks in the field of digital modeling of chemical substances and chemical reactions.

[0008] The first branch involves learning models from graph neural networks (GNNs). To compute atomic properties, any molecular input format can be used. This format is not easily scalable and struggles with smaller datasets common across all chemistry fields.

[0009] The second branch is based on NLP methods that use string literals (e.g., SMILES format), where chemical knowledge is learned entirely using this syntax. This approach has the advantage of data augmentation because the same molecule can be written in a new sentence using a different order of rules (sentence grammar).

[0010] The third branch is an image convolutional neural network that learns and predicts from molecular images.

[0011] Such methods require large datasets, which are not common in fields such as fragrance design and olfaction measurement, perfumery, advanced fragrance making, and flavor design. Without rich datasets, using neural network technology may lead to inefficient models due to the large number of parameters that need to be considered and the risk of memory loss in the network.

[0012] Furthermore, the input of such graph neural networks is to convert molecules into a specified input format, usually the SMILES format of molecular structures, which is insufficient to create effective chemical signatures for chemical reaction prediction models.

[0013] Ensembling is a technique that involves training multiple models (often called base models or weak learners) and aggregating their outputs at inference time using some kind of voting mechanism.

[0014] This technique is widely used by practitioners (notably leading to successful solutions in many machine learning competitions), and it is often a key step in improving the final performance.

[0015] Despite the popularity of ensemble training, finding the best ensemble process for building, training, and combining base models is often non-trivial. Traditionally, ensemble training techniques have attempted to generate a diverse (or complementary) collection of base models and combine them using some kind of voting technique, typically to reduce the bias and / or variance of the final system. Many different techniques can be used to train diverse models. For example, bagging (using bootstrapped resampling) introduces diversity by sampling the training dataset, while boosting introduces diversity by training models sequentially so that each model has an incentive to compensate for mistakes made by the previous models. Voting techniques can include simple averaging, majority voting (for classification), or stacking, where the final prediction is generated by a meta-model that is trained to combine the base models on some held-out dataset.

[0016] While these techniques work well in practice, they are mostly hand-crafted heuristics designed to increase diversity or complementarity between base models. While these heuristics can be used to make base models complementary—for example, during training (e.g., in boosting) or during inference (e.g., in stacking), these models do not directly learn how to best complement each other. In particular, these models cannot explicitly capture the fact that they are part of an ensemble.

[0017] To address this challenge, some have proposed jointly training all base models (or end-to-end). In the context of neural networks, this means treating each base model as part of a larger neural network and training them jointly using a common loss function.

[0018] Interestingly, this blurs the concepts of ensembles and multi-head (or multi-branch) networks, as each base model can now be viewed as an independent branch within a single neural network model. This end-to-end approach is attractive, but it is well known that blindly optimizing the global loss for the entire system usually does not lead to optimal results, and it is usually better to train the base models individually for a certain amount (usually controlled by a specific term in the loss).

[0019] Currently, the best-known approaches to training such end-to-end models generally involve trying different interpolations between individual and global (or diluted) loss terms, and in general, the best approach seems to depend on the problem and the model. Summary of the Invention

[0020] The present invention aims to address all or part of these disadvantages.

[0021] According to a first aspect, the present invention is directed to a method for predicting the value of at least one physicochemical and / or odor property of a chemical structure or composition, comprising the following steps:

[0022] - a definition step, which defines a digital representation of a chemical structure or composition on a computer interface,

[0023] - executing a step of executing, by means of a computing device, an end-to-end trained ensemble neural network or multi-branch neural network model according to the defined digital representation, in order to predict at least one physicochemical and / or odor property value of the chemical structure or composition,

[0024] - providing a step of providing on a computer interface at least one physicochemical and / or odor property value of a chemical structure or composition,

[0025] The following steps are also included:

[0026] - providing a step of providing a set of example data to an end-to-end integrated neural network or a multi-branch neural network, comprising: at least one set of inputs corresponding to a digital representation of a chemical structure or composition; and at least one set of outputs corresponding to physicochemical and / or odor properties associated with the set of inputs, the end-to-end integrated neural network or the multi-branch neural network comprising:

[0027] - a plurality of neural network sub-devices, each sub-device being configured to provide independent predictions based on example data,

[0028] - a layer configured to output at least one value based on or representing a distribution of the independent predictions, and

[0029] - the layer comprises a sampling device configured to output at least one random value according to a probability distribution representing a distribution of independent predictions, the outputted random value being calculated in a differentiable manner and used for backpropagation within an end-to-end integrated neural network or a multi-branch neural network device,

[0030] - an operating step of operating an end-to-end integrated neural network or a multi-branch neural network device based on the set of example data, and

[0031] - an obtaining step of obtaining a trained end-to-end integrated neural network or multi-branch neural network model configured to predict the physicochemical and / or odor properties of the input digital representation of the chemical structure or composition.

[0032] Such a specification allows for accurate prediction of the physicochemical and / or odor property values ​​of a defined chemical structure or composition.

[0033] Such a provision can also improve the stability and reliability of predictions, increase training speed and overall performance, and provide a variance measure that represents model uncertainty. Consequently, such an implementation can save resources, including computational time or power consumption, as well as model complexity. Current approaches typically require numerous models and iterations to obtain a reliable prediction model.

[0034] Furthermore, such a regulation enables the trained models to achieve higher accuracy than other competing methods, which either require designing diversity among base models or rely on fine-tuned loss functions to balance the objectives of training individual models and ensemble models.

[0035] This rule also provides a simple way to regularize the ensemble by introducing noise. Finally, this rule leads to more stable training dynamics and better individual base models. This approach does not require any additional tuning and does not introduce new learnable parameters.

[0036] In a particular embodiment, at least one set of input example data corresponds to a hash vector of at least one atomic property in a chemical structure or composition, and the method further comprises a conversion step upstream of the execution step, which converts the defined digitized chemical structure or composition into a set of hash vectors representing at least one atomic property of the digitized chemical structure or composition, wherein the set of hash vectors is used as input during the execution step.

[0037] Such a provision has proven to be particularly effective in increasing the reliability of prediction results in the prediction of physicochemical and / or odor properties.

[0038] In particular embodiments, the hash vector of at least one atomic property represents one of the following:

[0039] - the atomic number of the corresponding atom,

[0040] - the atomic symbol of the corresponding atom,

[0041] - atomic mass,

[0042] - explicit mapping number,

[0043] - the row index in the periodic table,

[0044] - column index in the periodic table,

[0045] - the total number of hydrogen atoms,

[0046] - the implied number of hydrogens on the atom,

[0047] - the exact number of hydrogen atoms,

[0048] - the degree of the atom,

[0049] - the total degree of the atom,

[0050] - the valence state of the atoms,

[0051] - the implicit valence of the atom,

[0052] - the definite valence of the atoms,

[0053] - formal charges on atoms,

[0054] - partial charges on atoms,

[0055] - the electronegativity of the atom,

[0056] - the number of keys calculated by key type,

[0057] - adjacent numbers calculated by atomic number, wildcard,

[0058] - the number of neighbors calculated by bond type plus atomic number, wildcard,

[0059] - the number of adjacent characters counted by wildcards,

[0060] - a value indicating aromaticity,

[0061] - represents the value of aliphatic atoms,

[0062] - represents the value of conjugated atoms,

[0063] - represents the value of the ring atom,

[0064] - represents the value of the macrocyclic atom,

[0065] - represents the value of the geometrically constrained atom,

[0066] - represents the value of electron-withdrawing atoms,

[0067] - represents the value of the electron-donating atom,

[0068] - represents the value of the reaction site,

[0069] - represents the value of hydrogen bond donor,

[0070] - represents the value of hydrogen acceptor,

[0071] - a value indicating multivalency as a hydrogen bond donor,

[0072] - a value indicating multivalency as a hydrogen bond acceptor,

[0073] - the number of cycles on the atom,

[0074] - ring size on the atom,

[0075] - the hybridization state of the atoms,

[0076] - values ​​representing atomic geometry,

[0077] - the number of electrons in the atomic orbital,

[0078] - number of lone pairs of electrons,

[0079] - the state of the atomic group,

[0080] - isotopes on atoms,

[0081] -atomic centrosymmetry function,

[0082] - as the relative stereochemical value of chi clockwise and chi counterclockwise,

[0083] - the value of the absolute stereochemistry,

[0084] - the value of the absolute stereochemistry,

[0085] - the value of the double bond stereochemistry,

[0086] - a value that determines stereochemical preference,

[0087] - a value representing the positive or negative impact that an atom has on the identified training target, thereby representing a rich knowledge-based contribution,

[0088] - a value representing the positive or negative impact of the atom on the identified training target, thereby representing the knowledge-based diluted contribution and / or

[0089] - Value of the ring stereochemistry.

[0090] In particular embodiments, at least one hash vector of a key characteristic represents one of the following:

[0091] - key sequence,

[0092] -Key type,

[0093] -bond stereochemistry:

[0094] - the bond orientation of the tetrahedral stereochemistry,

[0095] -bond direction of the double bond stereochemistry or

[0096] - spatially oriented bond directions,

[0097] - the atomic number of the "from" and / or "to" atom,

[0098] - the atomic symbols of the "from" and / or "to" atoms,

[0099] - dipole moment in the bond,

[0100] -Quantum chemical properties:

[0101] - electron density in the bond,

[0102] - the electronic configuration of the bond,

[0103] -Key Track,

[0104] -bond energy,

[0105] -attraction,

[0106] - repulsive force,

[0107] -key distance,

[0108] - aromatic bonds,

[0109] - aliphatic bonds,

[0110] -Ring characteristics of the bond:

[0111] - the number of rings on the bond,

[0112] - Ring size of the key,

[0113] - minimum ring size of the key,

[0114] - maximum ring size of the key,

[0115] - rotatable key,

[0116] - spatial constraint keys,

[0117] - Hydrogen bonding properties,

[0118] - ionic bond characteristics,

[0119] -The bond order of the reaction, including "empty" bonds, used to identify the bonds broken / formed in the reaction:

[0120] - the bond order in the reagent,

[0121] - the bond order in the intermediate, or

[0122] Bond order in the transition state.

[0123] In particular embodiments, the at least one output value representing the distribution represents a dispersion of the distribution.

[0124] In certain embodiments, an end-to-end ensemble neural network or a multi-branch neural network device is trained to minimize at least one value representing the dispersion of the distribution.

[0125] In certain embodiments, the at least one odor characteristic is representative of:

[0126] -Anti-insect ability value,

[0127] - sensory property values,

[0128] - biodegradability value,

[0129] - antibacterial value,

[0130] - Odor detection threshold,

[0131] - Odor intensity value,

[0132] - Top-Center-Bottom values,

[0133] - Danger value,

[0134] - Biological activity of taste

[0135] - biological activity of the sense of smell,

[0136] - Biologically augment or modulate taste activity,

[0137] - Biologically augment or modulate olfactory activity, and / or

[0138] - Olfactory smell description.

[0139] In certain embodiments, the at least one physical characteristic represents:

[0140] - boiling point value,

[0141] - melting point value,

[0142] - water solubility value,

[0143] - Henry's constant value,

[0144] - vapor pressure value,

[0145] - Volatility value or

[0146] -Headspace concentration value.

[0147] In certain embodiments, at least one neural network device is:

[0148] - Recurrent Neural Network Device,

[0149] -Graph Neural Network Device,

[0150] - Variational Autoencoder Neural Network Device or

[0151] -Autoencoder neural network device.

[0152] In a particular embodiment, the method object of the present invention comprises an atom or bond relationship vector expansion step upstream of the providing step.

[0153] This provision allows for the use of more limited datasets than currently required for neural network applications. In fact, a molecular structure represented by one or more expanded hash value sequences can be expanded up to a number of times corresponding to the number of hash value sequences. Thus, a single molecular structure can serve as multiple inputs for natural language processing applications.

[0154] In a particular embodiment, the atom or bond relationship vector expansion step includes a horizontal expansion step configured to provide several vectors representing a single digital representation of a molecular structure or composition, each vector representing a specific representation of a canonical representation of the molecular structure or composition, each vector being considered as a single input during the providing step.

[0155] This provision allows for the use of more limited datasets than currently required for neural network applications. In fact, a molecular structure represented by one or more expanded hash value sequences can be expanded up to a number of times corresponding to the number of hash value sequences. Thus, a single molecular structure can serve as multiple inputs for natural language processing applications.

[0156] In certain embodiments, the atom or bond relationship vector expansion step includes a vertical expansion step that creates several groups of several horizontal expansions representing unique molecular structures or compositions, each group being treated as a single input during the providing step.

[0157] This provision allows for the use of more limited datasets than currently required for neural network applications. In fact, a molecular structure represented by one or more expanded hash value sequences can be expanded up to a number of times corresponding to the number of hash value sequences. Thus, a single molecular structure can serve as multiple inputs for natural language processing applications.

[0158] According to a second aspect, the present invention is directed to a method for efficiently assembling a chemical structure or composition, comprising:

[0159] - executing steps which carry out the method object of the invention, and

[0160] - an assembling step, which assembles a chemical structure or composition related to the output obtained in the obtaining step.

[0161] Such a specification allows the physicalization of chemical structures for the prediction of odor properties.

[0162] According to a third aspect, the present invention is directed to a system for predicting the value of at least one physicochemical and / or odor property of a chemical structure or composition, comprising:

[0163] - a definition mechanism that defines a digital representation of a chemical structure or composition on a computer interface,

[0164] - an execution unit which, by means of a computing device, executes an end-to-end trained ensemble neural network or multi-branch neural network model according to the defined digital representation, thereby predicting at least one physicochemical and / or odor property value of a chemical structure or composition,

[0165] - providing means for providing on a computer interface the value of at least one physicochemical and / or odorous property of a chemical structure or composition,

[0166] The system includes the following mechanisms:

[0167] - providing a mechanism for providing a set of example data to an end-to-end integrated neural network or a multi-branch neural network, comprising: at least one set of inputs corresponding to digital representations of chemical structures or compositions; and at least one set of outputs corresponding to physicochemical and / or odor properties associated with the set of inputs, the end-to-end integrated neural network or the multi-branch neural network comprising:

[0168] - a plurality of neural network sub-devices, each sub-device being configured to provide independent predictions based on example data,

[0169] - a layer configured to output at least one value based on or representing a distribution of the independent predictions, and

[0170] - the layer comprises a sampling device configured to output at least one random value according to a probability distribution representing a distribution of independent predictions, the outputted random value being calculated in a differentiable manner and used for backpropagation within an end-to-end integrated neural network or a multi-branch neural network device,

[0171] - an operating mechanism that operates an end-to-end integrated neural network or a multi-branch neural network device based on the set of example data, and

[0172] - obtaining means for obtaining a trained end-to-end integrated neural network or multi-branch neural network model configured to predict the physicochemical and / or odor properties of an input digital representation of a chemical structure or composition.

[0173] The advantages of the system object of the present invention are similar to those of the method object of the present invention. In addition, all embodiments of the method object of the present invention can be reproduced in the system object of the present invention, but must be modified. BRIEF DESCRIPTION OF THE DRAWINGS

[0174] Other advantages, objects, and specific features of the present invention will be apparent from the following non-exhaustive description of at least one specific embodiment or series of steps of the present invention, taken in conjunction with the accompanying drawings, in which:

[0175] - Figure 1 A first specific sequence of steps of the method object of the present invention is schematically shown;

[0176] - Figure 2 A specific embodiment of the system object of the present invention is schematically shown;

[0177] - Figure 3 The overall overview of the machine learning system is schematically shown;

[0178] - Figure 4 A second specific sequence of steps of the method object of the present invention is schematically shown;

[0179] - Figure 5 schematically illustrates a detailed view of a novel neural network layer used during training of an end-to-end integrated neural network or a multi-branch neural network device according to a particular embodiment;

[0180] - Figure 6 A second specific sequence of steps of the method object of the present invention is schematically shown;

[0181] - Figure 7Schematically illustrates a specific sequence of steps for obtaining a hash vector used by the system or method object of the present invention;

[0182] - Figure 8 An example of an embodiment of the training method object of the present invention is schematically shown;

[0183] - Figures 9 to 11 Performance results for three odor characteristics are shown;

[0184] - Figures 12 to 14 Schematic illustration of a training architecture that can be used to select, classify, or predict odor properties and / or physicochemical properties of chemical structures; and

[0185] - Figure 15 A specific training architecture that can be used to classify chemical structures is schematically shown. DETAILED DESCRIPTION

[0186] This description is not exhaustive, as every feature of one embodiment may be combined in an advantageous manner with any other feature of any other embodiment.

[0187] Various inventive concepts can be embodied as one or more methods, examples of which are provided herein. The actions performed as part of a method can be ordered in any suitable manner. Thus, embodiments can be constructed that perform actions in an order different from that shown, and even if illustrated as sequential actions in an illustrative embodiment, some actions may be performed simultaneously.

[0188] The indefinite articles "a" and "an" as used in this specification and claims should be understood to mean "at least one" unless expressly stated otherwise.

[0189] The phrase "and / or" as used in this specification and claims should be understood to mean "either or both" of the elements so connected together, that is, these elements exist in a connected manner in some cases and in a separated manner in other cases. Multiple elements listed with "and / or" should be understood in the same way, that is, "one or more" of the elements so connected together. In addition to the elements explicitly specified by the "and / or" clause, other elements may optionally be present, whether or not these elements are related to the explicitly specified elements. Thus, as a non-limiting example, when referring to "A and / or B", when used in conjunction with open language (such as "comprising"), in one embodiment, it may refer to only A (optionally including elements other than B); in another embodiment, it may refer to only B (optionally including elements other than A); in yet another embodiment, it may refer to A and B (optionally including other elements); and so on.

[0190] As used in this specification and claims, "or" should be understood to have the same meaning as "and / or" defined above. For example, when separating items in a list, "or" or "and / or" should be understood to be inclusive, i.e., including at least one, but also including more than one of a plurality of elements or a list of elements, and optionally, including other unlisted items. Only terms that clearly indicate the opposite meaning, such as "only one" or "exactly one", or "consisting of..." used in the claims, refer to the inclusion of multiple elements or exactly one element in a list of elements. In general, the term "or" as used herein should be understood to represent an exclusive choice (i.e., "one or the other, but not both") only when it is preceded by an exclusive term (e.g., "any one", "one of which", "only one" or "exactly one"). "Mainly consisting of...", when used in the claims, should have its usual meaning in the field of patent law.

[0191] As used in this specification and claims, the phrase "at least one" referring to a list of one or more elements should be understood to mean at least one element selected from any one or more elements in the list of elements, but does not necessarily include at least one of each element explicitly listed in the list of elements, nor does it exclude any combination of elements in the list of elements. This definition also allows for the optional presence of other elements in addition to the elements explicitly named in the list of elements to which the phrase "at least one" refers, whether related or unrelated to the elements explicitly named. Thus, as a non-limiting example, "at least one of A and B" (or equivalently, "at least one of A or B," or equivalently, "at least one of A and / or B") may, in one embodiment, refer to at least one, optionally including more than one A, and no B (and optionally including elements other than B); in another embodiment, refer to at least one, optionally including more than one B, and no A (and optionally including elements other than A); in yet another embodiment, refer to at least one (optionally including more than one) A, and at least one (optionally including more than one) B (and optionally including other elements); and so on.

[0192] In the claims and the foregoing description, all conjunctions such as "include," "comprising," "carrying," "having," "containing," "involving," "having," "consisting of," etc., shall be understood as open-ended conjunctions, i.e., including but not limited to. Only the conjunctions "consisting of" and "consisting essentially of" shall be closed or semi-closed conjunctions, respectively.

[0193] It is important to note at this time that the figures are not drawn to scale.

[0194] As used herein, the term "ingredient" refers to any ingredient, preferably an ingredient with perfume or fragrance capabilities. "Compound" or "ingredient" refers to the same substance as "volatile ingredient". An ingredient can be composed of one or more chemical molecules.

[0195] The term "composition" refers to a liquid, solid or gaseous combination of at least two aroma or flavor components or one aroma or flavor component and a neutral solvent for dilution.

[0196] As used herein, "flavor" refers to the olfactory perception produced by at least one volatile component through the activation, expansion, and inhibition (if present) of odorant receptors via orthonasal and retronasal olfaction, as well as activation of taste buds containing taste receptor cells. Thus, for the purpose of illustrating and not limiting the scope of this disclosure, "flavor" refers to the olfactory and taste bud perceptions produced by a first volatile component activating an odorant receptor or taste bud associated with the flavor of coconut, a second volatile component activating an odorant receptor or taste bud associated with the flavor of celery, and a third volatile component inhibiting an odorant receptor or taste bud associated with the flavor of hay.

[0197] As used herein, "fragrance" refers to the olfactory perception resulting from the aggregation of at least one volatile component that activates, expands, and, when present, inhibits odorant receptors. Thus, for the purpose of illustrating, but not limiting, the scope of this disclosure, a "fragrance" is the olfactory perception resulting from the aggregation of the following components: a first volatile component that activates odorant receptors associated with the smell of coconut, a second volatile component that activates odorant receptors associated with the smell of celery, and a third volatile component that inhibits odorant receptors associated with the smell of hay.

[0198] As used herein, "odor properties" or "olfactory properties" refer to any psychophysical properties of an ingredient or composition. In other words, such properties refer to the body's response to the physical presence of an olfactory ingredient or composition, and such psychophysical properties are directly related to the ability of the ingredient or composition to readily penetrate and come into close proximity with the body's olfactory receptors.

[0199] As used herein, the term "input device" refers to, for example, a keyboard, mouse, and / or touch screen for interacting with a computing system to collect user input. In some variations, the input device is logical in nature, for example, a network port of a computing system configured to receive electronically transmitted input commands. Such input devices may be associated with a GUI (graphical user interface) or an API (application programming interface) displayed to the user. In other variations, the input device may be a sensor configured to measure a specific physical parameter relevant to the intended use case.

[0200] As used herein, the term "computing system" or "computer system" refers to any electronic computing device, whether modular or distributed, that is capable of receiving numerical input and providing numerical output through any type of digital and / or analog interface. Typically, a computing system refers to a computer that runs software and has access to data storage, or to a client-server architecture where data and / or computations are performed on the server side, with the client acting as the interface.

[0201] As used herein, the term "digital identifier" refers to any computerized identifier, such as an identifier used in a computer database to represent a physical object (e.g., a flavoring ingredient). A digital identifier can refer to a label representing the flavoring ingredient's name, chemical structure, or internal reference.

[0202] Throughout this specification, the term "materialized" means existing outside of the digital environment of the present invention. "Materialized" can mean, for example, readily found in nature or synthesized in a laboratory or chemical plant. Regardless, a materialized composition takes on a tangible reality. Terms such as "to be compounded" or "compounded" refer to the act of materializing a composition, whether by extracting and assembling the components or by synthesizing and assembling the components.

[0203] As used herein, the term "atomic properties" refers to properties of an atom and / or the bonds to which any atom is attached, regardless of the molecular context in which it is used. Thus, atomic properties refer to absolute descriptions of atomic characteristics, rather than relative descriptions of atoms within a molecule within the broader molecular context.

[0204] The term "activation function" as used in this article defines how a weighted sum of inputs is transformed into the output of one or more nodes in a layer of a neural network. These activation functions can be defined by the arithmetic solution of a layer in the network or by a loss function.

[0205] As used herein, "end-to-end integrated neural network or multi-branch neural network device" refers to a set of independent neural network devices that cooperate with each other to provide an output, as well as a single neural network device composed of independent branches that cooperate with each other to provide an output.

[0206] As used herein, the term "atomic properties" refers to properties of an atom and / or the bonds to which any atom is attached, regardless of the molecular context in which it is used. Thus, atomic properties refer to absolute descriptions of atomic characteristics, rather than relative descriptions of atoms within a molecule within the broader molecular context.

[0207] The embodiments disclosed below are presented in a general manner.

[0208] Figure 3 Shows an overview of the key components of machine learning.

[0209] Figure 5The two-layer specific implementation of the end-to-end integrated neural network or multi-branch neural network training device object of the present invention is schematically shown. Figure 5 It also helps to understand the technical contribution of the present invention. Figure 5 The underlying theory of the model shown is as follows.

[0210] Let ε be the K-based model M composed of an integrated (or multi-branch) neural network k , k = 1...K. This method can be considered a new neural network layer that combines the vector outputs of multiple base models into one. During training, the process is as follows:

[0211] Let x be an input to the neural network.

[0212] The k-th base model outputs o k =M k (x)∈R h , where h is the output dimension of the base model. This layer inputs k ,k=1...K,output o~D(g(o 1 ,...,o K )), where ~ represents differentiable sampling, D is a multivariate distribution, g is a function that maps the vector ok to the distribution parameters, and o∈R h The output o of this layer has the same dimensions as the input vectors. The final output of can be expressed as Where f is a function that provides the correct output format for the current task (e.g., softmax for classification). During inference, the layer outputs, for example, D(g(o 1 ,...,o K )), not the mean of a random sample.

[0213] Using the reparameterization trick, sampling can be performed in a differentiable manner, making it compatible with gradient descent-based neural network training. Therefore, unlike traditional ensemble methods (such as bagging or stacking), which separate the training of each model in the ensemble, this layer ensures that all base models for all training samples obtain the gradient, thus achieving an end-to-end training form.

[0214] There are different options for the computation performed by this layer, specified by D and g. For example, Figure 5 A simple variant is shown, where D is a Gaussian distribution parameterized by a diagonal covariance matrix, and the function g(o 1 ,...,o K ) calculate the parameters μ∈Rh and σ∈Rh so that oi~N(μ i ,σ 2 i ).

[0215] The following discloses the 1 ,...,o K ))Several different methods of construction and sampling.

[0216] - Mean: No sampling is performed, and the output is calculated by taking the mean of each base model. This can be summarized as training a multi-head neural network where the output of all heads is taken as the mean. More precisely,

[0217] -Diagonal: This corresponds to Figure 5 During training, the output o from the base model k Calculate the following values: and σ so that Then generate samples∈~N(0,1), and the output of this layer is o i =μ i +∈ i σ i , so o obeys N(μ,diag(σ 2 i )) distribution. It is inferred that the output of this layer is o = μ.

[0218] -Diagonal parameterization: This method is similar to diagonal, but it calculates and where l θ and l γ are two learnable functions (e.g., set as fully connected layers) that directly output μ and σ. This specified distribution D(g(o 1 ,...,o K )) in a manner reminiscent of VAEs, except that this layer computes μ and σ based on multiple underlying base models.

[0219] -Full covariance: This is an extension of diagonal covariance, where sampling is not done using a diagonal covariance matrix, but using a full covariance matrix. Also, during training, this layer computes In principle, the covariance matrix must be calculated as However, to apply the reparameterization trick in this setting, the layer needs to compute samples o = μ + R∈, where ∈ ≈ N(0, I) and R is typically decomposed via Cholesky decomposition Σ = RR TThis is problematic because the Cholesky decomposition requires Σ to be positive definite. One solution is to compute the decomposition on Σ′ = Σ + τI for some small τ∈R+. However, in practice, it has been observed that this solution leads to numerical problems that affect the results. Instead, a simpler method can be used that bypasses the computation of Σ and the Cholesky decomposition altogether. Note that where R' has (o1,...,o K )) as a column, you can calculate where ∈~N(0,I). Similarly, during inference, this layer simply returns o=μ.

[0220] The performance of this architecture can be evaluated by comparing our approach on the CIFAR-10 image classification task. Each competing model is trained using 5 random seeds and 120 epochs. The test loss is calculated on the entire test set, which follows the conventional CIFAR-10 split of 50,000 images for training and 10,000 for testing.

[0221] The training method object of the present invention can be compared with different ensemble methods to evaluate the effect of sampling as a new technique for end-to-end ensemble training. All ensemble methods use K=8 base models, which are standard CNNs with ReLU and batch normalization layers. Each base model contains 68,906 parameters, so each ensemble contains a total of 551,248 parameters. This paper evaluates the different variants mentioned above, which refer to parameterized isotropic variants of multilayer perceptrons for the function l(·). It is observed that when diagonal sampling is used, the training may be unstable at the beginning if the weights are initialized based on a uniform distribution. This is because the initial base models are not diverse enough at the beginning of training, resulting in a standard deviation close to zero, which makes Gaussian sampling prone to numerical instability. Therefore, Gaussian or orthogonal initialization can be used, which do not seem to be affected by this problem. In this particular embodiment, a bagging version based on random initialization of the network weights and random shuffling of the data points is used. Finally, negative correlation learning ("NCL") is used, as shown in the following equation:

[0222]

[0223] Where L is the loss function, K, y, and where represents the number of base models in the ensemble, the i-th target, and the i-th prediction of the ensemble, respectively. j is the index of the j-th subunit and the prediction of the k-th base model. The second term measures the diversity among ensemble members, and λ is a hyperparameter to be tuned. Intuitively, λ = 0 corresponds to individual training, while λ = 1 corresponds to end-to-end training.

[0224] In addition to the ensemble method, the results are also compared with a standalone CNN of similar capacity to the ensemble method, without dropout ("Simple") and with dropout ("Simple+Dropout"). The structure of this CNN is similar to the CNN used for the base model, but it has 506,290 parameters, which can be obtained by increasing the depth and number of channels.

[0225] The validation accuracy of different models can be used as a performance measure. The coefficient of variation measures the diversity among the members of the ensemble during training. It is calculated as the mean of the standard deviations of the elements, which is given by 1 ,...,o K Finally, the average test accuracy of the base models can be used as a performance metric. This measures the dilution effect during training, that is, the performance of each independent base model on the test set.

[0226] This comparison shows that:

[0227] - Mean outperforms Single, indicating that multiple heads can provide some benefit even when simply averaging the outputs of the base models during training. However, mean is deterministic and does not inject any noise. By comparing diagonal and mean, we can see the benefit of sampling during training over simple averaging, as diagonal clearly provides better results.

[0228] Full covariance sampling ultimately achieves higher validation accuracy than the simpler diagonal sampling method. The larger coefficient of variation for full covariance sampling compared to diagonal sampling appears to indicate that full covariance sampling benefits from greater diversity. Furthermore, comparing the two sampling methods, full covariance sampling achieves better aggregate test accuracy starting around the 30th training epoch. However, it is noteworthy that full covariance sampling only outperforms diagonal sampling in test accuracy around the 60th epoch. In other words, even with lower test accuracy, full covariance sampling achieves higher average individual test accuracy than diagonal sampling. Therefore, full covariance sampling exhibits better dilution properties. Overall, sampling from a richer distribution appears to lead to better results. However, full covariance sampling only outperforms the other methods after a few dozen epochs and is computationally more expensive. Overall, this gain comes at the expense of more computation for the same number of parameters.

[0229] - By comparing the diagonal MLP and the diagonal MLP, we can notice a net advantage of the latter, which indicates that using this MLP parameterization function from o 1 ,...,o KComputing the parameters D by concatenating the vectors of is worse than directly computing the mean and standard deviation of the elements of the vector. However, it is not clear whether the performance gap comes from the inductive bias of the direct (non-parametric) option or other factors. Therefore, some people have proposed to directly calculate the parameters of D, which also has the advantage of not adding new learnable parameters.

[0230] - It is assumed that this training setup acts primarily as a regularization mechanism, allowing end-to-end training of ensembles that would otherwise be prone to overfitting. In fact, even though averaging achieves a net improvement over singletons, it still underperforms bagging in terms of test accuracy and dilution. Nevertheless, once sampling is used, better test accuracy is achieved, and the dilution effect is also better than sampling. Furthermore, injecting noise in this way means that the sampling process adapts to the o during training. 1 ,...,o K In contrast, relying on a noise schedule adds some complexity and requires careful tuning.

[0231] The performance gap between diagonal and full covariance suggests that more complex distributions provide better expressive power for the ensemble, which implies that the sampling distribution can capture part of the data and therefore acts as more than just a noise mechanism providing variable magnitude regularization during training.

[0232] The single-step method with dropout is more regularized than the full covariance method (all others tend to overfit the training set). However, the full covariance method still performs best on the test set. Thus, our sampling-based method appears to address different regions of the bias-variance tradeoff.

[0233] NCL and bagging have similar performance in terms of test accuracy, but NCL has slightly better dilution and greater diversity, which reflects the advantages of end-to-end ensembles. However, in addition to these advantages, this method only requires choosing the right distribution, without having to adjust new hyperparameters like NCL.

[0234] - Training independent CNNs can lead to overfitting, which can be mitigated by dropout. However, single layers + dropout are more unstable than diagonal layers and full covariance layers, which can be a problem when stability is important (as is often the case in industrial deployments).

[0235] Finally, the following table shows that full covariance has an advantage over competing methods in terms of test accuracy.

[0236] Model Verification accuracy Bagging 81.6%±0.2 Complete covariance (this method) 83.1%±0.3 Diagonal MLP (this method) 76%±0.5 Diagonal (this method) 82.4%±0.2 mean 79.2%±0.3 NCL 81.7%±0.4 Single Point 77%±0.5 Single Click + Exit 82.8%±1.1

[0237] Therefore, this method is particularly suitable for combining multiple branches of a neural network, which can be viewed as a way to train an ensemble of neural networks end-to-end. It consists of a new neural network layer that takes multiple independent predictions from different base models (or branches) as input and uses differentiable sampling to produce a single output, while providing regularization and distributing gradients to all base models. This approach has multiple advantages.

[0238] First, it achieves higher accuracy than competing methods, which either require engineering diversity among base models or rely on fine-tuned loss functions to balance the objectives of training individual models and ensembles.

[0239] Second, it provides a simple way to normalize the collection by introducing noise.

[0240] Third, it can lead to more stable training dynamics and better individual base models. This approach does not require any additional tuning and does not introduce new learnable parameters.

[0241] Figure 1 The specific sequence of steps of the method 100 of the present invention is shown. The method 100 is used to predict at least one physicochemical and / or odor property value of a chemical structure or composition, comprising the following steps:

[0242] - a definition step 105 of defining a digital representation of a chemical structure or composition on a computer interface,

[0243] - executing step 110 of executing, by means of a computing device, an end-to-end trained ensemble neural network or multi-branch neural network model according to the defined digital representation, in order to predict at least one physicochemical and / or odor property value of the chemical structure or composition,

[0244] - providing a step 115 of providing on a computer interface at least one physicochemical and / or odor property value of a chemical structure or composition,

[0245] The method 100 of the present invention further comprises the following steps:

[0246] - providing step 120, which provides a set of example data to an end-to-end integrated neural network or a multi-branch neural network device, the set of example data comprising: at least one set of inputs corresponding to a digital representation of a chemical structure or composition; and at least one set of outputs corresponding to physicochemical and / or odor properties associated with the set of inputs, the end-to-end integrated neural network or the multi-branch neural network device comprising:

[0247] - a plurality of neural network sub-devices, each sub-device configured to provide independent predictions based on example data,

[0248] - a layer configured to output at least one value based on or representing a distribution of the independent predictions, and

[0249] - the layer comprises a sampling device configured to output at least one random value according to a probability distribution representing a distribution of independent predictions, the output random value being calculated in a differentiable manner and used for backpropagation within an end-to-end integrated neural network or a multi-branch neural network device,

[0250] - an operation 125 of operating an end-to-end integrated neural network or a multi-branch neural network device based on the set of example data, and

[0251] - An obtaining step 130 of obtaining a trained end-to-end ensemble neural network or multi-branch neural network model configured to predict the physicochemical and / or odor properties of an input digital representation of a chemical structure or composition.

[0252] It should be noted that a layer configured to output at least one value based on or representing the independently predicted distribution may be understood as a layer providing values ​​representing the distribution to be used by the sampling device, or a layer providing values ​​obtained from the sampling device.

[0253] "Differentiable" means that samples are drawn from the distribution in such a way that the gradients of the layer output with respect to the distribution's parameters can be computed. This also means that these parameters are computed as differentiable functions of the outputs of the neural network subunits. This results in a "proper" neural network layer whose output gradients with respect to its input can be computed, allowing it to be embedded in any larger neural network trained using backpropagation.

[0254] Defining step 105 is performed, for example, by using input device 240 coupled to I / O subsystem 220, such as with respect to Figure 2 The public.

[0255] During this definition step 105 a chemical structure or composition is defined.

[0256] Chemical structure is defined as the molecular geometry and, optionally, the electronic structure of a target molecule. Molecular geometry refers to the spatial arrangement of atoms in a molecule and the chemical bonds that hold them together. It can be represented by structural formulas and molecular models; a complete description of the electronic structure includes specifying the occupation of molecular orbitals. Structure determination can be applied to a range of targets, from very simple molecules (such as diatomic oxygen or nitrogen) to very complex molecules (such as proteins or DNA).

[0257] A composition is defined as the sum of molecules or compounds, often referred to as flavor or fragrance ingredients.

[0258] For example, in the definition step 105, a user can connect to the GUI and select an existing chemical structure, or design a chemical structure by specifying its constituent atoms and their associated geometric shapes. Alternatively, a user can connect to the GUI and select an existing fragrance or flavor component, each of which is associated with at least one chemical structure. The selection or definition of such a chemical structure or composition is accomplished through a digital representation of the material equivalent of the chemical structure or composition. The representation can be presented in textual form and associated with an entry in a computer database, with each representation storing a plurality of parameters.

[0259] Step 110 is performed, for example, by one or more hardware processors 210 (e.g. Figure 2 The hardware processor is configured to execute a set of instructions representing a trained end-to-end integrated neural network or multi-branch neural network model. Figure 5 Specific implementations for implementing step 110 are disclosed.

[0260] The inputs to step 110 depend on the parameters of the end-to-end integrated neural network or multi-branch neural network device, thereby obtaining an end-to-end integrated neural network or multi-branch neural network model. For example, these parameters may correspond to:

[0261] - a list of atoms that make up a chemical structure (e.g. a molecule),

[0262] - a list of atomic properties in a chemical structure,

[0263] - a list of ingredients in the composition,

[0264] - a list of molecules in the composition and / or

[0265] - A list of molecules corresponding to the ingredients in the composition.

[0266] An ensemble neural network or multi-branch neural network model is used to provide an output for a standardized input format. The standardized input format can correspond to a digital representation of the atoms, atomic properties, molecules, components, compositions, and / or chemical structures. Such digital representations can correspond to strings of characters. These strings can be concatenated to form a single input representing a larger-scale substance entry (e.g., multiple atoms comprising a molecule).

[0267] Figure 7 and Figures 12 to 14 An example of such input is shown.

[0268] The providing step 115 is performed using, for example, an output device 235 coupled to the I / O subsystem 220, e.g., Figure 2 The public.

[0269] In certain embodiments, this providing step 115 shows on a GUI the model's predicted results based on the defined chemical structures or compositions input to the model.

[0270] Providing step 120 can be performed via a computer interface (e.g., an API or any other digital input method). Providing step 120 can be initiated manually or automatically. The example data set can be manually assembled via the computer interface or automatically assembled from a larger example data set by a computing system.

[0271] Example data may include, for example:

[0272] - a digital representation of at least one chemical structure or composition, and

[0273] - a value expressing an odor quality or a physicochemical property associated with a chemical structure or composition.

[0274] Such odor characteristics can be, for example, the tonality of a chemical structure, the odor detection threshold of a chemical structure, the odor intensity of a chemical structure (e.g., categorizing odor intensity into four categories: no odor, weak odor, medium odor, and strong odor), and / or the "top-heart-base" of a chemical structure (e.g., categorizing the persistence of an ingredient or composition during evaporation into three ranges: top, heart, and base, where "top" refers to an ingredient or composition that can still be smelled within 15 minutes of evaporation or as determined by gas chromatography, "heart" refers to between 15 minutes and 2 hours, and "base" refers to more than 2 hours). This list is not limiting, and any odor characteristic known in the field of fragrance and flavor design and related industries can be associated with a hash vector.

[0275] Odor characteristics may correspond to:

[0276] -Anti-insect ability value,

[0277] - sensory property values,

[0278] - biodegradability value,

[0279] - antibacterial value,

[0280] - Odor detection threshold,

[0281] - Odor intensity value,

[0282] - Top-Center-Bottom values,

[0283] - Danger value,

[0284] - Biological activity of taste

[0285] - biological activity of the sense of smell,

[0286] - Biologically augment or modulate taste activity,

[0287] - Biological augmentation or modulation of olfactory activity and / or

[0288] - Olfactory smell description.

[0289] Physical chemistry may correspond to:

[0290] - boiling point value,

[0291] - melting point value,

[0292] - water solubility value,

[0293] - Henry's constant value,

[0294] - vapor pressure value,

[0295] - Volatility value or

[0296] -Headspace concentration value.

[0297] Operation 125 can be performed, for example, by a computer program executed on a computing system. In operation 125, the end-to-end integrated neural network or multi-branch neural network device is configured to be trained based on input data. In operation 125, each neural network sub-device of the end-to-end integrated neural network or multi-branch neural network device is configured with coefficients of an artificial neuron layer to provide outputs, which form an output distribution. A statistical parameter representing this distribution can be obtained and used in the activation function to be minimized.

[0298] Each neural network sub-device in a collection can be of the same type or of a different type.

[0299] In certain embodiments, at least one neural network sub-device is:

[0300] - Recurrent Neural Network Device,

[0301] -Graph Neural Network Device,

[0302] - Variational Autoencoder Neural Network Device or

[0303] -Autoencoder neural network device.

[0304] In certain embodiments, the at least two activation functions represent:

[0305] - the mean of the statistical distribution of multiple independent predictions,

[0306] - the variance of the statistical distribution of multiple independent predictions, and

[0307] - Optionally extended with additional activation functions, representing:

[0308] - Deviations in the statistical distribution of multiple independent forecasts, and / or

[0309] -The kurtosis of the statistical distribution of multiple independent predictions.

[0310] In particular embodiments, the at least one output value representing the distribution represents a dispersion of the distribution.

[0311] For example, the value may correspond to the standard deviation of the outputs of the neural network sub-device.

[0312] In certain embodiments, an end-to-end ensemble neural network or a multi-branch neural network device is trained to minimize at least one value representing the dispersion of the distribution.

[0313] The obtaining step 130 may be performed via a computer interface (eg, an API or any other digital output system). The obtained training model may be stored in a data storage device, such as a hard disk or a database.

[0314] In certain embodiments, the neural network device obtained during the obtaining step 130 is configured to additionally provide at least one value representative of the statistical dispersion of the outputs.

[0315] In a particular embodiment, at least one set of input example data corresponds to a hash vector of at least one atomic property in a chemical structure or composition, and the method further includes a conversion step 135 upstream of the execution step 110, converting the defined digitized chemical structure or composition into a set of hash vectors representing at least one atomic property of the digitized chemical structure or composition, wherein the set of hash vectors is used as input during the execution step.

[0316] A hash value corresponds to the result of a hash function, which is any function that can map data of any size to a fixed-size value. Many such functions are known to those skilled in the art, such as SHA-3, Skein, or Snefru.

[0317] This hash value can be organized into a vector that can be used by an end-to-end integrated neural network or a multi-branch neural network device.

[0318] To obtain a hash vector representing atomic properties in a chemical structure, a method comprising the following steps may be implemented:

[0319] - a receiving step of receiving, by means of a computing system, a digital representation of a chemical structure, the digital representation comprising at least one atomic property digital identifier and at least one atomic property digital identifier for said at least one atomic property digital identifier,

[0320] - a determination step of determining, by a computing system, at least one value corresponding to at least one atom or bond property of at least one atom or bond of the digital representation of the chemical structure,

[0321] - a hashing step of hashing at least one determined value by a computing system to form a unique string fingerprinting at least one atomic property digital identifier and at least one associated atomic property, and

[0322] - providing a step of providing at least one hash value to the set of neural network devices to be trained by the computing system.

[0323] The receiving step can be performed, for example, by any input device 240 suitable for a particular use case. For example, during this receiving step, a digital representation of at least one chemical, atomic, or bond structure is input into a computer interface. Such input can be purely logical, such as by connecting the computing system to another computing system using an API (application programming interface) or via a computer network. Such input can also rely on a human-computer interface, such as a keyboard, mouse, or touch screen. The mechanism used in the receiving step is not important to the scope of the present invention.

[0324] Ultimately, a digital representation of a chemical structure consists primarily of two types of data:

[0325] - an atom identifier, which corresponds to an atom in a molecular structure and is typically represented by at least one letter (e.g., "C" for carbon, "H" for hydrogen), and

[0326] - at least one relationship of at least one atom identifier, which defines whether and how at least one atom is connected to other atoms of the molecule.

[0327] This digital representation can take many forms, depending on the system. For example, the SMILES ("Simplified Molecular Input System") format is a line representation of a molecular structure that provides both of the aforementioned types of data. Another example is a molecular graph representation of a molecule. Another representation is the SDF ("Structure Data File") format, which defines atomic properties and a table of bonds. Yet another representation is a complete molecular matrix consisting of atomic numbers and an adjacency matrix defining the bonds.

[0328] Typically, the primary digital representation used in chemical reaction modeling and feature prediction is the SMILES format.

[0329] The determining step is performed by, for example, one or more hardware processors 210 (e.g. Figure 2The hardware processor 210 is configured to execute a set of instructions representing computer software. In this determining step, at least one atomic property can be read from a digital representation of the chemical structure, or determined by executing a dedicated algorithm, or obtained from dedicated third-party software.

[0330] The hashing step is performed, for example, by one or more hardware processors 210, such as Figure 2 As shown, the hardware processor 210 is configured to execute a set of instructions representing computer software.

[0331] The output of the hashing step is a given number of hash values, each representing an atom identifier and at least one associated atomic property of the identified atom. Thus, a chemical structure containing multiple atoms can be represented by a sentence containing multiple hash values. Each hash value acts as a unique fingerprint, which is particularly useful for neural network applications. This means that in a dataset, each atom can be represented by a corresponding hash key (unique fingerprint).

[0332] This hashing inherently enforces network sparseness.

[0333] In this hashing step, the hash can be composed of repeated values ​​of the property, thereby defining the value of the property in reagents, intermediates, transition states, and products.

[0334] In particular embodiments, at least one of the atomic properties of the hash represents one of the following:

[0335] - the atomic number of the corresponding atom,

[0336] - the atomic symbol of the corresponding atom,

[0337] - atomic mass,

[0338] - explicit mapping number,

[0339] - the row index in the periodic table,

[0340] - column index in the periodic table,

[0341] - the total number of hydrogen atoms,

[0342] - the implied number of hydrogens on the atom,

[0343] - the exact number of hydrogen atoms,

[0344] - the degree of the atom,

[0345] - the total degree of the atom,

[0346] - the valence state of the atoms,

[0347] - the implicit valence of the atom,

[0348] - the definite valence of the atoms,

[0349] - formal charges on atoms,

[0350] - partial charges on atoms,

[0351] - the electronegativity of the atom,

[0352] - the number of keys divided by key type,

[0353] - adjacent numbers by atomic number, wildcards,

[0354] -Adjacent number calculated by key type plus atomic number and wildcard, -Adjacent number calculated by wildcard,

[0355] - a value indicating aromaticity,

[0356] - represents the value of aliphatic atoms,

[0357] - represents the value of conjugated atoms,

[0358] - represents the value of the ring atom,

[0359] - represents the value of the macrocyclic atom,

[0360] - represents the value of the geometrically constrained atom,

[0361] - represents the value of electron-withdrawing atoms,

[0362] - represents the value of the electron-donating atom,

[0363] - represents the value of the reaction site,

[0364] - represents the value of hydrogen bond donor,

[0365] - represents the value of hydrogen acceptor,

[0366] - a value indicating multivalency as a hydrogen bond donor,

[0367] - a value indicating multivalency as a hydrogen bond acceptor,

[0368] - the number of cycles on the atom,

[0369] - ring size on the atom,

[0370] - the hybridization state of the atoms,

[0371] - values ​​representing atomic geometry,

[0372] - the number of electrons in the atomic orbital,

[0373] - the number of electrons in the lone pair,

[0374] - the state of the atomic group,

[0375] - isotopes on atoms,

[0376] -atomic centrosymmetry function,

[0377] - values ​​of relative stereochemistry as chi clockwise and chi counterclockwise,

[0378] - the value of the absolute stereochemistry,

[0379] - the value of the absolute stereochemistry,

[0380] - the value of the double bond stereochemistry,

[0381] - the preferred value of the stereochemistry,

[0382] - a value representing the positive or negative impact that an atom has on a determined training target, thereby representing a rich knowledge-based contribution,

[0383] - A value representing the positive or negative impact that an atom has on a determined training target, thereby representing the value of the knowledge-based dilution contribution and / or ring stereochemistry.

[0384] Values ​​representing the positive or negative impact of atoms on a given training target can be initialized by the user or trained as part of an auxiliary training method. These values ​​can be used at the atomic or molecular level.

[0385] Another method of hashing involves assigning a character to each atomic identity and concatenating the characters into "words." These characters may correspond to the characters in a SMILES string, with all characters not identified as chemical atom characters removed.

[0386] In particular embodiments, the at least one hashed key characteristic represents one of the following:

[0387] - key sequence,

[0388] -Key type,

[0389] -bond stereochemistry:

[0390] - the bond orientation of the tetrahedral stereochemistry,

[0391] - the bond direction of the double bond stereochemistry, or

[0392] - spatially oriented bond directions,

[0393] - for the atomic numbers of the "from" and / or "to" atoms,

[0394] - atomic symbols representing "from" and / or "to" atoms,

[0395] - dipole moment in the bond,

[0396] -Quantum chemical properties:

[0397] - electron density in the bond,

[0398] - the electronic configuration of the bond,

[0399] -Key Track,

[0400] -bond energy,

[0401] -attraction,

[0402] - repulsive force,

[0403] -key distance,

[0404] - aromatic bonds,

[0405] - aliphatic bonds,

[0406] -Ring characteristics of the bond:

[0407] - the number of rings on the bond,

[0408] - Ring size of the key,

[0409] - minimum ring size of the key,

[0410] - maximum ring size of the key,

[0411] - rotatable key,

[0412] - spatial constraint keys,

[0413] - hydrogen bonding properties,

[0414] - ionic bonding properties,

[0415] -The bond order of the reaction, including "empty" bonds, used to identify the bonds broken / formed in the reaction:

[0416] - the bond order in the reagent,

[0417] - the bond order in the intermediate product, or

[0418] -Bond order in the transition state.

[0419] The obtaining step is performed, for example, by using any output device 235 associated with the I / O subsystem 220, such as Figure 2 shown.

[0420] In a specific embodiment, the method for obtaining a hash vector may further include:

[0421] - a construction step of constructing, by a computing system, a chemical structure string fingerprint by associating at least two hash values ​​corresponding to at least two atomic properties in a single string, and

[0422] - at least one augmentation step of augmenting at least one chemical structure string fingerprint by a computing system, said augmented chemical structure string fingerprint being used during the providing step.

[0423] The construction step is performed by one or more hardware processors 210 (e.g. Figure 2 The hardware processor 210 is configured to execute a set of instructions representing computer software. During this construction step, at least two hash values ​​corresponding to at least two atom identifiers and their associated features are associated, typically by concatenating the corresponding hash values. The order of concatenation can follow a concatenation rule that prevents misinterpretation by the neural network.

[0424] Figure 7 A specific embodiment of the method is schematically shown for obtaining a hash vector representing a chemical structure. Figure 7 express:

[0425] - a receiving step 505, which receives an input chemical structure,

[0426] - a determination step 510 of determining the properties of the atoms and / or bonds in the received chemical structure, thereby annotating the vector of properties entered; the properties used here are [atomic number, degree, number of hydrogen atoms],

[0427] - a hashing step 515 which hashes the property vector into a single hash key; alternatively, a hash key identifying the atomic type may be formed by simply concatenating the values, e.g. [06, 1, 1] becomes "0611",

[0428] - a construction step 520 of constructing a vector of hash keys of atoms in atomic order, and

[0429] - A data augmentation step 525 which may apply data augmentation on the vector by changing the order of atoms; this order may comprise a vector of canonical order of atoms.

[0430] In a specific embodiment, the method 100 object of the present invention includes an expansion step 140 upstream of the providing step 120 of providing input data to the end-to-end integrated neural network or multi-branch neural network device, which expands the atom or bond relationship vector.

[0431] At least one expansion step 140 is performed, for example, by computer software running on a computing device (e.g., a microprocessor). During the expansion step 140, the order of the constituent hash values ​​of a given molecular structure is shifted by one or more factors. That is, for example, the last hash value becomes the second-to-last, the second-to-last becomes the third-to-last, the first becomes the last, or vice versa, depending on the desired expansion order.

[0432] This augmentation can increase the number of samples from the same chemical structure, thereby greatly improving the quality of the output of the neural network device.

[0433] In a particular embodiment, the expansion step 140 of expanding the atom or bond relationship vector includes a horizontal expansion step 145, which is configured to provide multiple vectors representing a single digital representation of a molecular structure or chemical reaction, each vector representing a specific representation of a canonical representation of the molecular structure or chemical reaction, and each vector is considered as a single input during the providing step.

[0434] In certain embodiments, the atom or bond relationship vector expansion step 140 includes a vertical expansion step 150 to create several sets of horizontal expansions representing unique molecular structures or chemical reactions, each set being treated as a single input in the providing step.

[0435] Such a vertical expansion step 150 can be performed by, for example, computer software executed by a computing system. This vertical expansion step 150 can be performed by grouping the horizontal expansions into individual inputs, typically by concatenating hash keys representing the atoms and / or bond characteristics of the chemical structure. These individual inputs can be the same or different, for example, by changing the order of concatenation.

[0436] Figure 6 Another representation of a specific embodiment of the method 600 object of the present invention is shown.

[0437] Figure 6 The training of the model is specifically shown, including a providing step 120 , an operating step 125 and an obtaining step 130 .

[0438] During the providing step 120 , a digital representation of the chemical structure 605 and known odor property values ​​or physicochemical property values ​​610 are used as input.

[0439] During operation 125 , the training end-to-end integrated neural network or multi-branch neural network device 615 outputs two values ​​620 and 625 representing the distribution of the individual outputs of the neural network sub-devices that make up the end-to-end integrated neural network or multi-branch neural network device 615 , such as the mean and standard deviation.

[0440] Figure 2A specific embodiment of the system 200 of the present invention is schematically shown. The system 200 is used to predict at least one physicochemical and / or odor property value of a chemical structure or composition, comprising the following mechanism 205:

[0441] - a definition mechanism that defines a digital representation of a chemical structure or composition on a computer interface,

[0442] - an execution unit which, by means of a computing device, executes an end-to-end trained ensemble neural network or multi-branch neural network model according to the defined digital representation, thereby predicting at least one physicochemical and / or odor property value of a chemical structure or composition,

[0443] - providing means for providing on a computer interface the value of at least one physicochemical and / or odorous property of a chemical structure or composition,

[0444] Also included are the following institutions 205:

[0445] - providing a mechanism for providing a set of example data to an end-to-end integrated neural network or a multi-branch neural network, comprising: at least one set of inputs corresponding to digital representations of chemical structures or compositions; and at least one set of outputs corresponding to physicochemical and / or odor properties associated with the set of inputs, the end-to-end integrated neural network or the multi-branch neural network comprising:

[0446] - a plurality of neural network sub-devices, each sub-device being configured to provide independent predictions based on example data,

[0447] - a layer configured to output at least one value based on or representing a distribution of the independent predictions, and

[0448] - the layer comprises a sampling device configured to output at least one random value according to a probability distribution representing a distribution of independent predictions, the outputted random value being calculated in a differentiable manner and used for backpropagation within an end-to-end integrated neural network or a multi-branch neural network device,

[0449] - an operating mechanism that operates an end-to-end integrated neural network or a multi-branch neural network device based on the set of example data, and

[0450] - obtaining means for obtaining a trained end-to-end ensemble neural network or multi-branch neural network model configured to predict the physicochemical and / or odor properties of an input digital representation of a chemical structure or composition.

[0451] Figure 2 FIG2 is a block diagram showing an example computer system that can be used to implement the embodiments. Figure 2In the example of FIG, a computer system 205 and instructions for implementing the disclosed techniques in hardware, software, or a combination of hardware and software are represented schematically (e.g., boxes and circles) with the same level of detail that is typically used by ordinary technicians in the field to which this disclosure relates when communicating about computer architecture and computer system implementations.

[0452] Computer system 205 includes an input / output (IO) subsystem 220, which may include a bus and / or other communication mechanism for transferring information and / or instructions between components of computer system 205 via electronic signal paths. I / O subsystem 220 may include an I / O controller, a memory controller, and at least one I / O port. Electronic signal paths are represented schematically in the figure, for example, as lines, single-direction arrows, or double-direction arrows.

[0453] At least one hardware processor 210 is coupled to the I / O subsystem 220 for processing information and instructions. The hardware processor 210 may include, for example, a general-purpose microprocessor or microcontroller, and / or a specialized microprocessor, such as an embedded system, a graphics processing unit (GPU), a digital signal processor, or an ARM processor. The processor 210 may include an integrated arithmetic logic unit (ALU) or may be coupled to a separate ALU.

[0454] The computer system 205 includes one or more memory units 225, such as main memory, coupled to the I / O subsystem 220 for electronically storing data and instructions to be executed by the processor 210. The memory 225 may include volatile memory, such as various forms of random access memory (RAM) or other dynamic storage devices. The memory 225 may also be used to store temporary variables or other intermediate information during the execution of instructions by the processor 210. When these instructions are stored in a non-transitory computer-readable storage medium accessible to the processor 210, the computer system 205 can be made into a special-purpose machine customized to perform the operations specified in the instructions.

[0455] Computer system 205 also includes nonvolatile memory, such as read-only memory (ROM) 230 or other static storage device, coupled to I / O subsystem 220 for storing information and instructions for processor 210. ROM 230 may include various forms of programmable ROM (PROM), such as erasable PROM (EPROM) or electrically erasable PROM (EEPROM). Persistent memory unit 215 may include various forms of nonvolatile RAM (NVRAM), such as flash memory, solid-state memory, magnetic disks, or optical disks (e.g., CD-ROMs or DVD-ROMs), and may be coupled to I / O subsystem 220 for storing information and instructions. Memory 215 is an example of a non-transitory computer-readable medium that may be used to store instructions and data that, when executed by processor 210, result in the execution of a computer-implemented method, thereby performing the techniques described herein.

[0456] The instructions in memory 225, ROM 230, or storage device 215 may include one or more groups of instructions organized into modules, methods, objects, functions, routines, or calls. These instructions may be organized into one or more computer programs, operating system services, or applications (including mobile applications). These instructions may include: operating system and / or system software; one or more libraries for supporting multimedia, programming, or other functionality; data protocol instructions or stacks for implementing TCP / IP, HTTP, or other communication protocols; file format processing instructions for parsing or rendering files encoded using HTML, XML, JPEG, MPEG, or PNG; user interface instructions for rendering or interpreting commands for a graphical user interface (GUI), command line interface, or text user interface; and application software, such as an office suite, an internet access application, a design and manufacturing application, a graphics application, an audio application, a software engineering application, an educational application, a game, or other application. These instructions may implement a web server, a web application server, or a web client. The instructions may be organized into a presentation layer, an application layer, and a data storage layer, such as a relational database system using Structured Query Language (SQL) or no-SQL, an object store, a graph database, a flat file system, or other data storage.

[0457] The computer system 205 can be coupled to at least one output device 235 via the I / O subsystem 220. In one embodiment, the output device 235 is a digital computer display. Examples of displays that can be used in various embodiments include a touch screen display, a light emitting diode (LED) display, a liquid crystal display (LCD), or an electronic paper display. The computer system 205 can also include other types of output devices 235 as an alternative to or in addition to the display device. Examples of other output devices 235 include a printer, a receipt printer, a plotter, a projector, a sound card or video card, a speaker, a buzzer or piezoelectric device or other sound device, a light or LED or LCD display, a tactile device, an actuator, or a servo device.

[0458] At least one input device 240 is coupled to I / O subsystem 220 for communicating signals, data, command selections, or gestures to processor 210. Examples of input device 240 include touch screens, microphones, still and video digital cameras, alphanumeric and other keys, keypads, keyboards, graphics tablets, image scanners, joysticks, clocks, switches, buttons, dials, and sliders.

[0459] Another type of input device is a control device 245, which can perform cursor control or other automated control functions, such as navigating a graphical interface on a display screen, as an alternative to or in addition to the input functions. The control device 245 can be a touchpad, mouse, trackball, or cursor direction keys, which communicate directional information and command selections to the processor 210 and control the movement of the cursor on the display 235. The input device can have at least two degrees of freedom in two axes, namely a first axis (e.g., x-axis) and a second axis (e.g., y-axis), which enables the device to specify a position within a plane. Another type of input device is a wired, wireless, or optical control device, such as a joystick, control rod, console, steering wheel, pedals, gear shift mechanism, or other type of control device. The input device 240 can include a combination of multiple different input devices, such as a camera and a depth sensor.

[0460] In another embodiment, the computer system 205 may comprise an Internet of Things (IoT) device, in which one or more of the output device 235, the input device 240, and the control device 245 are omitted. Alternatively, in such an embodiment, the input device 240 may comprise one or more cameras, motion detectors, thermometers, microphones, earthquake detectors, other sensors or detectors, measuring devices, or encoders, while the output device 235 may comprise a dedicated display (e.g., a single-line LED or LCD display), one or more indicators, display panels, meters, valves, solenoids, actuators, or servo devices.

[0461] The computer system 205 can implement the techniques described herein using custom hardwired logic, at least one ASIC or FPGA, firmware, and / or program instructions or logic that, when loaded and used or executed in conjunction with the computer system, cause or program the computer system to operate as a special-purpose machine. According to one embodiment, the techniques described herein are performed by the computer system 205 in response to the processor 210 executing at least one sequence of at least one instruction contained in the main memory 225. These instructions can be read into the main memory 225 from another storage medium (e.g., the memory 215). Execution of the sequences of instructions contained in the main memory 225 causes the processor 210 to perform the process steps described herein. In alternative embodiments, hardwired circuitry can be used in place of or in combination with software instructions.

[0462] As used herein, the term "storage media" refers to any non-transitory medium for storing data and / or instructions that cause a machine to operate in a specific manner. Such storage media may include non-volatile media and / or volatile media. Non-volatile media include, for example, optical or magnetic disks, such as memory 215. Volatile media include dynamic memory, such as memory 225. Common forms of storage media include, for example, hard disks, solid-state drives, flash drives, magnetic data storage media, any optical or physical data storage media, memory chips, and the like.

[0463] Storage media are distinct from transmission media, but can be used in conjunction with them. Transmission media participate in the transfer of information between storage media. Examples of transmission media include coaxial cables, copper wire, and optical fiber, including the wiring that makes up the bus of I / O subsystem 220. Transmission media can also take the form of acoustic or optical waves, such as those generated during radio and infrared data communications.

[0464] Various forms of media may be used to transmit at least one sequence of instructions to processor 210 for execution. For example, the instructions may initially be stored on a disk or solid-state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and use a modem to send the instructions over a communication link (e.g., optical fiber, coaxial cable, or telephone line). A modem or router local to computer system 205 may receive the data on the communication link and convert the data into a format readable by computer system 205. For example, a receiver such as a radio frequency antenna or an infrared detector may receive data carried in a wireless or optical signal, and appropriate circuitry may provide the data to I / O subsystem 220, such as by placing the data on a bus. I / O subsystem 220 transmits the data to memory 225, from which processor 210 retrieves and executes the instructions. The instructions received by memory 225 may optionally be stored in memory 215 before or after execution by processor 210.

[0465] Computer system 205 also includes a communication interface 260 coupled to bus 220. Communication interface 260 provides a bidirectional data communication coupling with network links 265, which are directly or indirectly connected to at least one communication network, such as network 270 or a public or private cloud on the Internet. For example, communication interface 260 can be an Ethernet network interface, an integrated services digital network (ISDN) card, a cable modem, a satellite modem, or a modem for providing a data communication connection to a corresponding type of communication line, such as an Ethernet cable or any type of metallic cable, fiber optic line, or telephone line. Network 270 broadly represents a local area network (LAN), a wide area network (WAN), a campus network, the internet, or any combination thereof. Communication interface 260 can include a LAN card for providing a data communication connection to a compatible LAN; a cellular radiotelephone interface for sending or receiving cellular data over a wired connection according to a cellular radiotelephone wireless network standard; or a satellite radio interface for sending or receiving digital data over a wired connection according to a satellite wireless network standard. In any such implementation, communication interface 260 sends and receives electrical, electromagnetic, or optical signals over a signal path that carries digital data streams representing various types of information.

[0466] Network link 265 typically provides electrical, electromagnetic, or optical data communication to other data devices, directly or through at least one network, using, for example, satellite, cellular, Wi-Fi, or Bluetooth technology. For example, network link 265 may be connected to host 250 via network 270 .

[0467] In addition, network link 265 can provide connectivity through network 270 or establish connections with other computing devices through interconnected devices and / or computers operated by Internet Service Provider (ISP) 275. ISP 275 provides data communication services through a global packet data communication network represented by Internet 280. Server computer 255 can be connected to Internet 280. Server 255 refers broadly to any computer, data center, virtual machine or virtual computing instance (whether or not with a hypervisor), or a computer executing a containerized program system (such as Docker or Kubernetes). Server 255 can represent an electronic digital service implemented using multiple computers or instances and accessed and used by transmitting Web service requests, uniform resource locator (URL) strings containing parameters in HTTP payloads, API calls, application service calls, or other service calls. Computer system 205 and server 255 can form elements of a distributed computing system that includes other computers, processing clusters, server farms, or other computer organizations that collaborate to perform tasks or execute applications or services. Server 255 may include one or more sets of instructions organized as modules, methods, objects, functions, routines, or calls. These instructions may be organized as one or more computer programs, operating system services, or applications (including mobile applications). These instructions may include: operating system and / or system software; one or more libraries for supporting multimedia, programming, or other functionality; data protocol instructions or stacks for implementing TCP / IP, HTTP, or other communication protocols; file format processing instructions for parsing or rendering files encoded using HTML, XML, JPEG, MPEG, or PNG; user interface instructions for rendering or interpreting graphical user interfaces (GUIs), command-line interfaces, or text user interface commands; and application software, such as office suites, internet access applications, design and manufacturing applications, graphics applications, audio applications, software engineering applications, educational applications, games, or other applications. Server 255 may include a web application server that hosts a presentation layer, an application layer, and a data storage layer, such as a relational database system using Structured Query Language (SQL) or non-SQL, an object store, a graph database, a flat file system, or other data storage.

[0468] Computer system 205 can send messages and receive data and instructions, including program code, through the network, network link 265, and communication interface 260. In the Internet example, server 255 can transmit the requested code for an application program through Internet 280, ISP 275, local network 270, and communication interface 260. The received code can be executed by processor 210 upon receipt and / or stored in memory 215 or other non-volatile storage for later execution.

[0469] The execution of instructions described in this section can implement a process, which exists as an instance of an executing computer program and consists of the program code and its current activity. Depending on the operating system (OS), a process can consist of multiple threads of execution that execute instructions concurrently. In this context, a computer program is a passive collection of instructions, while a process is the actual execution of those instructions. Multiple processes can be associated with the same program; for example, opening multiple instances of the same program often means that multiple processes are executing. Multitasking can be implemented, allowing multiple processes to share the processor 210. Although each processor 210 or processor core executes a single task at a time, the computer system 205 can be programmed to implement multitasking, allowing each processor to switch between executing tasks without waiting for each task to complete. In one embodiment, switching can be performed when a task performs an input / output operation, when a task indicates that a switch is possible, or when a hardware interrupt is issued. Time-sharing can be implemented, allowing multiple processes to execute concurrently by rapidly performing context switches, thereby achieving fast responsiveness in interactive user applications. In one embodiment, to ensure security and reliability, the operating system can prevent direct communication between independent processes, providing strictly coordinated and controlled inter-process communication capabilities.

[0470] about Figure 1 Specific uses of the system 200 object of the present invention are disclosed.

[0471] Figure 4 A series of steps of the method 400 of the present invention are schematically shown. The method 400 is used to efficiently assemble a chemical structure or composition, comprising:

[0472] - Execute step 405, which is performed as follows Figure 1 The method shown, and

[0473] An assembling step 410 of assembling a chemical structure or composition related to the output obtained during the obtaining step 115 .

[0474] The assembly step 410 is used to achieve the physicalization of the composition. The assembly step 410 can be performed in a variety of ways, such as in a laboratory or a chemical plant.

[0475] Figure 8 A specific implementation example of the method 800 object of the present invention is schematically shown. The method 800 for training an integrated neural network or a multi-branch neural network device is similar to the training performed by the end-to-end integrated neural network or multi-branch neural network device used in the method 100 object of the present invention. The method 800 includes:

[0476] - an input step 805 of inputting an augmented atom and / or bond property hash key from a sample set of chemical structure digital identifiers associated with known outputs representing at least one physicochemical and / or odor property associated with the atom and / or bond property hash key,

[0477] - an embedding or marking step 810, which embeds or marks the input,

[0478] - an operation 815 of operating a set of recurrent neural network devices based on the input,

[0479] - an operation 820 of operating an attention layer at the output of the operation 815 of operating a set of recurrent neural network devices,

[0480] - an operation step 825 of operating a flattening layer on the output of the operation step 820 of operating the attention layer,

[0481] - an operation 830 of operating a multi-layer perceptron ("MLP") layer on the output of the flattening layer, and

[0482] - An output step 835 which outputs the value of the target odor property and / or physicochemical property.

[0483] Implementation parameters for a particular embodiment may be:

[0484]

[0485] The number N in the table represents the number of points (e.g., the input batch size). In this architecture, chemical structures are represented as expanded 2D chips and then transformed using embedding and recurrent neural network layers. Attention layers perform feature selection. The MLP portion of the network is a fully connected neural network with an activation function.

[0486] Figures 9 to 11 Shown Figure 8 Performance of the shown architecture relative to three different targets:

[0487] - Figure 9 The performance of the odor detection threshold ("ODT") is shown.

[0488] - Figure 10 The performance of volatility and

[0489] - Figure 11 The properties are shown as LogVP (ie the logarithm of the vapor pressure).

[0490] It will be understood that the present invention also aims to provide a computer-implemented integrated neural network or multi-branch neural network device, wherein the integrated neural network or multi-branch neural network device is obtained by any deformation of the object of the computer-implemented method 300 of the present invention.

[0491] It should be understood that the present invention is also directed to a computer program product comprising instructions for performing the steps of the method 300 object of the present invention when executed on a computer.

[0492] It is understood that the present invention is also directed to a computer-readable medium storing instructions that, when executed on a computer, perform the steps of the method 300 object of the present invention.

[0493] Figure 12 A training architecture 1200 is schematically shown for selecting chemical structures from a set of chemical structures that provide a specific characteristic, such as an insect repellency value above a determined threshold.

[0494] This architecture 1200 includes:

[0495] - as input 1205, a hash vector of at least one atomic property in a chemical structure, and at least one set of outputs corresponding to odor properties associated with the set of inputs,

[0496] - an ensemble neural network or multi-branch neural network 1210 comprising a set of recurrent neural networks for generating embeddings 1215,

[0497] - Two alternative and compatible routes can then be implemented:

[0498] - In the first route, a multivariate statistical algorithm 1220 may be used, supplemented by a numerical domain eccentricity estimation algorithm 1225,

[0499] - In the second route, temporal distribution of embedding 1230 is performed to obtain alternative input 1235, supplemented by an ensemble 1240 of neural networks using Tanimoto neural networks, such as disclosed in the present invention.

[0500] Figure 13 A training architecture 1300 is schematically shown for classifying chemical structures providing a specific characteristic (eg, a biodegradability value) from a set of chemical structures.

[0501] This architecture 1300 includes:

[0502] - as input 1305, a hash vector of at least one atomic property in a chemical structure, and at least one set of outputs corresponding to odor properties associated with the set of inputs,

[0503] - an ensemble neural network or multi-branch neural network 1310 comprising a set of recurrent neural networks for generating embeddings 1315,

[0504] - Two alternative and compatible routes can then be implemented:

[0505] - In the first route, a multivariate statistical algorithm 1320 may be used, supplemented by a numerical domain eccentricity estimation algorithm 1325,

[0506] - In the second route, temporal distribution of embedding 1330 is performed, supplemented by an ensemble 1335 of neural networks using classification neural networks, such as disclosed in the present invention.

[0507] Figure 14 A training architecture 1400 is schematically shown for predicting the value of a specific feature, such as an odor detection threshold, for a set of chemical structures.

[0508] This architecture 1400 includes:

[0509] - as input 1405, a hash vector of at least one atomic property in a chemical structure, and at least one set of outputs corresponding to odor properties associated with the set of inputs,

[0510] - an ensemble neural network or multi-branch neural network 1410 comprising a set of recurrent neural networks for generating embeddings 1415,

[0511] - Two alternative and compatible routes can then be implemented:

[0512] - In the first route, a multivariate statistical algorithm 1420 may be used, supplemented by a numerical domain eccentricity estimation algorithm 1425,

[0513] - In the second route, an ensemble 1430 of single-task or multi-task recurrent neural networks is used, such as disclosed in the present invention.

[0514] It will be appreciated that the present invention can be used as a filtering technique, using any predicted physicochemical and / or odorant properties to tag digital identifiers of molecules or ingredients in a database that are selected by flavorists and perfumers as worthy of exploration.

[0515] It will be appreciated that the present invention can use molecular pairs as input to predict the proximity of molecules in a molecular pair, or use observed differences in a molecular pair for regression or classification.

[0516] It will be appreciated that the present invention may be used as a classifier relating values ​​of physicochemical and / or odor properties of chemical structures or compositions.

[0517] Figure 15A particular architecture is shown, highlighting the performance of this classifier.

[0518] In machine learning, the chi-squared test is often used to evaluate the performance of classification models. For example, suppose we have a binary classification problem where we need to predict whether a patient has a disease or not. We can use the chi-squared test to determine if our model performs better than chance by comparing the predicted class distribution to the expected class distribution.

[0519] In ensemble learning, multiple models are combined to improve overall performance. The chi-squared test can be used to evaluate the ensemble's performance. Ensemble learning is a popular technique in machine learning that trains and combines multiple models to improve overall performance. Using multiple models can reduce the risk of overfitting and improve model robustness.

[0520] In a classification model ensemble, each model makes independent predictions on the input data, and the final prediction is the combination of the predictions from all models. The chi-squared test can be used to evaluate the performance of the ensemble by comparing the distribution of the ensemble's predicted categories with the expected distribution of the categories. If the ensemble performs better than any individual model, the ensemble is considered effective.

[0521] Overall, the chi-squared test is a powerful tool for evaluating the performance of machine learning models and ensembles. By using the chi-squared test, we can make informed decisions about which models to use and how to improve them.

[0522] Forced-choice models are an example of a contrastive classification task, where the goal is to identify the correct example from a set of alternatives. These tasks are common in the real world, such as finding the correct answer on a multiple-choice exam or identifying a specific object from a group of similar objects. In scientific research, results are often evaluated in a relative context, that is, by comparing two or more candidate examples. Therefore, it has been hypothesized that contrastive neural networks trained to select the more promising example from a set of alternatives might provide valuable models.

[0523] By creating pairs, triplets, or other alternatives in contrastive neural networks, data can be augmented. In fact, for regression tasks, the data can be expanded from N to N 2 -N pairs, or expand to (N 2 -N) / 2 pairs, only the lower or upper half of the matrix needs to be considered. Alternatively, for problems with a low hit rate, the hit result can be combined with one or more non-hit results in a forced-choice classification. In the latter experiment, the model is trained to detect hit molecules from the suggested options. Another benefit of these contrastive networks is that they can create balanced sets. Indeed, one would expect the lower values ​​to be evenly distributed across the number of alternatives.

[0524] To solve this task, a neural network ensemble with independent voting can be used, where each model in the ensemble makes an independent prediction for the input data. The final prediction is then made by combining the predictions of all models. By using a model ensemble, the risk of overfitting can be reduced and the robustness of the model can be improved.

[0525] After making a prediction, the statistical significance of the decision can be measured using a chi-squared test. In this case, the predicted class distribution can be compared with the expected class distribution, which is evenly distributed across the three samples. If the chi-squared test shows that the predicted class distribution is significantly different from the expected class distribution, then we can conclude that the ensemble model performed well and was able to correctly identify the correct examples from the input X.

[0526] Overall, using ensemble neural networks for individual voting and chi-square tests is an effective approach for comparative classification tasks, helping to improve the accuracy and robustness of the model. This approach allows us to make informed decisions about which samples are correct and which are incorrect, thereby improving our ability to recognize and classify objects in real-world scenarios.

[0527] To perform such calculations, a SMILES string containing explicit-implicit hydrogen atoms can be used. For example, take the toluene molecule. The explicit SMILES for toluene is written as "[CH3][c]1[cH][cH][cH][cH][cH][cH]1", which can be labeled by grouping atoms defined by characters in square brackets from [to]. All other characters, such as ring index 1, can be labeled as a single character. Therefore, the tokenized SMILES for toluene is "[CH3][c]1[cH][cH][cH][cH][cH][cH]1". Similarly, the explicit SMILES for glutamic acid is "[NH2][CH]([CH2][CH2][C](=[O])[OH])[C](=[O])[OH]", which can be labeled by individually labeling bonds, for example = represents a double bond, and branches, i.e. (and). The labeled SMILES for glutamic acid is "[NH2][CH]([CH2][CH2][C](=[O])[OH])[C](=[O])[OH]".

[0528] Forced-choice classification is run using a network layout where the same embedding, GRU, and attention and latent layers are applied to all input entries, followed by a learnable contrastive layer that creates differences between all pairs.

[0529] Figure 15This figure shows the architectural layout of a contrastive classifier, which selects the molecule with the lowest molecular weight. The input is a labeled vector of integers, followed by a Keras Embedding layer, a Keras GRU layer, and an Attention layer. The equal sign between these layers indicates that the same layer is applied to both entries. The trainable contrastive layer generates the difference between the outputs of Attention 1 and Attention 2. A multilayer perceptron with dropout is used for the classification task. This model is repeated N times to create an ensemble neural network. The values ​​(None, x) and (None, x, y) indicate the output shape of the layer.

[0530] A dataset obtained from NIST can be used to train such a contrastive classifier. The data is divided into a training set containing 8,518 molecules, a validation set containing 819 molecules, and a test set containing 772 molecules. To track the effectiveness of training at each epoch, the training and validation datasets are used. To train the network, 45,458 pairs of molecules are constructed with a maximum difference of 14.02 g / mol, corresponding to the mass of a CH2 group in the molecule. The classifier can be trained to detect the molecule with the highest molecular weight. Note that any numerical objective can be trained, including linear retention index, volatility, or odor detection threshold. The validation set can contain multiple pairs of molecules with a maximum difference of 14.02 g / mol. Training can be performed using 46 iterations per epoch, with a batch size of 1,000 pairs per iteration. The model is trained using the average binary cross-entropy calculated across all models in the ensemble.

[0531] Once completed, performance can be tested on a test set consisting of pairs of molecules with a maximum intermolecular difference of 14.02 g / mol. The performance results are shown in Table 1 below. Given that this model performs a relative classification task and the user requires the position of the lowest molecular weight, only the accuracy of the results is reported. In Table 1, the p-values ​​are calculated using a chi-squared test on the votes generated by the set. If the p-value for the proportion of votes is below 0.05, the result is considered deterministic. As can be clearly seen in Table 1, the results with deterministic entries significantly outperform those without deterministic entries.

[0532] Table 1

[0533] quantity accuracy All data 4554(100%) 98.0+ / -0.2% Certainty p < 0.05 4227(92.8%) 99.8+ / -0.1% Uncertainty p>0.05 327(8.4%) 74.9+ / -2.4%

[0534] In summary, using an ensemble model for classification can significantly improve the confidentiality of the results. The prediction results are composed of multiple class voting results and prediction confidence indicators (Table 2).

[0535] Table 2

[0536]

[0537] In summary, the methods presented in this article can be applied to both relative and absolute classification tasks. The example above illustrates a relative classification task, where the task is to learn to select molecules with higher molecular weights. In this type of classification, the regression task is transformed into a comparative classification task. In an absolute classifier, the ensemble model is required to predict the class defined in the data, such as the task of detecting digits in images, as performed in the MNIST dataset.

Claims

1. A method (100) for predicting at least one physicochemical and / or odor property value of a chemical structure or composition, comprising the following steps: - a definition step (105) of defining a digital representation of a chemical structure or composition on a computer interface, - performing a step (110) of executing, by means of a computing device, an end-to-end trained ensemble neural network or multi-branch neural network model according to the defined digital representation, in order to predict at least one physicochemical and / or odor property value of the chemical structure or composition, - a providing step (115) of providing on a computer interface at least one physicochemical and / or odor property value of a chemical structure or composition, It is characterized in that it comprises the following steps: - providing a step (120) of providing a set of example data to an end-to-end integrated neural network or a multi-branch neural network, comprising: at least one set of inputs corresponding to a digital representation of a chemical structure or composition; and at least one set of outputs corresponding to physicochemical and / or odor properties associated with the set of inputs, the end-to-end integrated neural network or the multi-branch neural network comprising: - a plurality of neural network sub-devices, each sub-device being configured to provide independent predictions based on example data, - a layer configured to output at least one value based on or representing a distribution of the independent predictions, and - the layer comprises a sampling device configured to output at least one random value according to a probability distribution representing a distribution of independent predictions, the outputted random value being calculated in a differentiable manner and used for backpropagation within an end-to-end integrated neural network or a multi-branch neural network device, - an operating step (125) of operating an end-to-end integrated neural network or a multi-branch neural network device based on the set of example data, and - an obtaining step (130) of obtaining a trained end-to-end integrated neural network or multi-branch neural network model configured to predict the physicochemical and / or odor properties of the input digital representation of the chemical structure or composition.

2. The method (100) according to claim 1, wherein: At least one set of inputs of the example data corresponds to a hash vector of at least one atomic property in the chemical structure or composition, the method further comprising a conversion step (135) upstream of the execution step (110), which converts the defined digitized chemical structure or composition into a set of hash vectors representing at least one atomic property of the digitized chemical structure or composition, the set of hash vectors being used as input during the execution step.

3. The method (100) according to claim 2, wherein: The hash vector of at least one atomic property represents one of the following: - the atomic number of the corresponding atom, - the atomic symbol of the corresponding atom, - atomic mass, - explicit mapping number, - the row index in the periodic table, - column index in the periodic table, - the total number of hydrogen atoms on the atom, - the implied number of hydrogens on the atom, - the exact number of hydrogen atoms, - the degree of the atom, - the total degree of the atom, - the valence state of the atoms, - the implicit valence of the atom, - the definite valence of the atoms, - formal charges on atoms, - partial charges on atoms, - the electronegativity of the atom, - the number of keys calculated by key type, - adjacent numbers calculated by atomic number, wildcard, - the number of neighbors calculated by bond type plus atomic number, wildcard, - the number of adjacent characters counted by wildcards, - a value indicating aromaticity, - represents the value of aliphatic atoms, - represents the value of conjugated atoms, - represents the value of the ring atoms, - represents the value of the macrocyclic atom, - represents the value of the geometrically constrained atom, - represents the value of electron-withdrawing atoms, - represents the value of the electron-donating atom, - represents the value of the reaction site, - represents the value of hydrogen bond donor, - represents the value of hydrogen acceptor, - a value indicating multivalency as a hydrogen bond donor, - a value indicating multivalency as a hydrogen bond acceptor, - the number of cycles on the atom, - ring size on the atom, - the hybridization state of the atoms, - values ​​representing atomic geometry, - the number of electrons in the atomic orbital, - number of lone pairs, - the state of the atomic group, - isotopes on atoms, -atomic centrosymmetry function, - as the relative stereochemical value of chi clockwise and chi counterclockwise, - the value of the absolute stereochemistry, - the value of the absolute stereochemistry, - the value of the double bond stereochemistry, - values ​​that determine stereochemical preferences, - a value representing the positive or negative impact that an atom has on the identified training target, thereby representing a rich knowledge-based contribution, - a value representing the positive or negative impact of the atom on the identified training target, thereby representing the knowledge-based diluted contribution and / or - Value of the ring stereochemistry.

4. The method (100) according to any one of claims 2 or 3, wherein: At least one hash vector of the key feature represents one of the following: - key sequence, -Key type, -bond stereochemistry: - the bond orientation of the tetrahedral stereochemistry, -bond direction of the double bond stereochemistry or - spatial orientation of the bond direction, - the atomic number of the "from" and / or "to" atom, - the atomic symbols of the "from" and / or "to" atoms, - dipole moment in the bond, -Quantum chemical properties: - electron density in the bond, - the electronic configuration of the bond, -Key track, -bond energy, -attraction, - repulsive force, -key distance, - aromatic bonds, - aliphatic bonds, -Ring characteristics of the bond: - the number of rings on the bond, - Ring size of the key, - minimum ring size of the key, - maximum ring size of the key, - rotatable key, - spatial constraint keys, - Hydrogen bonding properties, - ionic bond characteristics, -The bond order of the reaction, including "empty" bonds, used to identify the bonds broken / formed in the reaction: - the bond order in the reagent, - the bond order in the intermediate, or Bond order in the transition state.

5. The method (100) according to any one of claims 1 to 4, wherein: The at least one output value representing the distribution represents a dispersion of the distribution.

6. The method (100) according to claim 5, wherein: The end-to-end ensemble neural network or the multi-branch neural network device is trained to minimize at least one value representing the dispersion of the distribution.

7. The method (100) according to any one of claims 1 to 6, wherein: At least one odor characteristic indicating: -Anti-insect ability value, - sensory property values, - biodegradability value, - antibacterial value, - Odor detection threshold, - Odor intensity value, - Top-Center-Bottom values, - Danger value, - Biological activity of taste - biological activity of the sense of smell, - Biologically augment or modulate taste activity, - Biologically augment or modulate olfactory activity, and / or - Olfactory smell description.

8. The method (100) according to any one of claims 1 to 7, wherein: At least one physical property represents: - boiling point value, - melting point value, - water solubility value, - Henry's constant value, - vapor pressure value, - Volatility value or -Headspace concentration value.

9. The method (100) according to any one of claims 1 to 8, wherein At least one neural network device is: - Recurrent Neural Network Device, -Graph Neural Network Device, - Variational Autoencoder Neural Network Device or -Autoencoder neural network device.

10. The method (100) according to any one of claims 1 to 9, comprising: An atom or bond relationship vector expansion step (140) is upstream of a providing step (120) of providing input data to an end-to-end integrated neural network or a multi-branch neural network device.

11. The method (100) according to claim 10, wherein: The atom or bond relationship vector expansion step (140) includes a horizontal expansion step (145) configured to provide several vectors representing a single digital representation of the molecular structure or composition, each vector representing a specific representation of the canonical representation of the molecular structure or composition, and during the providing step, each vector is considered as a single input.

12. The method (100) according to claim 11, wherein: The atom or bond relationship vector expansion step (140) includes a vertical expansion step (150) that creates several groups of several horizontal expansions representing unique molecular structures or compositions, each group being treated as a single input during the providing step.

13. A method (400) for efficiently assembling a chemical structure or composition, characterized in that It includes: - performing a step (405) of performing a method according to any one of claims 1 to 12, and - an assembling step (410) of assembling a chemical structure or composition related to the output obtained in the obtaining step (115).

14. A system (200) for predicting at least one physicochemical and / or odor property value of a chemical structure or composition, comprising the following means (205): - a definition mechanism that defines a digital representation of a chemical structure or composition on a computer interface, - an execution unit which, by means of a computing device, executes an end-to-end trained ensemble neural network or multi-branch neural network model according to the defined digital representation, thereby predicting at least one physicochemical and / or odor property value of a chemical structure or composition, - providing means for providing on a computer interface the value of at least one physicochemical and / or odorous property of a chemical structure or composition, It is characterized in that It includes the following institutions (205): - providing a mechanism for providing a set of example data to an end-to-end integrated neural network or a multi-branch neural network, comprising: at least one set of inputs corresponding to digital representations of chemical structures or compositions; and at least one set of outputs corresponding to physicochemical and / or odor properties associated with the set of inputs, the end-to-end integrated neural network or the multi-branch neural network comprising: - a plurality of neural network sub-devices, each sub-device being configured to provide independent predictions based on example data, - a layer configured to output at least one value based on or representing a distribution of the independent predictions, and - the layer comprises a sampling device configured to output at least one random value according to a probability distribution representing a distribution of independent predictions, the outputted random value being calculated in a differentiable manner and used for backpropagation within an end-to-end integrated neural network or a multi-branch neural network device, - an operating mechanism that operates an end-to-end integrated neural network or a multi-branch neural network device based on the set of example data, and - obtaining means for obtaining a trained end-to-end integrated neural network or multi-branch neural network model configured to predict the physicochemical and / or odor properties of an input digital representation of a chemical structure or composition.