Ascertaining the stability of packaged formulations

EP4588053A1Pending Publication Date: 2025-07-23BAYER AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2023765235
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-14
Filing Date
2023-09-05
Publication Date
2025-07-23

AI Technical Summary

Technical Problem

Current methods for determining the stability of formulations in contact with materials are time-consuming and expensive, and existing machine learning models, such as those for predicting shelf life, do not account for the specific influences between formulations and materials, limiting their applicability across different types of formulations.

Method used

A computer-implemented method using a machine learning model trained on reference formulations with measured stability characteristics to predict stability features for new systems comprising formulations and materials, by generating numerical representations of components based on their amounts and physical/chemical properties, and adjusting model parameters to minimize prediction deviations.

Benefits of technology

This approach enables efficient prediction of stability characteristics for formulations and materials, reducing the need for extensive testing and improving the suitability assessment of materials as containers or seals, thereby optimizing storage and use conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000029_0000
    Figure 00000029_0000
  • Figure 00000030_0000
    Figure 00000030_0000
  • Figure 00000031_0000
    Figure 00000031_0000
Patent Text Reader

Abstract

The invention concerns the process of ascertaining the stability of systems comprising a formulation and at least one material which is in contact with a formulation using a machine learning model. The invention relates to a method for training the machine learning model, to a computer-implemented method for predicting at least one stability characteristic of the system using the trained model, and to a computer system and a computer program product for carrying out the prediction method.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Determination of the stability of packaged formulations

[0002] The present invention relates to determining the stability of systems comprising a formulation and at least one material in contact with the formulation using a machine learning model. The present invention relates to a method for training the machine learning model, a computer-implemented method for predicting at least one stability characteristic of the system using the trained model, and a computer system and a computer program product for executing the prediction method.

[0003] Formulations are ubiquitous. We encounter them in the form of cosmetic products, pharmaceuticals, pesticides, and many other preparations, excipients, intermediates, and / or finished products.

[0004] A formulation is manufactured from defined quantities of substances according to a defined recipe. A formulation is usually used to bring one or more components of the formulation into a form that fulfils a specific purpose. In the case of crop protection products (such as insecticides, fungicides and / or herbicides), one or more active ingredients are usually brought into a form with formulation aids that improves the biological effect of the active ingredient(s) in a target organism compared to the pure active ingredient(s) and / or makes the finished product usable for a defined purpose (e.g. application of a crop protection product using a specific application technique) and / or ensures the chemical stability of the active ingredient(s) under defined conditions. There are other purposes for which a formulation is created.

[0005] Formulations are stored and / or transported in various containers and come into contact with various materials. With newly developed formulations, the question arises as to which materials are suitable for a container, seal, or similar formulation. It is conceivable that components of the formulation could damage a material, for example, by penetrating it and causing it to swell. It is also conceivable that substances from a material (container, seal, etc.) could migrate into the formulation and contaminate it and / or have negative effects on the formulation.

[0006] Tests are typically conducted to determine the suitability of materials for formulations. Many tests follow standardized procedures. Such tests can be time-consuming and expensive. Since such tests often involve storage studies, providing sufficient suitable storage facilities can also be a problem.

[0007] U. Siripatrawan and P. Jantawat disclose a method for predicting the shelf life of packaged, moisture-sensitive meals using a trained machine learning model (A novel method for shelflife prediction of a packaged moisture-sensitive snack using multilayer perceptron neural network, Expert Systems with Applications, 34 (2008) 1562-1567). Shelf life is primarily influenced by moisture penetration into the packaging. The meals are represented only by generic components such as ash content, fat content, carbohydrate content, and the like; the chemical compounds actually present in the meals and their chemical and / or physical properties are not taken into account. Thus, while the method is suitable for predicting the shelf life of certain packaged, moisture-sensitive meals, a transfer of the results to other formulations is not possible.For example, it is not possible to predict the influence of a formulation on a material in contact with the formulation and / or the influence of a material on a formulation in contact with the material using this method.

[0008] The present invention addresses these and other problems. The present invention provides means by which at least one stability characteristic for a system comprising a formulation and at least one material in contact with the formulation can be predicted using a machine learning model. A first aspect of the present invention is a computer-implemented method for training the machine learning model. The training method comprises the steps:

[0009] Providing and / or receiving information about a plurality of reference formulations, wherein the information for each reference formulation comprises (i) a recipe and (ii) at least one measured stability characteristic, wherein the stability characteristic is a measure of the stability of a system comprising the reference formulation and at least one material in contact with the reference formulation,

[0010] Determination of components of the reference formulation based on the recipe,

[0011] Determination of physical and / or chemical properties of the identified components of each reference formulation,

[0012] Generating a numerical representation for each reference formulation, wherein the numerical representation of the reference formulation comprises for each determined component a numerical representation of the determined component, wherein the numerical representation of each determined component comprises a value for the amount of the component in the reference formulation and values ​​for the determined physical and / or chemical properties of the component,

[0013] Training the machine learning model, wherein the training for each reference formulation comprises the following steps: o Inputting the numerical representation of the reference formulation into the machine learning model, wherein the machine learning model is configured to predict at least one stability feature based on the numerical representation of the reference formulation and model parameters, o Receiving the at least one predicted stability feature from the machine learning model, o Quantifying a deviation between the at least one preferably measured stability feature and the at least one predicted stability feature, o Modifying the model parameters with a view to reducing the deviation,

[0014] Storing and / or outputting the trained machine learning model and / or transmitting the trained machine learning model to a separate computer system and / or using the trained machine learning model to predict at least one stability characteristic of a new system comprising a formulation and at least one material in contact with the formulation.

[0015] Another object of the present invention is a computer-implemented method for determining at least one stability characteristic of a system comprising a test formulation and at least one material in contact with the test formulation using the trained machine learning model. The method comprises the steps:

[0016] Receiving a recipe of the test formulation,

[0017] Determination of components of the test formulation based on the recipe of the test formulation,

[0018] Determination of physical and / or chemical properties of the identified components,

[0019] Generating a numerical representation of the test formulation, wherein the numerical representation of the test formulation comprises, for each determined component, a numerical representation of the determined component, wherein the numerical representation of each determined component comprises a value for the amount of the component in the test formulation and values ​​for the determined physical and / or chemical properties of the component,

[0020] Entering the numerical representation of the test formulation into the trained machine learning model,

[0021] Receiving at least one predicted stability feature from the trained machine learning model,

[0022] Outputting and / or storing the at least one predicted stability feature and / or transmitting the at least one predicted stability feature to a separate computer system.

[0023] Another object of the present invention is a computer system comprising

[0024] • an input unit,

[0025] • a control and computing unit and

[0026] • an output unit, wherein the control and computing unit is configured to provide a trained machine learning model, o wherein the trained machine learning model is configured and has been trained with the aid of training data to predict at least one stability feature for a system comprising a formulation and at least one material in contact with the formulation based on a numerical representation of the formulation, o wherein the training data for each reference formulation of a plurality of reference formulations comprises (i) a numerical representation of the reference formulation and (ii) at least one measured stability feature, wherein the at least one measured stability feature is a measure of the stability of a system comprising the reference formulation and at least one material in contact with the reference formulation,o where training the machine learning model for each reference formulation comprises:,

[0027] ■ Inputting the numerical representation of the reference formulation into the machine learning model,

[0028] ■ Receiving at least one predicted stability feature from the machine learning model,

[0029] ■ Quantifying a deviation between the at least one measured stability characteristic and the at least one predicted stability characteristic,

[0030] ■ Modifying the model parameters to reduce the deviation, causing the input unit to receive a recipe of a test formulation,

[0031] To determine components of the test formulation based on the recipe, to determine physical and / or chemical properties of the determined components, to generate a numerical representation of the test formulation, wherein the numerical representation of the test formulation comprises a numerical representation of the determined component for each determined component, wherein the numerical representation of each determined component comprises a value for the amount of the component in the test formulation and values ​​for the determined physical and / or chemical properties of the component, to input the numerical representation of the test formulation into the trained machine learning model, to receive at least one predicted stability characteristic from the trained machine learning model, to cause the output unit,to output and / or store and / or transmit at least one predicted stability characteristic to a separate computer system.

[0032] Another object of the present invention is a computer-readable storage medium comprising instructions which, when executed by a computer system, cause the computer system to perform the following steps:

[0033] Providing a trained machine learning model, o wherein the trained machine learning model is configured and has been trained using training data to predict at least one stability characteristic for a system comprising a formulation and at least one material in contact with the formulation based on a numerical representation of the formulation, o wherein the training data for each reference formulation of a plurality of reference formulations comprises (i) a numerical representation of the reference formulation and (ii) at least one measured stability characteristic, wherein the at least one measured stability characteristic is a measure of the stability of a system comprising the reference formulation and at least one material in contact with the reference formulation, o wherein training the machine learning model for each reference formulation comprises:

[0034] ■ Inputting the numerical representation of the reference formulation into the machine learning model,

[0035] ■ Receiving at least one predicted stability feature from the machine learning model,

[0036] ■ Quantifying a deviation between the at least one measured stability characteristic and the at least one predicted stability characteristic,

[0037] ■ Modifying the model parameters to reduce the deviation,

[0038] Receiving a recipe of a test formulation,

[0039] Determination of components of the test formulation based on the recipe,

[0040] Determination of physical and / or chemical properties of the identified components,

[0041] Generating a numerical representation of the test formulation, wherein the numerical representation of the test formulation comprises, for each determined component, a numerical representation of the determined component, wherein the numerical representation of each determined component comprises a value for the amount of the component in the test formulation and values ​​for the determined physical and / or chemical properties of the component,

[0042] Entering the numerical representation of the test formulation into the trained machine learning model,

[0043] Receiving at least one predicted stability feature from the trained machine learning model,

[0044] Outputting and / or storing the at least one predicted stability feature and / or transmitting the at least one predicted stability feature to a separate computer system.

[0045] Further objects of the present invention as well as preferred embodiments can be found in the dependent claims, the present description and the drawings.

[0046] The invention is explained in more detail below, without distinguishing between the subject matters of the invention (training method, prediction method, computer system, computer-readable storage medium). Rather, the following explanations are intended to apply analogously to all subject matters of the invention, regardless of the context in which they are described (training method, prediction method, computer system, computer-readable storage medium).

[0047] If steps are mentioned in a particular order in this description or in the claims, this does not necessarily mean that the invention is limited to that order. Rather, the steps may be performed in a different order or even in parallel, unless a step builds on another step, which requires that the subsequent step be performed subsequently (which will become clear in individual cases). The specified order thus represents preferred embodiments of the invention.

[0048] The invention is explained in more detail at some points with reference to drawings. The drawings depict specific embodiments with specific features and combinations of features, which primarily serve for illustrative purposes; the invention should not be understood as being limited to the features and combinations of features shown in the drawings. Furthermore, statements made in the description of the drawings regarding features and combinations of features are intended to apply generally, meaning they are also transferable to other embodiments and are not limited to the embodiments shown.

[0049] The present disclosure describes means for predicting at least one stability characteristic of a system comprising a formulation and at least one material in contact with the formulation using a machine learning model.

[0050] The formulation is preferably a formulation of one or more plant protection products. The term "plant protection product" refers to a product used to protect plants or plant products from pests or to prevent their effects, to destroy undesirable plants or parts of plants, to inhibit undesirable plant growth or to prevent such growth, and / or to influence plant life processes in a way other than nutrients (e.g., growth regulators). Growth regulators are used, for example, to increase lodging in cereals by shortening stalk length (stalk shorteners or, more accurately, intermodia shorteners), to improve the rooting of cuttings, to reduce plant height by stunting in horticulture, or to prevent the germination of potatoes. Other examples of plant protection products are herbicides, fungicides, and pesticides (e.g., insecticides).The formulation may also include one or more nutrients for plants. The term "nutrients" refers to those inorganic and organic compounds from which plants can obtain the elements that make up their bodies. These elements themselves are often also referred to as nutrients. These are usually simple inorganic compounds such as nitrate (NO), phosphate (PO), and potassium (K). + In addition to the core elements of organic matter (C, O, H, N, and P), K, S, Ca, Mg, Mo, Cu, Zn, Fe, B, Mn, CI in higher plants, Co, and Ni are also essential for life. Different compounds can be present for the individual nutrients; for example, nitrogen can be supplied as nitrate, ammonium, or amino acid.

[0051] The at least one material is a material that comes into contact with the formulation during manufacture, packaging, transport, storage, and / or use of the formulation. Preferably, the at least one material is a material that is in contact with the formulation for a period of time greater than one hour, preferably greater than one day, even more preferably greater than one week, most preferably greater than one month. The period may also be greater than one year.

[0052] The at least one material is preferably a material from which a container for a formulation can be produced and / or a material that is a component of a container for a formulation and / or a container itself and / or a material for a lid of a container or a lid of a container and / or a material for a seal of a container and / or a seal for a container.

[0053] Preferably, the at least one material is a plastic or a composite of several plastics. Materials frequently used for contact with a formulation include, for example, polyethylene (PE), polypropylene (PP), polyvinyl chloride (PVC), polyethylene terephthalate (PET), polyamide (PA), ethylene-vinyl alcohol copolymer (EVOH), coextrudate of polyethylene and ethylene-vinyl alcohol copolymer (COEX PE / EVOH), and coextrudate of polyethylene and polyamide (COEX PE / PA).

[0054] The at least one stability characteristic is a measure of the stability of the system comprising the formulation and the at least one material in contact with the formulation. Stability is understood to mean the property that at least one of the components of the system, i.e. the formulation or the at least one material, or both components of the system, when in contact with each other, retain the properties (including their chemical composition) they had prior to contact. High stability can be demonstrated, for example, by the fact that both components of the system retain their properties after contact or change them to an extent that is irrelevant for the intended use of the formulation and / or the at least one material (no loss of quality).

[0055] Low stability may, for example, be manifested in the fact that at least one property of at least one component of the system is altered after contacting in such a way that the altered property has a negative impact on the intended use of the formulation and / or the at least one material.

[0056] If stability is low, the at least one material is typically not suitable as a container, lid, sealant, or the like for the formulation. If stability is low, the at least one material should not be used as a material in contact with the formulation.

[0057] In one embodiment, system stability is understood to mean the resistance of the material to the formulation with which the material is in contact. The formulation can, for example, influence the strength and / or impermeability of the material. A high level of system stability can therefore mean that the material in contact with the formulation maintains its impermeability to the formulation and / or to external moisture and / or to other externally penetrating substances (e.g., oxygen). A high level of system stability can mean that the material in contact with the formulation maintains its strength and does not soften or become brittle and / or retains its shape.

[0058] Examples of further stability features are listed further down in the description.

[0059] The prediction of at least one stability characteristic is carried out using a machine learning model.

[0060] A "machine learning model" can be understood as a computer-implemented data processing architecture. The model can receive input data and produce output data based on this input data and model parameters. Through training, the model can learn a relationship between the input data and the output data. During training, the model parameters can be adjusted to produce a desired output for a given input.

[0061] When training such a model, the model is presented with training data from which it can learn. The trained machine learning model is the result of the training process. The training data includes input data and the correct output data (target data) that the model is supposed to generate based on the input data. During training, patterns are recognized that map the input data to the target data.

[0062] During the training process, the input data of the training data is fed into the model, and the model generates output data. The output data is compared with the target data. Model parameters are adjusted to reduce the deviations between the output data and the target data to a (defined) minimum.

[0063] During training, a loss function can be used to guide the training process and evaluate the model's predictive quality. The loss function can be chosen to assume low values ​​for a desired relationship between output data and target data and / or high values ​​for an undesired relationship between output data and target data. Such a relationship can be, for example, a similarity, a dissimilarity, or another relationship.

[0064] An error function can be used to calculate an error (loss) for a given pair of output and target data. The goal of the training process may be to modify (adjust) the parameters of the machine learning model so that the error is reduced to a (defined, desired) minimum for all pairs in the training dataset.

[0065] For example, an error function can quantify the deviation between the model's output data for a given input and the target data. For example, if the output and target data are numbers, the error function can be the absolute difference between these numbers. In this case, a high absolute value of the error function may indicate that one or more model parameters need to be significantly changed.

[0066] For example, for output data in the form of vectors, difference metrics between vectors such as the mean square error, a cosine distance, a norm of the difference vector such as a Euclidean distance, a Chebyshev distance, an Lp norm of a difference vector, a weighted norm, or any other type of difference metric between two vectors can be chosen as the error function.

[0067] For higher-dimensional outputs, such as two-dimensional, three-dimensional, or higher-dimensional outputs, an element-wise difference metric can be used. Alternatively or additionally, the output data can be transformed, e.g., into a one-dimensional vector, before calculating an error value. In the present case, the machine learning model is trained to predict at least one stability characteristic for a system comprising a formulation and at least one material in contact with the formulation.

[0068] The machine learning model can be trained to predict a single stability feature. To predict multiple stability features, multiple machine learning models can be used, each trained to predict a (single) stability feature.

[0069] However, it is also possible for a machine learning model to be trained to predict more than one stability feature (e.g., two or three or four or five or six or more than six stability features).

[0070] The machine learning model can be trained to predict one or more stability characteristics for a defined material. To predict one or more stability characteristics for multiple materials, multiple machine learning models can be trained independently, each predicting one or more stability characteristics for a defined material. However, it is also possible to train a machine learning model to predict one or more stability characteristics for multiple materials.

[0071] The machine learning model is trained using training data. The training data includes, for a variety of reference formulations, a formulation and at least one preferably measured stability characteristic. In this description, to distinguish between training data and data used for prediction, the suffix "reference" is used for those formulations for which information is used to train the machine learning model. The formulation for which at least one stability characteristic is predicted using the trained machine learning model is also referred to as the "test formulation" in this description. The terms "reference formulation" and "test formulation" therefore serve only to distinguish between the training phase and the prediction phase. The suffix "reference" and "test" have no other (restrictive) meanings.Statements made in this description regarding formulations apply to both test formulations and reference formulations. In some places, the term "test formulation / reference formulation" is used to clarify that the respective statement applies to both test formulations and reference formulations.

[0072] The term “multiplicity of reference formulations” means more than 10, preferably more than 100 reference formulations.

[0073] In addition to a formulation and at least one preferably measured stability characteristic, the training data can also include further data, for example information on the material used (e.g. which material is in contact with the formulation), data on the conditions under which the at least one preferably measured stability characteristic was determined, e.g. measured. The at least one stability characteristic can, for example, be a property of the material and / or the formulation after storage under defined conditions. Further training data can, for example, be the storage time and / or data on the defined conditions (temperature, pressure, intensity of mechanical stress and / or the like). From the formulation and, if applicable,Further data can be used to generate input data for training the machine learning model; data on the at least one stability feature acts as target data for training the machine learning model.

[0074] In the case of the training data, the at least one stability characteristic has preferably been measured, i.e., it is preferably the result of a metrological analysis of each reference formulation and / or the at least one material. Preferably, the at least one stability characteristic has been determined using a standardized testing procedure (e.g., a measurement procedure regulated by a standard (DIN or ISO)). An example of a standard for determining stability characteristics is EN ISO 13274 in its currently applicable version.

[0075] Typically, in such a testing method, the at least one material (e.g. in the form of a test specimen and / or in the form of an intended consumer article (e.g. as a container, lid, seal and / or the like)) is subjected to a first measurement, then brought into contact with the reference formulation for a defined period of time under defined conditions, and then subjected to a second measurement. It is possible that more than two measurements are carried out to determine the at least one stability characteristic. It is conceivable that no measurement is carried out before the at least one material is brought into contact with the reference formulation.It is also conceivable that one or more measurements are carried out on the formulation before and / or after the formulation comes into contact with the material, for example to identify one or more substances in the formulation that have migrated from the material into the formulation.

[0076] It is possible that no measured stability characteristic is available for one or more reference formulations, or that only some of the stability characteristics are measured. It is possible that one or more stability characteristics of the training data are the result of a calculation.

[0077] Examples of stability characteristics are: the swelling of the at least one material in contact with the formulation, e.g. measured using the method BAM-GGR 004, BAM-GGR 015, DIN EN ISO 13724 using an analytical balance (e.g. from Mettler Toledo), the permeation of components of the formulation into and / or through the at least one material, e.g. measured using the method BAM-GGR 015 using an analytical balance (e.g. from Mettler Toledo),

[0078] Tensile impact strength of at least one material after contact with the formulation, e.g. measured using the BAM-GGR 015 method using a device for testing tensile impact hardness using the pendulum impact method (e.g. the HIT5P device from Zwick Roell GmbH),

[0079] Melt flow rate of at least one material after contact with the formulation, e.g. measured using the method BAM-GGR 015, DIN EN ISO 13724 using a melt flow index tester (e.g. the flow tester A4 7ow from Zwick Roell GmbH),

[0080] Residual tensile strength of at least one material after contact with the formulation, e.g. measured using the method BAM-GGR 004, DIN EN ISO 13724 using a device for testing tensile force (e.g. the Universal Testing Device Z010 from Zwick Roell GmbH).

[0081] B AM-GGR004: https: / / tes.bam.de / TES / Content / DE / Downloads / ggr-004

[0082] B AM-GGR015: https: / / tes.bam.de / TES / Content / DE / Downloads / ggr-015

[0083] In addition to at least one stability feature, the training data includes a recipe for each reference formulation.

[0084] A recipe contains information about the components of a formulation. A recipe specifies which components are contained in a formulation or from which components it is composed.

[0085] In the case of a plant protection product formulation, common components are: one or more active ingredients (see e.g. https: / / www.proplanta.de / Pflanzenschutzmittel / Wirkstoffe / ), natural oils (e.g. sunflower oil optionally dewaxed, mineral oils, silicone oils), fatty acids and fatty acid derivatives (e.g.), aliphatic solvents (e.g. butyl lactic acid, cyclohexanone, dimethyl sulfoxide, ethanol), aromatic solvents (e.g. toluene, Solvesso 100, Solvesso 150, Solvesso 200, Solvarex, Caromax 28 LN, Caromax 20 LN, xylene, xylene), surfactants (e.g. ATLOX 3467, Enviomet EM 5665, Enviomet WT 8519, Rhodacal 60-BE, Genapol X-060, Genapol XM 060), Waxes (e.g. Joncryl Wax 4), other additives (e.g. dyes such as Laicril P-1530).

[0086] The recipe also specifies the quantities of the ingredients. These quantities can be specified, for example, as relative quantities (e.g., weight percent, mole percent, volume percent, concentration) or as absolute quantities (e.g., mass, mole, volume).

[0087] The recipe may include information on how the formulation is prepared, ie it may, for example, include information on the order in which the ingredients are combined and / or the conditions under which they are combined (e.g., temperature, pressure, stirring speed), how they are combined (e.g., dropwise, with stirring, and / or the like), and / or the further steps taken to produce the formulations (e.g., filtration, sedimentation, distillation, heating, cooling, and / or the like).

[0088] The recipe preferably includes information on the type of (reference) formulation. Examples of types of (reference) formulations are: emulsion concentrate formulations (EC formulations), suspension concentrate formulations (SC formulations), suspoemulsions (SE formulations), oil dispersions (OD formulations), emulsions for seed treatment (ES formulations), oil-in-water emulsions (EW formulations), suspension concentrates for seed treatment (FS formulations), technical active ingredients (TC formulations), and mixed formulations of CS and SC (ZC formulations).

[0089] Further examples can be found at https: / / www.bvl.bund.de / EN / Tasks / 04_Plant_protection_products / 01_ppp_tasks / 08_ProductChemistr y / 01_ppp_coformulants_formulationChemistry / 01_ppp_formulation_types / ppp_formulation_types_no de .html.

[0090] In a first step of the training and prediction procedures, components of the test formulation / reference formulations are determined based on the recipe. Typically, the components are specified in the form of a name (e.g., IUPAC name, common name, or brand name), a code (e.g., Chemical Abstract Service Registry Number (CAS No.)), and / or a formula (e.g., molecular formula, chemical structural formula).

[0091] Preferably, at least the main constituents of the respective formulation(s) are determined. A main constituent can, for example, be a constituent that is present in the formulation in a minimum amount (e.g., at least 1%, or at least 5%, or at least 10%, where the percentage can be, for example, weight percent, volume percent, or mole percent).

[0092] Some components of a recipe do not consist of a single chemical compound, but rather a mixture of different chemical compounds (e.g. sunflower oil). For components of a formulation that consist of more than one chemical compound, the chemical compounds contained in the component are preferably determined. The term “chemical compound” is understood to mean a pure substance that consists of atoms of two or more chemical elements, where - in contrast to mixtures - the atom types are in a fixed stoichiometric ratio to one another. Preferably, at least those chemical compounds of a component are determined that are present in the component in a minimum amount (e.g. at least 1%, or at least 5% or at least 10%, whereby the percentage can be, for example, percent by weight, percent by volume or mole percent).Preferably, all chemical compounds contained in each ingredient are identified. Preferably, all ingredients of the respective formulation(s) are identified. Preferably, the chemical compound(s) involved are identified for each ingredient. In other words, when reference is made to a "ingredient" in this description, this preferably refers to a chemical compound (and not a generic ingredient).

[0093] For many components containing more than one chemical compound, the chemical compounds contained can be determined from publicly accessible databases and / or from the literature. Component manufacturers often provide information about the chemical compounds contained, e.g., in a product specification. It is also possible to analyze components for the chemical compounds they contain (or have them analyzed).

[0094] Some ingredients may contain impurities. Impurities are preferably determined for ingredients. Impurities are preferably determined for at least the main ingredients of the respective formulation(s). Preferably, at least those impurities that occur in a minimum amount in a (main) ingredient are determined. For many ingredients, the impurities present can be determined from publicly accessible databases and / or from the literature. Manufacturers of ingredients often provide information on the impurities present, e.g. in the form of a product specification. It is also possible to analyze ingredients for the impurities they contain or to have them analyzed. Any impurities present are preferably taken into account in the form of the respective chemical compound(s).

[0095] Depending on the formulation, the constituents identified may be more than one hundred, particularly when the chemical compounds constituting the constituents and / or impurities are identified.

[0096] Preferably, molecular descriptors are determined for preferably all determined components (chemical compounds) of the respective formulation(s), in particular for those components (chemical compounds) for which no measured values ​​of physical and / or chemical properties are available.

[0097] The term “molecular descriptor” is understood to mean a representation of a chemical compound that reflects the chemical structure of the chemical compound at the molecular level or from which the chemical structure at the molecular level can be derived.

[0098] Examples of molecular descriptors are IUPAC name (IUPAC: International Union of Pure and Applied Chemistry), SMILES code (SMILES: Simplified Molecular Input Line Entry Specification), InChl code, CML code (CML: Chemical Markup Language), WLN code (WLN: Wiswesser Line Notation), molecular graph.

[0099] Molecular descriptors can be obtained, for example, from publicly accessible databases and / or literature based on the name of a component or a component code. Component manufacturers often also provide molecular descriptors of chemical compounds, e.g., in the form of a product specification. Molecular descriptors can be generated manually using suitable programs (e.g., ChemDraw, ChemSketch). Molecular descriptors can usually be converted into one another, e.g., a SMILES code can be generated from an lUPAC name, and vice versa.

[0100] Preferably, molecular descriptors are determined that can be used in computer programs to determine physical and / or chemical properties of chemical compounds based on the structure of the chemical compounds. These are typically SMILES codes and / or molecular graphs. More on computer programs for determining physical and / or chemical properties is provided further down in the description. Preferably, molecular descriptors are determined at least for the main components of a formulation and / or for chemical compounds and / or impurities that are present in at least a minimum amount in a (main component). Preferably, molecular descriptors are determined for all chemical compounds of a formulation, including impurities.

[0101] Physical and / or chemical properties are determined for components of the test formulation / reference formulation. Examples of physical and / or chemical properties are:

[0102] molecular weight,

[0103] Partition coefficient, e.g. for w-octanol / water or a derived value such as the decimal logarithm of the partition coefficient logD e.g. at pH=7 and / or pH=2.3 and / or the logP value

[0104] Melting point (e.g. under normal conditions),

[0105] Water solubility, e.g. in the form of solubility as mass of substance of the compound that can be dissolved in 100 g of water,

[0106] Proportion of polar / apolar functional groups of one or more contained molecules pKa / pKb values ​​of one or more contained molecules,

[0107] Charges in the molecule at corresponding pH value, pH value of the formulation.

[0108] The physical and / or chemical properties can be measured values. The physical and / or chemical properties can be retrieved from one or more databases and / or determined from the literature. Examples of databases that store physical and / or chemical properties are PubChem (https: / / pubchem.ncbi.nlm.nih.gov / ), NIST Chemistry WebBook (https: / / webbook.nist.gov / chemistry / ), and CRC Handbook of Chemistry and Physics (https: / / hbcp.chemnetbase.com / ). The physical and / or chemical properties can be determined experimentally.

[0109] The physical and / or chemical properties can be calculated values. For example, the physical and / or chemical properties can be calculated based on the molecular structure of a component. There are numerous methods for calculating physical and / or chemical properties based on molecular structure. These can be found in the literature under the term QSPR (Quantitative Structure Property Relationship). There are also commercially and freely available computer programs for calculating the physical and / or chemical properties of substances based on their molecular structure (e.g., QSAR-Co: J. Chem. Inf. Model. 2019, 59, 6, 2538-2544; Molgen-QSPR: https: / / www.researchgate.net / publication / 266470632).

[0110] A method for calculating the decimal logarithm of the partition coefficient for n-octanol / water is described, for example, in: AJ Leo: Calculating logP oct from structures, Chem. Rev. 1993, 93, 4, 1281-1306.

[0111] A method for calculating the melting point is described, for example, in: H. Modarresi et al.: QSPR Correlation of Melting Point for Drug Compounds Based on Different Sources of Molecular Descriptors, J. Chem. Inf. Model. 2006, 46, 2, 930-936.

[0112] A method for calculating water solubility is described, for example, in: N. Meftahi et al. : Predicting aqueous solubility by QSPR modeling, Journal of Molecular Graphics and Modelling, Volume 106, July 2021, 107901. Further methods for calculating physical and / or chemical properties can be found in: AR Katritzky et al.: QSPR as a means of predicting and understanding chemical and physical properties in terms of structure, Pure and Applied Chemistry 69(2):245-248).

[0113] Preferably, for all components (chemical compounds) of a formulation for which a molecular descriptor has been determined, physical and / or chemical properties of the components are calculated on the basis of the respective molecular descriptor.

[0114] In a further step, a numerical representation is generated for each formulation.

[0115] Many machine learning models require numerical values ​​as input data. The input data representing a formulation includes a numerical representation of each component (chemical compound) identified in the formulation. It is possible for each component (chemical compound) in the formulation to have a numerical representation. It is possible for only major components of the formulation to have a numerical representation.

[0116] A numerical representation of a component (chemical compound) comprises the amount of the component (chemical compound) in the formulation and values ​​of the determined physical and / or chemical properties of the component (chemical compound). Preferably, the numerical representation consists exclusively of the amount of the component (chemical compound) in the formulation and values ​​of the determined physical and / or chemical properties of the component (chemical compound). Preferably, the numerical representation of a component (chemical compound) does not comprise any categorical variables. The amount of the component (chemical compound) and the values ​​of the determined physical and / or chemical properties of the component (chemical compound) can, for example, be in the form of a vector.

[0117] In other words, in a preferred embodiment, the prediction of the at least one stability characteristic is based on numerical representations of the components (chemical compounds) of a formulation, which only include information about the amount of the components (chemical compounds) and their physical and / or chemical properties. Further information, such as information about the nature of the components (which can be encoded with categorical variables) or about the chemical structures of the chemical compounds of the components, is not required. This further information is implicitly contained in the physical and / or chemical properties. In a preferred embodiment, all physical and / or chemical properties are calculated based on the molecular descriptors (which represent the chemical structure).It is therefore possible to represent a formulation solely by the physical and / or chemical properties of its components and the amounts of the components. Such a representation is sufficient to predict a stability characteristic for a system comprising the formulation and a material in contact with the formulation.

[0118] A numerical representation of a formulation can consist exclusively of the numerical representations of its (main) components. For example, it can be a matrix in which each row or column is a numerical representation of a component.

[0119] Preferably, the numerical representation of each formulation comprises, in addition to numerical representations of components of the representation, only one or more categorical variables that specify the type of formulation. For example, an SC formulation can be represented by a 1, an SE formulation by a 2, and other formulations by other digits. The type of formulation can also be represented by a 1-of-n code (one-hot encoding). When training the machine learning model, input data representing a reference formulation is fed into the machine learning model.The machine learning model can be configured to independently subject the numerical representations of the components to feature extraction, in which an optionally compressed numerical representation of the component is generated from each numerical representation of the component. Each numerical representation of a component can, for example, be fed to an artificial neural network configured to generate an optionally compressed numerical representation based on the supplied numerical representation and on the basis of trainable (learnable) model parameters. The artificial neural network can, for example, be a convolutional neural network (CNN).

[0120] The numerical representation of a formulation in feature space contains information relevant for predicting at least one stability feature. However, it may also contain other, irrelevant, or redundant information. A "compressed numerical representation" is a representation that contains or emphasizes only the relevant information, or in which irrelevant and / or redundant information is suppressed and / or reduced.

[0121] A compressed numerical representation can be generated, for example, by an encoder of an autoencoder.

[0122] An autoencoder is typically used to learn efficient data encodings in an unsupervised learning process. In general, the goal of an autoencoder is to learn a representation for a dataset, typically for dimensionality reduction, by training the machine learning model to identify the relevant data and ignore noise.

[0123] The encoder compresses the input data into a latent-space representation, and the decoder attempts to reconstruct the input data from this latent-space representation. The latent-space representation is the compressed numerical representation of the input data.

[0124] A key feature of an autoencoder is an information bottleneck between the encoder and decoder. This bottleneck, a continuous vector of fixed length, causes the machine learning model to learn a compressed representation that captures the most statistically relevant information in the data.

[0125] Further details on autoencoders and the generation of compressed representations can be found in several publications (see, for example, D. Bank et al. '. Autoencoders, arXiv:2003.05991).

[0126] The numerical representations of formulations can be of different sizes (i.e., have different numbers of dimensions) for different formulations, simply because they can have a different number of components.

[0127] Therefore, in a first step, an optionally compressed numerical representation of the component can be generated for each identified component of a formulation based on the numerical representation of the component. In a second step, the optionally compressed numerical representations of all identified components of a formulation can be pooled into a numerical representation of the formulation. The result can be a representation (a vector) that has the same size (the same dimension) for all formulations.

[0128] It is possible to use an attention mechanism to aggregate numerical representations of components. In machine learning, attention is a technique that mimics cognitive attention in humans. The effect is to enhance some parts of the input data and attenuate others—the idea being that the machine learning model should pay more attention to an important part of the data than to a less important one. The model learns which parts are important and which are less important during the training phase.

[0129] Although model architectures based on convolutional networks can describe arbitrary functions on mathematical sets (see, for example, M. Zaheer et al., 'Deep Sets, arXiv: 1703.06114v3), such architectures can only implicitly model individual interactions between specific formulation components from the same recipe. Neural networks with an attention mechanism, on the other hand, are able to explicitly model interactions between formulation components of the same recipe. Attention functions, in general terms, combine a query vector with pairs of (key, value) vectors to form an output vector. The output vector, in turn, is calculated from the weighted sum of the value vectors, with the weights being predicted by the network architecture based on the query and 'key' vectors.By providing physical / chemical properties of formulation ingredients from a recipe as query, key, and value vectors, interactions between formulation ingredients of the same recipe can be directly processed. It has been found that the attention mechanism can improve the prediction accuracy of the trained models compared to convolutional networks.

[0130] It is possible to use a transformer network to generate a numerical representation of a formulation. Such a transformer network can account for pairwise interactions between the elements of the input. The transformer network can process sequential inputs, differently weighting the contributions of each piece of input data. By using multiple self-attention units, the transformer network is able to process relevant interactions between two or more components of the formulation. Details on transformer networks can be found, for example, in: J. Lee et al.: A Framework for Attention-based Permutation-Invariant Neural Networks, 2019, arXiv: 1810.00825v3; A. Vaswani et al.: Attention is all you need, arXiv: 1706.03762v5.

[0131] Further details on attention mechanisms are described in scientific articles and patents / patent applications (see e.g.: A. Vaswani et al:. Attention Is All You Need, arXiv: 1706.03762v5, EP3166049A1, US20170262996, WO2020 / 222985).

[0132] Finally, the recipe of a formulation can be encoded in a single vector using a global pooling operation (see, for example, J. Lee et al.: Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks, arXiv: 1810.00825v3). An example of a global pooling operation is the summation of the representations (or embeddings pre-processed with the attention mechanism) of the individual components. Attention pooling is a more flexible extension of the summation. Inspired by the attention mechanism, a single learnable vector is provided as a query vector, with which a single output vector can be generated from the representations (or embeddings) of the formulation components as key and value vectors. Here, too, representations of formulation components are processed in a weighted sum, but the weights from the representations (or embeddings) areEmbeddings) of the formulation components.

[0133] The bundling operation creates a permutation-invariant representation of a formulation. This means that the order in which components are listed is irrelevant. In other words, for a formulation with components A, B, and C, the same representation is always generated, regardless of whether the components of the formulation are listed in the order A, B, C or A, C, B or B, A, C or B, C, A or C, A, B or C, B, A. The (permutation-invariant) representation of the formulation can be fed into a regression model that is configured to calculate at least one stability feature based on the (permutation-invariant) representation of the formulation and on the basis of trainable (learnable) model parameters.

[0134] The regression model can be, for example, an artificial neural network or a random forest model or the like. Numerous examples of regression models can be found in the literature (see, for example, H. Saleh, JA Layous: Machine Learning - Regression, 2022, DOI: 10.13140 / RG.2.2.35768.67842). In one embodiment, the regression model is an artificial neural network that is trained together with the upstream networks to pool the representations of the formulation components. In terms of architecture, such a regression network can be a "fully connected network" (also referred to as a "multi-layer perceptron"). In this architecture, individual neurons (with one input and one output) are organized into layers, so that the total output of one layer represents the input of every neuron in the following layer.

[0135] The at least one predicted stability characteristic can be compared with the at least one (preferably measured) stability characteristic of the training data. The deviations between the at least one predicted stability characteristic and the stability characteristic of the training data can be quantified using an error function. Examples of error functions are root mean square deviation (RMSD), mean absolute deviation (MAD), log cosine distances, cosine distances, Huber loss, etc.

[0136] The model parameters can be modified, for example, in a gradient method or in another optimization method to reduce the deviations to a (defined) minimum.

[0137] The training is continued with further training sets, each comprising a numerical representation of a reference formulation and at least one (preferably measured) stability feature.

[0138] Once the machine learning model achieves a desired level of prediction accuracy, training can be terminated. The trained machine learning model can be stored in a data store and / or transmitted to a separate computer system.

[0139] To validate the model, the training data can be split into two datasets. The first dataset can be used for training; the second dataset can be used for validation.

[0140] Once the machine learning model is trained, it can be used to predict at least one stability characteristic of a new system comprising a new formulation (test formulation) and at least one material in contact with the new formulation (test formulation). The term "new" means that a dataset describing the formulation has not already been used during training.

[0141] The prediction process is analogous to the training process. This means that the steps performed for the reference formulations are also performed for the new formulation (test formulation):

[0142] In a first step, a recipe for the test formulation is received.

[0143] In a second step, components of the test formulation are determined based on the recipe.

[0144] In a third step, physical and / or chemical properties of the components of the test formulation are determined.

[0145] In a fourth step, a numerical representation of the test formulation is generated. The numerical representation includes a numerical representation of each identified component. Each numerical representation of a identified component includes a value for the amount of the component in the test formulation and values ​​for the identified physical and / or chemical properties of the component.

[0146] In a fifth step, the numerical representation of the test formulation is input into the trained machine learning model.

[0147] In a sixth step, at least one predicted stability feature is received from the trained machine learning model.

[0148] In a seventh step, the at least one predicted stability characteristic is output (e.g. displayed on a screen and / or printed on a printer) and / or stored in a data storage device and / or transmitted to a separate computer system.

[0149] In a further step, the prediction can be verified experimentally. It is possible to use a measurement to check whether at least one predicted stability characteristic was correctly predicted.

[0150] It is possible that not every predicted stability criterion is tested experimentally. It is possible that (initially) it is checked whether the at least one predicted stability characteristic meets predefined requirements. It is possible that the at least one predicted stability characteristic is compared with a reference value or with several reference values. Such a reference value can, for example, be a minimum value that must be reached for the at least one material for which the prediction was made to even be considered as a material for packaging, a container, a seal and / or the like for the test formulation. If a predicted stability characteristic is, for example, smaller than the minimum value, it is possible that the at least one material cannot be considered as a packaging material, container material, seal material and / or the like for the test formulation.Another material must be found for the test formulation. If a predicted stability characteristic is greater than or equal to the minimum value, the material is a candidate to be considered. It is conceivable that the predicted stability characteristic is verified by experimental testing in a laboratory. In other words: if the at least one predicted stability characteristic meets predefined requirements, testing of the at least one material in a laboratory can be initiated, for example by automatically generating and transmitting a message to the laboratory and / or automatically ordering one or more components of the test formulation and / or automatically ordering the test formulation and / or automatically ordering the at least one material.

[0151] The invention is explained in more detail below with reference to drawings, without wishing to limit the invention to the features and combination of features shown in the drawings.

[0152] Fig. 1 shows schematically and by way of example the generation of a numerical representation of a test formulation / reference formulation based on a recipe.

[0153] The starting point is a recipe R of the respective formulation. In a first step (110), components of the formulation are determined based on the recipe R, and a first list LI of the determined components is generated. In a second step (120), chemical compounds of the components or impurities are determined for components in the first list LI that comprise more than one chemical compound, and / or for components that comprise impurities. The information on chemical compounds and / or impurities contained in components can be obtained from one or more databases DB1. The determined chemical compounds and / or impurities are included in the first list LI of components, whereby a second list L2 of components is obtained.

[0154] In a third step (130), physical and / or chemical properties are calculated for one or more components of the formulation based on one or more molecular descriptors of the component(s). The calculation is performed using one or more QSPR models that relate structural properties to physical and / or chemical properties. In addition to or instead of calculating physical and / or chemical properties, in a fourth step (140), physical and / or chemical properties are read from one or more DB2 databases for one or more components of the formulation.

[0155] In a fifth step (150), a numerical representation RF of the formulation is generated. The numerical representation RF includes, for each determined component of the second list L2, a value for the quantity of the component in the formulation and values ​​for the determined (calculated and / or read) physical and / or chemical properties.

[0156] Fig. 2 schematically shows an exemplary embodiment for training the machine learning model. Training is carried out using training data. The training data is obtained based on information on a large number of reference formulations and preferably measured stability characteristics. In the example shown in Fig. 2, only one training data set TD is shown, which has a numerical representation R Fa reference formulation and a stability characteristic SM. The numerical representation RF of the reference formulation can be generated as described with reference to Fig. 1. The numerical representation RF comprises the numerical representations Rn, R|2 and Rn of three components of the reference formulation. However, the number of components is not limited to three. Each reference formulation comprises at least one component, preferably at least ten components and in some cases even more than 50 components (the number of components is unlimited). The stability characteristic SM is a measure of the stability of a system comprising the reference formulation and at least one material in contact with the reference formulation. The stability characteristic SM is preferably the result of a metrological investigation.

[0157] The numerical representation R FThe reference formulation is fed as input data to the MLM machine learning model.

[0158] In the architecture of the machine learning model (convolutional network) shown in Fig. 2, the numerical representations Rn, R12, and R13 of the components of the reference formulation are independently subjected to a feature extraction (FE). During the feature extraction (FE), a numerical representation R is generated from the numerical representation Rn. c n, which can be compressed (ie, comprises fewer dimensions than the input representation) but, in a preferred implementation, is expanded (ie, comprises more dimensions than the input representation). Likewise, the numerical representation R12 is transformed into a numerical representation R c [2 and from the numerical representation Rß a numerical representation R CB is generated. In Fig. 2, the feature extraction unit (FE) is shown three times to process the three numerical representations of the three components of the reference representation. However, this does not mean that the machine learning model has one feature extraction unit for each component of a formulation. This would be impractical, since different formulations typically have a different number of components. It is therefore possible to have only one feature extraction unit, which processes the numerical representations of the components of a formulation one after the other.

[0159] The (optionally compressed) numerical representations R c n, R C I2 and R C B are combined with a bundling operation GP to form a preferably permutation-invariant representation R PIthe reference formulation. This bundling operation can, in a simple implementation of the architecture, be a conventional reduction operation applied to the processed representations R c n, R c [2 and R C B (e.g., the sum, the mean, the element-wise minimum, the element-wise maximum, etc.). In a preferred implementation of the described architecture, the global pooling operation consists of multi-attention pooling. The global pooling operation is not limited to the options mentioned and can be replaced by other global pooling operations.

[0160] The preferably permutation-invariant representation R PI The reference formulation is fed into a regression model RM, which is based on the preferably permutation-invariant representation R PIA stability feature Sp is determined (predicted). The predicted stability feature Sp can be compared with the stability feature SM of the training dataset. An error function LF is used to quantify the deviations between the predicted stability feature Sp and the stability feature SM of the training dataset. The determined error L can be used to modify model parameters with a view to reducing the error L. Model parameters can be contained in the feature extraction unit FE, in the pooling operation, and in the regression model.

[0161] Fig. 3 shows a schematic and exemplary use of the trained machine learning model for prediction. In this example, a numerical representation R F * a test formulation in the trained model MLM T of machine learning. The numerical representation R F* of the test formulation can be obtained from a recipe R* of the test formulation as described in relation to Fig. 1. The numerical representation R F * The test formulation comprises, as in the example in Fig. 2, three numerical representations Rn*, Rn* and Rn* of three components of the test formulation. From these numerical representations Rn*, Rn* and Rn*, analogous to the example in Fig. 2, three optionally compressed numerical representations R c n*, R c i2* and R c i3* generated by feature extraction FE. The optionally compressed numerical representations R c n*, R c i2* and R c if are converted by the bundling operation GP into a preferably permutation-invariant representation R PI * the test formulation. The preferably permutation-invariant representation R PI* the test formulation is fed into the regression model RM, which determines (predicts) a stability characteristic Sp* for the test formulation.

[0162] Fig. 4 shows an example and schematically a computer system according to the invention.

[0163] A "computer system" is an electronic data processing system that processes data using programmable computing instructions. Such a system typically includes a "computer," the unit containing a processor for performing logical operations, and peripherals.

[0164] In computer technology, "peripherals" refers to all devices connected to a computer that serve to control the computer and / or act as input and output devices. Examples include monitors, printers, scanners, mice, keyboards, drives, cameras, microphones, speakers, etc. Internal connectors and expansion cards are also considered peripherals in computer technology.

[0165] The computer system (1) shown in Fig. 4 comprises an input unit (10), a control and computing unit (20) and an output unit (30).

[0166] The control and computing unit (20) serves to control the computer system (1), to coordinate the data flows between the units of the computer system (1) and to carry out calculations.

[0167] The control and computing unit (20) is configured to provide the trained machine learning model, to cause the input unit (10) to receive a recipe of a test formulation,

[0168] to determine components of the test formulation based on the recipe, to determine physical and / or chemical properties of the determined components, to generate a numerical representation of the test formulation, wherein the numerical representation of the test formulation comprises a numerical representation of the determined component for each determined component, wherein the numerical representation of each determined component comprises a value for the amount of the component in the test formulation and values ​​for the determined physical and / or chemical properties of the component, to input the numerical representation of the test formulation into the trained machine learning model, to receive at least one predicted stability feature from the trained machine learning model, to cause the output unit (30),to output and / or store and / or transmit the at least one predicted stability characteristic to a separate computer system.

[0169] Fig. 5 shows an exemplary and schematic representation of a further embodiment of the computer system according to the invention.

[0170] The computer system (1) comprises a processing unit (21) connected to a memory (22). The processing unit (21) and the memory (22) form a control and calculation unit, as shown in Fig. 4.

[0171] The processing unit (21) may comprise one or more processors alone or in combination with one or more memories. The processing unit (21) may be conventional computer hardware capable of processing information such as digital images, computer programs, and / or other digital information. The processing unit (21) typically consists of an arrangement of electronic circuits, some of which may be embodied as an integrated circuit or as multiple interconnected integrated circuits (an integrated circuit is sometimes referred to as a "chip"). The processing unit (21) may be configured to execute computer programs that may be stored in a main memory of the processing unit (21) or in the memory (22) of the same or another computer system.

[0172] The memory (22) may be ordinary computer hardware capable of storing information, data, computer programs, and / or other digital information either temporarily and / or permanently. The memory (22) may comprise volatile and / or non-volatile memory and may be permanently installed or removable. Examples of suitable memories include RAM (Random Access Memory), ROM (Read-Only Memory), a hard disk, flash memory, a removable computer diskette, an optical disc, magnetic tape, or a combination of the above. Optical discs may include read-only compact discs (CD-ROM), read / write compact discs (CD-R / W), DVDs, Blu-ray discs, and the like.

[0173] In addition to the memory (22), the processing unit (21) can also be connected to one or more interfaces (11, 12, 31, 32, 33) for displaying, transmitting, and / or receiving information. The interfaces can comprise one or more communication interfaces (32, 33) and / or one or more user interfaces (11, 12, 31). The one or more communication interfaces can be configured to send and / or receive information, e.g., to and / or from other computer systems, networks, data storage devices, or the like. The one or more communication interfaces can be configured to transmit and / or receive information via physical (wired) and / or wireless communication connections. The one or more communication interfaces can include one or more interfaces for connecting to a network, e.g.,using technologies such as cellular, Wi-Fi, satellite, cable, DSL, fiber optic, and / or the like. In some examples, the one or more communication interfaces may include one or more short-range communication interfaces configured to connect devices using short-range communication technologies such as NFC, RFID, Bluetooth, Bluetooth LE, ZigBee, infrared (e.g., IrDA), or the like.

[0174] The user interfaces may comprise a display (31). A display (31) may be configured to display information to a user. Suitable examples include a liquid crystal display (LCD), a light-emitting diode display (LED), a plasma display (PDP), or the like. The user input interface(s) (11, 12) may be wired or wireless and may be configured to receive information from a user into the computer system (1), e.g., for processing, storage, and / or display. Suitable examples of user input interfaces include a microphone, an image or video recording device (e.g., a camera), a keyboard or keypad, a joystick, a touch-sensitive surface (separate from or integrated with a touchscreen), or the like.In some examples, the user interfaces may include automatic identification and data capture (AIDC) technology for machine-readable information. This may include barcodes, radio frequency identification (RFID), magnetic stripes, optical character recognition (OCR), integrated circuit cards (ICC), and the like. The user interfaces may further include one or more interfaces for communicating with peripheral devices such as printers and the like.

[0175] One or more computer programs (40) can be stored in the memory (22) and executed by the processing unit (21), which is thereby programmed to perform the functions described in this description. The retrieval, loading, and execution of instructions of the computer program (40) can occur sequentially, such that one instruction is retrieved, loaded, and executed at a time. However, the retrieval, loading, and / or execution can also occur in parallel.

[0176] The machine learning model according to the invention can also be stored in the memory (22).

[0177] The computer system according to the invention can be designed as a laptop, notebook, netbook and / or tablet PC.

[0178] Fig. 6 schematically shows, in the form of a flowchart, an embodiment of the method for training the machine learning model. The training method (200) comprises the following steps:

[0179] (210) Providing and / or receiving information about a plurality of reference formulations, wherein the information for each reference formulation comprises a recipe and at least one preferably measured stability characteristic, wherein the stability characteristic is a measure of the stability of the system comprising the reference formulation and at least one material in contact with the reference formulation,

[0180] (220) Determination of components of the reference formulation based on the recipe,

[0181] (230) Determination of physical and / or chemical properties of the identified components of each reference formulation,

[0182] (240) generating a numerical representation for each reference formulation, wherein the numerical representation of the reference formulation comprises for each determined component a numerical representation of the determined component, wherein the numerical representation of each determined component comprises a value for the amount of the component in the reference formulation and values ​​for the determined physical and / or chemical properties of the component,

[0183] (250) Training the machine learning model, wherein the training for each reference formulation comprises the following steps: (251) Inputting the numerical representation of the reference formulation into a machine learning model, wherein the machine learning model is configured to predict at least one stability feature based on the numerical representation of the reference formulation and model parameters,

[0184] (252) receiving the at least one predicted stability feature from the machine learning model,

[0185] (253) quantifying a deviation between the at least one preferably measured stability characteristic and the at least one predicted stability characteristic,

[0186] (254) Modifying the model parameters to reduce the deviation,

[0187] (260) storing and / or outputting the trained machine learning model and / or transmitting the trained machine learning model to a separate computer system and / or using the trained machine learning model to predict at least one stability characteristic of a new system comprising a formulation and at least one material in contact with the formulation.

[0188] Fig. 7 schematically shows, in the form of a flowchart, an embodiment of the method for predicting at least one stability characteristic of a system comprising a test formulation and at least one material in contact with the test formulation using the trained machine learning model. The prediction method (300) comprises the steps:

[0189] (310) Receiving a recipe of the test formulation,

[0190] (320) Determination of components of the test formulation based on the recipe,

[0191] (330) Determination of physical and / or chemical properties of the identified components,

[0192] (340) generating a numerical representation of the test formulation, wherein the numerical representation of the test formulation comprises, for each determined component, a numerical representation of the determined component, wherein the numerical representation of each determined component comprises a value for the amount of the component in the test formulation and values ​​for the determined physical and / or chemical properties of the component,

[0193] (350) Inputting the numerical representation of the test formulation to the trained machine learning model,

[0194] (360) receiving at least one predicted stability feature from the trained machine learning model,

[0195] (370) Outputting and / or storing the at least one predicted stability feature and / or transmitting the at least one predicted stability feature to a separate computer system.

[0196] Example

[0197] An example for implementing the present invention is described below. An artificial neural network was used as the machine learning model. It processes physical / chemical representations of individual formulation components (based on logD values ​​measured at neutral and acidic pH, water solubility, melting point, and molecular mass in combination with information on the formulation type) in an attention mechanism to create material-specific embeddings, pools these using attention pooling, and then predicts the desired stability characteristics using a fully connected neural network. The machine learning model was trained based on recipes of various crop protection product formulations. The machine learning model was trained to predict the following stability characteristics based on a numerical representation of the respective recipe.

[0198] Fig. 8 shows a comparison of measured and predicted stability characteristics. A total of seven different stability characteristics (residual tensile strength, mass change, permeability, impact tensile strength, and standard deviation of the impact tensile strength measurement) were predicted for three different packaging materials (GF 4750: polyethylene, PEEVOH: COEX PE / EVOH, and PEPA: COEX PE / PA).

[0199] The actual measurements and predictions are close to the diagonals shown and thus show good agreement.

[0200] Based on these predicted stability characteristics, the suitability of the materials as packaging materials for the respective formulation can now be assessed. For example, the tensile strength should be 100%. However, for many materials, this stability characteristic is below 50%. Such low values ​​should be avoided.

[0201] For certain stability characteristics (e.g., tensile impact strength), not only the absolute value but also the range of the measurements (“standard deviation”) is important for making a statement about suitability. A prediction of the range of variation is also possible using the machine learning model.

[0202] Based on the corresponding predicted stability characteristics, tests can be prioritized to verify the predictions. For this purpose, a threshold or criterion for evaluating a predicted stability characteristic can be defined based on experience and / or the stability characteristics of known materials. Measurements can be performed with materials for which a predicted stability criterion lies above the defined threshold and / or measurements can be performed with those materials for which the best prediction results were achieved. Materials for which a predicted stability characteristic lies below the defined threshold do not need to be subjected to practical storage and / or stability tests, as they are obviously not suitable as containers for the respective formulation.

Claims

Patent claims 1. Computer-implemented method comprising: Providing a trained model (MLM T ) of machine learning, where the trained model (MLM T ) of machine learning and has been trained using training data, at least one stability feature (Sp, Sp*) for a system comprising a formulation and at least one material in contact with the formulation on the basis of a numerical representation (R F , R F *) of the formulation, o where the training data for each reference formulation of a plurality of reference formulations (i) a numerical representation (R F) of the reference formulation and (ii) at least one measured stability characteristic (SM), wherein the at least one measured stability characteristic (SM) is a measure of the stability of a system comprising the reference formulation and at least one material in contact with the reference formulation, o wherein training the machine learning model (MLM) for each reference formulation comprises: ■ Entering the numerical representation (R F ) the reference formulation into the machine learning model (MLM), ■ Receiving at least one predicted stability feature (Sp) from the machine learning model (MLM), ■ Quantifying a deviation between the at least one measured stability characteristic (SM) and the at least one predicted stability characteristic (Sp), ■ Reducing the deviation by modifying the model parameters, Receiving a recipe (R*) of a test formulation, Determination of components of the test formulation based on the recipe (R*), Determination of physical and / or chemical properties of the identified components, Creating a numerical representation (R F *) of the test formulation, where the numerical representation (R F *) the test formulation comprises for each identified component a numerical representation (Rn*, Rn*, R13*) of the identified component, wherein the numerical representation (Rn*, R12*, R13*) of each identified component comprises a value for the amount of the component in the test formulation and values ​​for the identified physical and / or chemical properties of the component, Entering the numerical representation (R F *) the test formulation into the trained model (MLM T ) of machine learning, Receiving at least one predicted stability feature (Sp*) from the trained model (MLM T ) of machine learning, Outputting and / or storing the at least one predicted stability feature (Sp*) and / or transmitting the at least one predicted stability feature (Sp*) to a separate computer system.

2. The method according to claim 1, wherein training the machine learning model (MLM) comprises: Providing and / or receiving information about the plurality of reference formulations, wherein the information for each reference formulation comprises (i) a recipe (R) and (ii) the at least one measured stability characteristic (SM), Determination of components of the reference formulation based on the recipe (R) of the reference formulation, Determination of physical and / or chemical properties of the identified components of each reference formulation, Creating a numerical representation (R F ) for each reference formulation, where the numerical representation (Ri ) of the reference formulation for each determined component is a numerical representation (Rn, R [2 , Ru) of the determined component, wherein the numerical representation (Rn, R [2 , Rn) for each identified component, a value for the amount of the component in the reference formulation and values ​​for the identified physical and / or chemical properties of the component.

3. The method according to claim 1 or 2, wherein the test formulation and each reference formulation is a formulation of a plant protection agent.

4. The method according to any one of claims 1 to 3, wherein the at least one material is a material of a container for a formulation, a material of a lid of a container for a formulation and / or a material of a seal of a container for a formulation.

5. The method according to any one of claims 1 to 4, wherein the at least one material is a plastic or a composite of several plastics.

6. The method according to any one of claims 1 to 5, wherein the at least one material for all reference formulations and the test formulation is the same material or several of the same materials.

7. The method according to any one of claims 1 to 6, wherein the at least one measured stability characteristic (SM) and the at least one predicted stability characteristic (Sp. Sp*) comprise one or more of the following stability characteristics: Swelling of at least one material in contact with the reference formulation / test formulation, Permeation of components of the reference formulation / test formulation into and / or through the at least one material, Tensile impact strength of at least one material after contact with the reference formulation / test formulation, Melt flow rate of the at least one material after contact with the reference F ormulation / Te st-F ormulation, Residual tensile strength of at least one material after contact with the reference F ormulation / Te st-F ormulation .

8. The method according to any one of claims 1 to 7, wherein the recipe (R*, R) of the test formulation and / or reference formulation comprises information about the type of the test formulation and / or reference formulation.

9. The method according to any one of claims 1 to 8, wherein determining components of the test formulation and / or reference formulation based on the recipe (R*, R) comprises: determining impurities of components.

10. The method according to any one of claims 1 to 9, wherein determining components of the test formulation and / or reference formulation based on the recipe (R*, R) comprises: for components of the test formulation and / or reference formulation that consist of more than one chemical compound, determining chemical compounds that are contained in the component.

11. The method according to any one of claims 1 to 10, further comprising: Determine all chemical compounds of all components of the test formulation / reference formulations with a minimum content, Determining a molecular descriptor for each of the identified chemical compounds, wherein determining physical and / or chemical properties of the identified components comprises: Calculating physical and / or chemical properties of the identified chemical compounds based on the molecular descriptors.

12. A method according to any one of claims 1 to 11, wherein the determined physical and / or chemical properties comprise one or more of the following properties: molecular weight, Partition coefficient for w-octanol / water or a value derived therefrom, melting point, Water solubility.

13. The method according to any one of claims 1 to 12, wherein the machine learning model (MLM, MLM T) is configured to generate for each numerical representation (Rn, R12- R13, Rn*, R12*, R13*) of each determined component an optionally compressed numerical representation (R c n, R c i2, R c i3, R c n*, R c i2*, R c i3*) and the optionally compressed numerical representations (R c n, R c i2, R c i3, R c n*, R c i2*, R c i3*) of the components to a common, preferably permutation-invariant representation (R PI , R PI *) of the respective reference formulation / test formulation.

14. Computer system (1) comprising • an input unit (10), • a control and computing unit (20) and • an output unit (30), wherein the control and computing unit (20) is configured to output a trained model (MLM T) of machine learning, where the trained model (MLM T ) of machine learning and has been trained with the aid of training data to predict at least one stability feature (Sp, Sp*) for a system comprising a formulation and at least one material in contact with the formulation on the basis of a numerical representation (RF, RF*) of the formulation, o wherein the training data for each reference formulation of a plurality of reference formulations comprises (i) a numerical representation (RF) of the reference formulation and (ii) at least one measured stability feature (SM), wherein the at least one measured stability feature (SM) is a measure of the stability of a system comprising the reference formulation and at least one material in contact with the reference formulation, o where training the machine learning model (MLM) for each reference formulation comprises: ■ Input the numerical representation (RF) of the reference formulation into the machine learning model (MLM), ■ Receiving at least one predicted stability feature (Sp) from the machine learning model (MLM), ■ Quantifying a deviation between the at least one measured stability characteristic (SM) and the at least one predicted stability characteristic (Sp), ■ Reducing the deviation by modifying the model parameters, causing the input unit (10) to receive a recipe (R*) of a test formulation, To determine components of the test formulation based on the recipe (R*), to determine physical and / or chemical properties of the determined components, to generate a numerical representation (RF*) of the test formulation, wherein the numerical representation (RF*) of the test formulation comprises for each determined component a numerical representation (Rn*, Rn*, R13*) of the determined component, wherein the numerical representation (Rn*, R12*, R13*) of each determined component comprises a value for the amount of the component in the test formulation and values ​​for the determined physical and / or chemical properties of the component, the numerical representation (RF*) of the test formulation is incorporated into the trained model (MLM T ) of machine learning, at least one predicted stability feature (Sp*) from the trained model (MLM T) of the machine learning, to cause the output unit (30) to output and / or store and / or transmit the at least one predicted stability feature (Sp*) to a separate computer system.

15. A computer-readable storage medium comprising instructions which, when executed by a computer system (1), cause the computer system (1) to perform the following steps: Providing a trained model (MLM T ) of machine learning, where the trained model (MLM T) of machine learning and has been trained with the aid of training data to predict at least one stability feature (Sp, Sp*) for a system comprising a formulation and at least one material in contact with the formulation on the basis of a numerical representation (RF, RF*) of the formulation, o wherein the training data for each reference formulation of a plurality of reference formulations comprises (i) a numerical representation (RF) of the reference formulation and (ii) at least one measured stability feature (SM), wherein the at least one measured stability feature (SM) is a measure of the stability of a system comprising the reference formulation and at least one material in contact with the reference formulation, o wherein the training of the machine learning model (MLM) for each reference formulation comprises: Entering the numerical representation (RF) of the reference formulation into the machine learning model (MLM), ■ Receiving at least one predicted stability feature (Sp) from the machine learning model (MLM), ■ Quantifying a deviation between the at least one measured stability characteristic (SM) and the at least one predicted stability characteristic (Sp), ■ Modifying the model parameters to reduce the deviation, Receiving a recipe (R*) of a test formulation, Determination of components of the test formulation based on the recipe (R*), Determination of physical and / or chemical properties of the identified components, Generating a numerical representation (RF*) of the test formulation, wherein the numerical representation (RF*) of the test formulation comprises for each determined component a numerical representation (Rn*, Rn*, R13*) of the determined component, wherein the numerical representation (Rn*, R12*, R13*) of each determined component comprises a value for the amount of the component in the test formulation and values ​​for the determined physical and / or chemical properties of the component, Entering the numerical representation (RF*) of the test formulation into the trained model (MLM T ) of machine learning, Receiving at least one predicted stability feature (Sp*) from the trained model (MLM T ) of machine learning, Outputting and / or storing the at least one predicted stability feature (Sp*) and / or transmitting the at least one predicted stability feature (Sp*) to a separate computer system.