Method for specifying target synthesis specifications

The method addresses the challenges of developing specific models for identifying molecular technical application characteristics by using a trained model parameterized with both searched and unsearched datasets, achieving efficient and accurate results with reduced resource and expertise requirements.

JP2025516307APending Publication Date: 2025-05-27BASF SE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024564901
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-05
Filing Date
2023-05-05
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Developing a specific model for identifying technical application characteristics of molecules is often a difficult, error-prone, and time-consuming process, requiring extensive laboratory resources and expertise.

Method used

A method that uses a trained identification model parameterized based on both searched and unsearched datasets to effectively and objectively identify target molecules with specific technical application characteristics, reducing resource consumption and expertise dependence.

Benefits of technology

The method enables more accurate and efficient identification of molecules with desired technical application characteristics, reducing trial-and-error processes and laboratory resource consumption, while making the model more generally applicable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025516307000001_ABST
    Figure 2025516307000001_ABST
Patent Text Reader

Abstract

The present invention refers to a method for identifying target molecules. A target property and a digital representation of a potential target molecule are provided. A model is then utilized to identify properties of the potential target molecule. The model is parameterized based on a searched dataset and an unsearched dataset. The searched dataset comprises properties for a plurality of searched molecules and molecular feature parameter values ​​for a plurality of searched molecules. The unsearched dataset comprises feature parameter values ​​for a plurality of unsearched molecules. The identified properties of the potential target molecule are compared to the target property. Based on the comparison, either i) the potential target molecule is identified as a target molecule or ii) a new potential molecule is identified and the property identification is repeated. The identified target molecule is then provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Field of the Invention The present invention relates to a method, an apparatus, and a computer program product for identifying a synthesis specification and / or a molecular structure of a target molecule having technical application characteristics. Further, the present invention refers to a method, an apparatus, and a computer program product for generating a machine learning-based identification model that can be used to identify technical application characteristics of a molecule. Moreover, the present invention refers to a method, an apparatus, and a computer program product for identifying molecular feature parameters for training a machine learning-based identification model that can be used to identify technical application characteristics of a molecule. Further, the present invention refers to an interface method, an interface apparatus, and an interface computer program product that provide an interface for the above methods, apparatuses, and computer program products.

Background Art

[0002] Background of the Invention Generally, various molecules, particularly polymers, produced by the chemical industry are widely used in industrial products and / or daily necessities having a wide range of application characteristics. In order to develop new molecules or utilize known molecules in new application situations, it is often useful to have prior knowledge of at least some technical application characteristics, such as the heat insulation coefficient, hardness, or reflectivity of the molecule. For such purposes, in many cases, an identification model trained to identify respective technical application characteristics for a given molecule is utilized, for example, based on machine learning principles.

Summary of the Invention

Problems to be Solved by the Invention

[0003] Summary of the Invention Developing a specific model that enables high accuracy in a particular application scenario is often a difficult, error-prone, and time-consuming process. In particular, providing a training dataset that enables a specific model to be trained to be applicable in a given application scenario often relies solely on expertise and a trial-and-error process where potential training datasets are used to train the model and then the training dataset is revised multiple times based on the results of the training, i.e., based on the likelihood of the trained specific model. Providing such a training dataset is always associated with measuring multiple features and application characteristics for the molecules that are part of the training dataset, so this process also leads to high consumption of laboratory resources such as material resources, human resources, computing resources, etc. Moreover, the well-known training process for specific models consumes a great deal of time and computing resources so as to be extremely useful in avoiding non-productive training, i.e., training that leads to an inappropriate model. Therefore, it would be advantageous to be able to train a specific model more effectively, i.e., with less resource consumption, and more objectively, i.e., based on less expertise, especially to be able to provide training data for training the specific model. Further, typical training processes are based on historical molecular data for a particular application. However, for new applications, the models trained for each are often not very appropriate and completely new models have to be trained by the time- and computing resource-intensive training processes described above for each. Therefore, it would also be advantageous to be able to take into account the respective applicability to new applications already during the training process and make the model more generally applicable.

[0004] An object of the present invention is to provide a method, an apparatus, and a computer program product for a) identifying a target molecule having one or more target application characteristics, b) training a specific model utilized for identifying the target molecule, and c) providing a training dataset for training the specific model, wherein the specific model utilized can be provided to be trained more effectively and more objectively.

Means for Solving the Problem

[0005] In a first aspect of the present invention, there is provided a method executed by a computer for specifying a synthesis specification and / or molecular structure of a target molecule, particularly a target polymer, having target technical application characteristics, the method comprising: i) providing the target technical application characteristics; ii) providing a digital representation of potential target molecules indicating or associated with characteristic parameters; iii) using a trained identification model for identifying the technical application characteristics of the potential target molecules, the trained identification model being parameterized based on a searched dataset and an unsearched dataset such that the trained identification model is adapted to identify the technical application characteristics of a molecule based on one or more molecular characteristic parameters of the molecule, the searched dataset comprising: a) the technical application characteristics for a plurality of searched molecules; and b) a plurality of molecular characteristic parameter values for a plurality of molecular characteristic parameters associated with the synthesis specification and / or molecular structure corresponding to each of the plurality of searched molecules, the unsearched dataset comprising a plurality of molecular characteristic parameter values for a plurality of molecular characteristic parameters associated with the synthesis specification and / or molecular structure corresponding to a plurality of unsearched molecules; iv) comparing the identified technical application characteristics of the potential target molecules with the target technical application characteristics, and based on the comparison, either i) identifying the potential target molecule as the target molecule, or ii) providing a new potential target molecule and repeating the identification of the technical application characteristics using the new potential target molecule; and v) providing the identified target molecule. Since the trained identification model is parameterized based not only on the searched dataset, i.e., the dataset having the technical application characteristics, but also on the unsearched dataset, the inventors have found that the identification model can be trained more effectively and more objectively. In particular, the unsearched dataset can be easily generated with few computational resources and can be used to define possible technical applications of the identification model, for example, the characteristic parameters that can be used to identify the technical application characteristics and / or the molecules to which the trained identification model can be applied.Such generated unexplored datasets can then be utilized to adapt a training dataset based on the explored dataset to such identified technical application scenarios. This can also reduce the amount of trial and error execution for training a specific model, as well as the amount of additional measurements for identifying further explored data. Moreover, based on the unexplored dataset, the applicable area of the trained model can be increased, and thus such trained model is more generally applicable. Furthermore, the influence of expertise on the adjustment of the training dataset can be significantly reduced, leading to more objective training.

[0006] The method refers to a computer-implemented method and can thus be implemented by a general-purpose or special-purpose computer, or a network of computers adapted to implement the method, for example by executing respective computer programs. The method is configured to identify a target molecule having target technical application characteristics. In general, the term "molecule" relates to a group of two or more atoms held together by chemical bonds, and all bonds between non-hydrogen atoms are defined. The bonds may have probabilistic characteristics, such as in a random copolymer. Moreover, each molecule as used herein is associated with a specific molecular structure and / or synthesis specification such that the molecular structure and / or synthesis specification defines the molecule. However, a molecule may be associated with different molecular structures. For example, this is the case for a polymer having a molar mass distribution. Thus, providing a molecule generally also includes providing each molecular structure and / or synthesis specification of the molecule. In one embodiment, the molecule is a small molecule, and herein, a small molecule is defined as a molecule that exists in the environment in a form that enables the complete description of the molecule using a simple structural formula containing relevant information. A simple molecular formula refers to a molecule that can be described by the covalent bonds between the atoms of the molecule. However, the molecule may consist of isomers or may be partially protonated. Examples where this does not apply are systems having a dynamic equilibrium between several forms such as monomers and oligomers, such as in the case of some inorganic acids, or ionic species having highly localized charges that interact strongly with the solvent, for example via hydrogen bonds. Preferably, a small molecule is defined by a molecular weight of less than 600 g / mol, more preferably less than 300 g / mol. Preferably, the target molecule is a target polymer, and the method is configured to identify the target synthesis specification of the target polymer. In general, a synthesis specification is defined as an instruction on how a molecule, particularly a polymer, can be synthesized. In particular, for a polymer, the synthesis specification indicates the starting materials and the respective parameters for the polymerization from the starting materials.

[0007] The technical application characteristics can generally refer to any characteristics of a molecule and / or a substance at least partially composed of molecules, such as a formulation or mixture comprising molecules, which enable the evaluation of the technical applicability of each molecule as provided after synthesis. Preferably, the technical application characteristics are characteristics associated with a substance comprising or consisting of molecules. In particular, the technical application characteristics can be defined by the intended technical application of the molecule and / or the substance at least partially composed of molecules. In particular, the technical application characteristics can directly enable the identification of whether a molecule is suitable for a technical use, while non-technical application characteristics, such as the bond strength between two atoms of a molecule, do not directly enable the identification of the technical application characteristics, and in this regard, the technical application characteristics can be distinguished from non-technical application characteristics. Preferably, the technical application characteristics include at least one of mechanical characteristics, optical characteristics, physicochemical characteristics, chemical characteristics, and biological characteristics. Generally, the mechanical characteristics can refer to any of adhesion, tensile strength, rigidity, hardness, shrinkage, elongation, crack, tear strength, resilience, compressibility, wear, leakage, morphology, tactile characteristics, breaking stress, breaking elongation, particle size distribution, and filling degree. The optical characteristics can generally include any of coloring, turbidity, opacity, gloss, reflection, appearance, absorption, scattering, color strength, color tone, color saturation, chromaticity, cloud point, matting degree, optical density, spectrum, refractive index. Moreover, the physicochemical characteristics can refer to any of density, viscosity, K value, molar weight, dispersity, molar mass distribution, particle size distribution, solubility, partition coefficient, interfacial characteristics, surface tension, dispersibility, storage stability, odor, separation, coagulation, electrical conductivity, electrical capacity, surface area, flow time, vapor pressure, VOC, solids content, hygroscopicity, magnetism, miscibility, thixotropy, phase transition characteristics, glass transition temperature, corrosion inhibition, solvent separation, aggregation, self-heating property, shock sensitivity, loss on drying, reaction angle, electrostatic charge, minimum film-forming temperature, and charge density.The chemical properties may include any of chemical resistance, reaction timing, demolding time, growth, hard / soft segment content, crystallinity, reaction temperature, reaction pressure, decomposition, thermal decomposition, photodegradation, acidity, pKa, pH, moisture / water content, flammability, combustion rate, self-ignition, flash point, generation of flammable gas, reaction to fire, deflagration rate, residual monomer count, generation of by-products, degree of polymerization, salt content, temperature resistance, oxidation characteristics, reduction characteristics, reactivity, ash content, non-volatile substance content, stability, chelating ability, calorific value, saponification value. Further, the biological properties may include biodegradability, biological resistance, particularly resistance to pathogenic viruses, bacteria, fungi, plant or animal origin or the growth stages of said pathogens, resistance to environmental parameters, for example drought resistance, resistance to enzymatic degradation, for example protease resistance, lipase resistance, amylase resistance, hydrolase resistance, pesticide resistance, toxicity, in vivo changes, ecotoxicology, sensitization, particularly allergenicity, bacterial count, enzyme activity, substrate specificity, cofactor dependence, product specificity, inhibition of substrate and / or product, dissociation constant, Michaelis-Menten kinetic values, activity / stability in different pH, temperature, pressure, organic solvent concentration, carrier formulation, encapsulated formulation; distribution, compartmentalization, bioaccumulation, biological exposure LD50, mutagenicity in the environment.

[0008] In a first step, the method includes providing target technical application characteristics. Generally, in the present invention, unless otherwise stated, providing a quantity also refers to quantifying the quantity, for example, providing a value or a range of values. Thus, providing target technical application characteristics also includes quantifying the target technical application characteristics, for example, providing a value or a range of values of the target technical application characteristics. In particular, providing can refer to, for example, receiving the target application characteristics from the input of a user who applies each input unit. Moreover, providing can also refer to accessing a storage unit in which the target application characteristics are already stored. Further, providing can also include, for example, receiving the target application characteristics via a network connection from another source and providing the received target application characteristics. Generally, the target application characteristics can refer to one target value, for example, the intrinsic hardness of a molecule, or a range of values to be satisfied by the molecule. Moreover, the target application characteristics can also refer to any kind of target function, for example, a time series of characteristics under changing environmental conditions, such as hardness under changing temperature conditions. Such more complex target application characteristics can be advantageous when the application of the molecule includes different environmental conditions, such as different temperatures. The target molecule then refers to a molecule that provides each target technical application characteristic when provided in each form, for example, as a pure substance or as a mixture. In particular, the target molecule provides each target technical application characteristic when synthesized according to the corresponding synthesis specification and / or molecular structure.

[0009] Furthermore, the method includes providing a digital representation of a potential target molecule that indicates or is associated with the characteristic parameters of the potential target molecule. Generally, a digital representation is defined as data provided in a form and structure that enables processing in a digital system, particularly a computing system. For example, a digital representation can be a text string, and can be an ID, an image, a chemical formula, etc. Preferably, the digital representation comprises the characteristic parameters of the potential target molecule. Generally, the characteristic parameters of a molecule quantify one or more properties inherent to the molecule itself, particularly the physico-chemical characteristics of the molecule itself. Thus, while the characteristic parameters refer to the molecule itself, the technical application characteristics refer to the properties of the molecule in its technical application form, for example, as a substance comprising or consisting of the molecule. Preferably, the characteristic parameters are associated with the synthesis specification and / or the molecular structure. Here, "associated" means that the characteristic parameters can directly refer to the synthesis specification and / or the molecular structure, for example, as a structural or process parameter, or can be derived from the synthesis specification and / or the molecular structure as a calculated physico-chemical property that can be calculated from the synthesis specification and / or the molecular structure. Preferably, for molecules that can be defined by a simple structural formula, since the molecular structure clearly defines the molecule, the characteristic parameters are associated with the molecular structure. However, even in this case, the synthesis specification enables the derivation of appropriate characteristic parameters. In many cases, for more complex molecules, which may exist in multiple different structures or for which it is often difficult to identify the exact molecular structure, particularly for polymers, the characteristic parameters are preferably associated with the synthesis specification. However, even in these cases, in many cases, at least some parts of the molecular structure can be defined and then used to identify the characteristic parameters. Generally, the characteristic parameters can comprise any of the process parameters, recipe parameters, synthesis route, or physico-chemical parameters of the molecule. The process parameters, in this context, quantify the manufacturing process of each molecule and can be derived from the synthesis specification of the molecule.Certain process parameters lead to the production of specific molecules, so the process parameters also characterize the molecules themselves. For example, the process parameters can refer to any of a temperature profile, pressure, stirring force, synthesis step, etc. Recipe parameters generally quantify the production of the molecules themselves, in particular the substances used to produce the molecules. Such substances can refer to starting materials from which the molecules are produced, as in the case where the molecules are polymers and their respective prepolymers. However, the substances can also refer to auxiliary substances such as catalysts. In this context, the recipe parameters quantify the influence of these substances on the produced molecules and thus also characterize the molecules themselves. For example, the recipe parameters can refer to any of the amount of a specific chemical substance, the amount of a specific additive, the mixing ratio between different substances, or the solvent used. Preferably, the digital representation preferably comprises molecular descriptors, i.e., molecular physicochemical parameters that refer to calculated physicochemical molecular properties, and the molecular physicochemical parameters indicate the physicochemical characteristics of the molecules. In particular, the molecular physicochemical parameters indicate parameters that quantify the physicochemical characteristics of the molecules. In this context, the term "physicochemical characteristics" refers to the physical and / or chemical characteristics of the molecules. For example, the physicochemical parameters can refer to any of molar mass distribution, viscosity, degree of protonation, etc. However, the digital representation can also be provided to enable the derivation of molecular physicochemical parameters, for example, by providing a representation of the molecules in which each molecular physicochemical parameter is already stored or can be identified by, for example, each molecular descriptor calculation. Preferably, the digital representation comprises at least one of a recipe of the molecule, a structural formula, a brand name, an IUPAC name, a chemical identifier, and a CAS number.

[0010] In a preferred embodiment, the target molecule is a polymer, and the characteristic parameter indicates a parameter that quantifies the physicochemical characteristics of a subgroup of the polymer. In this embodiment, the digital representation can also be provided to enable the derivation of characteristic parameters by identifying subgroups of the polymer and the identification of characteristic parameters based on the physicochemical characteristics of the identified subgroups. Generally, a subgroup refers to a part of the polymer, and all subgroups of the polymer together form the polymer. For example, a subgroup can refer to a part of the polymer, and subgroups can be continuously linked together along a chain or network to form the polymer. Preferably, a subgroup of the polymer refers to a repeating unit that describes a part of the polymer that, when repeated, generates a complete polymer chain. However, in some cases, a subgroup can also refer to a single non-repeating part of the polymer. Moreover, a subgroup preferably includes a repeating part. For example, a subgroup of the polymer can include a repeating core that is also present in other subgroups and an additional part that is not present in other subgroups. Preferably, a subgroup refers to at least one polymerized monomer or oligomer fragment. More preferably, a subgroup refers to a polymerized monomer. In this context, polymerized monomers refer to the monomers after their polymerization and may also be referred to as "mer units" or "mers". In particular, polymerized monomers do not refer to the monomers present in the reaction mixture before polymerization, i.e., the raw materials, but rather to repeating units derived from monomers that have changed during or after polymerization. Therefore, the subgroup descriptors specified for polymerized monomers are different from the subgroup descriptors specified for unreacted monomers before polymerization. In particular, the inventors have found that polymerized monomers enable the identification of polymer descriptors from subgroup descriptors of polymerized monomers that enable a specific model to accurately identify the technical application characteristics of the polymer. In a preferred embodiment, the digital representation of the polymer comprises subgroups provided as molecular models showing the chemical structure of the subgroups after their polymerization.Even more preferably, the molecular model of the subgroup is identified by a method suitable for quantum chemical calculations regarding the number of atoms representing the characteristics of the subgroup in the polymer and their connectivity. Moreover, in addition to or instead of the molecular model of the subgroup that treats the subgroup as a monomer structure, a molecular model referring to an oligomer model that takes into account the influence of the adjacent molecular structure of the subgroup in the polymer may also be used.

[0011] In one embodiment, when the digital representation of the molecule does not directly include the characteristic parameter, the characteristic parameter is preferably identified by forming the indicated molecular structure and / or synthetic specification of the provided molecule. If the molecular structure and / or synthetic specification are not provided as part of the digital representation, the molecular structure and / or synthetic specification can be derived from information about the molecule itself by using known methods for identifying synthetic specifications, such as respective molecular databases or data-driven retrosynthetic pathway planners. Moreover, in this case, the molecular structure and / or synthetic specification can also be required to be provided by the user. When the molecule is a polymer, the characteristic parameter is preferably identified based on the provided synthetic specification instead of the molecular structure, which cannot often be identified for polymers. Further, in this case, when the characteristic parameter is based on the characteristic parameters of a subgroup, the characteristic parameter is preferably identified by first identifying the subgroup of the polymer. For example, each subgroup of the polymer can be identified using known methods. In particular, between atoms of different subgroups in the polymer, the subgroup is preferably identified such that the polarization in the bond is as small as possible, and preferably the bond order is as small as possible (e.g., CC single bond). In addition, the subgroup representing the polymer preferably contains the same number of active non-hydrogen atoms as the polymer. In addition to the active atoms, the subgroup can also contain additional atoms that can be ignored during the calculation of the physicochemical parameters of the subgroup. Further, the subgroup is preferably identified such that polymers containing portions constructed using different polymerization techniques are sufficiently covered and meet the above conditions. One example is a polyether used as a component of polyurethane. Generally, a database or archive having a plurality of reactions between polymer parts can be created, and subgroups can be obtained from the structure of each reaction. For example, specific chemical languages such as SMILES and SMARTS can be used to easily derive subgroups of the polymer.For example, a database of reaction SMARTS can be generated, and then, based on the polymerization of each polymer, the corresponding reaction SMARTS can be selected. From the selected reaction SMARTS, the SMILES of the monomers of the polymer can be directly derived. For example, using RDkit, from the SMILES of the monomers, the SMILES of the subgroups, i.e., the number of atoms and the bonding properties, can be identified.

[0012] The identified subgroups of the polymer are particularly associated with subgroup characteristic parameters that indicate parameters quantifying the physicochemical characteristics of the subgroups in the polymer. In particular, when the characteristic parameters are not directly provided by a digital representation, the characteristic parameters are preferably identified by identifying respective subgroup characteristic parameters for each of the subgroups and then identifying the characteristic parameters, for example, by averaging, based on the subgroup characteristic parameters of the subgroups. Thus, in this embodiment, the method preferably first provides or identifies subgroups of the polymer from a digital representation, then identifies or provides subgroup characteristic parameters, i.e., the values of the parameters quantifying the characteristic parameters of the subgroups, and then identifies polymer characteristic parameters based on the subgroup characteristic parameters of each polymer.

[0013] Preferably, the molecular feature parameter includes at least one physicochemical parameter indicating a descriptor, a count descriptor, a list of structural fragments, a fingerprint, a graph invariant, a 3D descriptor, and / or a higher-dimensional descriptor that quantifies the physicochemical characteristics of the molecule. Moreover, in the case of a polymer, the molecular feature parameter preferably includes, as a physicochemical parameter, the molar mass distribution of the polymer. In the case of a polymer, the molecular feature parameter more preferably includes, as a process parameter, a temperature profile. Further, the molecular feature parameter preferably includes, as a recipe parameter, the amounts of the respective components. Generally, the molecular feature parameter can also preferably include any combination of possible parameters described above or below.

[0014] In a preferred embodiment, the molecular feature parameter includes a 3D descriptor, particularly a calculated physicochemical property, and / or a fingerprint. When the molecule is a polymer, the molecular feature parameter can be derived from the parameters of the subgroups, and thus, the subgroup feature parameter can also refer to the same feature parameters as described above. Generally, the feature parameter can also be derived without using subgroups, for example, by simulation of the entire molecule. Possible feature parameters are defined in more detail below. Also, in these cases, the defined feature parameter can directly refer to the molecular feature parameter or, optionally, for polymer embodiments, the subgroup feature parameter.

[0015] The compositional descriptor can refer to any of potential, average molecular weight, polydispersity, charge, spin, boiling point, melting point, enthalpy of fusion, dissociation constant, Hansen parameters, protonicity, polar and dispersive contributions, Abraham parameters, retention index, TPSA, twist angle, degree of branching, nucleotide and / or amino acid composition, nucleotide and / or amino acid sequence, nucleotide and / or amino acid sequence conservation, isoelectric point, glycosylation pattern, receptor binding constant, inhibitor constant. The count descriptor can refer to any of the sum of atomic electronegativities, sum of atomic polarizabilities, amount of component, ratio of amounts of components, number of atoms and non-hydrogen atoms, number of H, B, C, N, O, P, S, Hal, and heavy atoms, number of hydrogen donor and hydrogen acceptor atoms, number of bonds, number of non-hydrogen or multiple bonds, number of double bonds, triple bonds, and aromatic bonds, number of functional groups, ratio of functional groups, sum of bond orders, aromatic ratio, number of rings or circuits, number of unpaired electrons, number of rotatable bonds, fraction of rotatable bonds, and number of conformers. The molecular descriptor that refers to a list of structural fragment descriptors can refer to at least one of a list of molecular fractions, a list of functional groups, a list of bonds, and a list of atoms. The fingerprint descriptor preferably includes at least one of MACCS keys, preferably in bit format or total format, Morgan fingerprints and other circular fingerprints, preferably in bit format or total format, topological twists, atom pairs, infrared spectra and related spectra, fingerprint numbers, PubChem fingerprints, substructure fingerprints, and Klekota-Roth fingerprints. The graph invariance / topology index descriptor preferably includes at least one of a topostructural index and a topochemical index.In a preferred embodiment, the characteristic parameter includes at least one or more of the total volume of all atoms, the average volume per atom, the total area of all atoms, the average area per atom, the area of all atoms, the average area per atom, the solvent-accessible surface, the dispersion energy, the dielectric energy, H donors, H acceptors, polar and non-polar surface areas, atom-resolved H donors, H acceptors, polar and non-polar surface areas, shape, spherical, dipole and higher-order electric moments, polarizability, dielectric energy, protonicity, polar and non-polar surface areas, orbital energy and orbital gap, ionization energy, electron affinity, hardness, electronegativity, electrophilicity, excitation energy and intensity, infrared and ultraviolet absorption bands, reactivity measurements, redox potential, binding reference point, partial charge, charge surface area, atomic orbital contribution, bond order, atomic radius, etc., and comprises a 3D descriptor. In particular, the molecular characteristic parameter preferably refers to a 3D descriptor including at least one of the total volume of all atoms, the average volume per atom, the total area of all atoms, the average area per atom, the solvent-accessible surface, the dispersion energy, the dielectric energy, H donors, H acceptors, polar and / or non-polar surface areas, atom-resolved H donors, H acceptors, polar and / or non-polar surface areas, shape, sphericity, cone angle, polarizability, dielectric energy, protonicity, polar and / or non-polar surface areas, excitation energy and intensity, infrared and / or ultraviolet absorption bands, reactivity measurements, particle charge, and / or charge surface area. Preferably, the high-dimensional descriptors used can include at least one or more of the conformational partition function, solubility, vapor pressure, activity coefficient, diffusion coefficient, partition coefficient, surface activity, rotation constant, moment of inertia, radius of gyration, compositional drift of the molecule, density, viscosity, conformation-weighted volume and area, conformation-weighted H donors, H acceptors, protonicity, polar and / or non-polar surface areas, charge distribution, conformational dipole moment, molecular refraction, etc. Preferably, high-dimensional descriptors including at least one of solubility, vapor pressure and activity coefficient, surface activity, conformation-weighted H donors, H acceptors, protonicity, polar and non-polar surface areas, and charge distribution are used.

[0016] In the next step, the trained specific model is utilized to identify the technical application characteristics of the molecule. Preferably, the utilization of the specific model can include, for example, providing or receiving the specific model from each storage unit and utilizing the provided or received specific model. The trained specific model is parameterized based on the explored dataset and the unexplored dataset such that the trained specific model is adapted to identify the technical application characteristics of the molecule based on one or more molecular feature parameters of the molecule. Generally, the parameterization of the trained specific model refers to the identification of the parameters of the specific model in the training process, and thus the specific model is trained based on the explored dataset and the unexplored dataset. The term "such that" herein should be interpreted as parameterization, i.e., training, adapting the specific model to identify the technical application characteristics of the molecule based on one or more molecular feature parameters of the molecule, and thus enabling the specific model to identify the technical application characteristics of the molecule. In particular, the specific model is a data-driven model, and the term "data-driven" is used to emphasize that the model is mainly based on each data input and not, for example, on intuition, personal experience, or knowledge. Preferably, the specific model is based on known machine learning algorithms such as neural networks, regression models, machine learning algorithms, etc. In this regard, for most applications, in particular, regression models based on linear regression, random forest, Lasso, boost tree, ridge regression, and MARS algorithms are appropriate, while for classification models, in particular, random forest, logistic regression, and SVM algorithms have been found to be appropriate. Generally, the specific model is parameterized during the training process, and in the training process, the explored dataset and the unexplored dataset are utilized for the training of the specific model.

[0017] Generally, an explored dataset comprises: a) the technical application characteristics for a plurality of explored molecules; and b) a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with the synthesis specifications and / or molecular structures corresponding to each of the plurality of explored molecules. Thus, an explored dataset refers to a dataset comprising only molecules for which at least one technical application characteristic has been identified at least once. Generally, an explored molecule refers to a molecule for which at least one technical application characteristic has been identified. In the expression “molecular feature parameters associated with the synthesis specifications and / or molecular structures”, the term “associated with” should be construed herein as being “defined by” or “derivable from” each synthesis specification and / or molecular structure, as already explained above with respect to the feature parameters of potential target molecules. The technical application characteristics of an explored molecule can be identified in any known way, for example, based on experimental data and measurements, or based on simulations or theoretical considerations. Preferably, the technical application characteristics refer to the measured technical application characteristics. For example, to measure the technical application characteristics, each molecule can be synthesized and each technical application characteristic can be measured with an appropriate measurement process. However, the technical application characteristics can also be identified based on each physical or chemical calculation, or based on each simulation of the molecule. The molecular feature parameter values associated with the molecule can also be measured, but can also be identified by any other known method, for example, calculated as already explained above. In one embodiment, the explored dataset can further comprise producibility information for one or more of the plurality of molecules in the explored dataset, and the producibility information indicates which technical application characteristics of the molecule can be measured. For example, during the synthesis of a molecule, it can be identified that the molecule cannot be produced by generally known synthesis processes, or that the synthesized molecule is not suitable for a particular measurement process that yields each technical application characteristic. Such information can be provided as part of the producibility information. Generally, the producibility information can refer to category information indicating whether the molecule refers to one or more categories.For example, the category can refer to whether a molecule can be produced or whether a particular application property can be measured for the molecule. As part of the explored dataset, in addition to or instead of the technical application properties, by providing information about at least some of the explored molecules, a specific model can be further trained based on these explored molecules to, for example, at least pay attention when a molecule may not be suitable for an application due to synthesis problems. Generally, for this embodiment, it is preferable that the technical application properties are provided for more explored molecules than the explored molecules for which productivity information is provided instead of the technical application properties.

[0018] The unexplored dataset comprises a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with synthetic specifications and / or molecular structures corresponding to a plurality of unexplored molecules. Preferably, the unexplored dataset consists of a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with synthetic specifications and / or molecular structures corresponding to a plurality of unexplored molecules. Thus, the unexplored dataset does not have technical application characteristics associated with the unexplored molecules. This enables the easy and rapid generation of a large number of unexplored molecules, especially the molecular structures and / or synthetic specifications, and the corresponding molecular feature parameter values for generating the unexplored dataset. For example, in the case of a polymer based on a starting synthetic specification, the parameters of the synthetic specification can be changed to generate a new synthesis. In the case of small molecules, a database query of similar molecules based on a fingerprint can be performed. An alternative would be to select a diverse subset of molecules from a larger set of possible molecules spanning a certain chemical application range. Moreover, any variations from the starting synthetic specification and / or molecular structure can also be utilized to generate an enormous amount of unexplored synthetic specifications and / or molecular structures. Based on these synthetic specifications and / or molecular structures, for example, known methods as described above can then be utilized to identify the respective molecular feature parameter values for the unexplored molecules associated with each synthetic specification and / or molecular structure. In this way, for example, a huge unexplored dataset comprising thousands of unexplored molecules can be generated. In particular, the explored dataset and the unexplored dataset each comprise at least some, preferably all, of the molecular feature parameter values for the same molecular feature parameters for each molecule. For example, the unexplored dataset can be generated such that the molecular feature parameters for which each value is provided in the explored dataset are calculated. However, in the explored dataset, the molecular feature parameters referring to the molecular feature parameters provided in the unexplored dataset can also be calculated. In particular, the generated unexplored dataset can be utilized to define the application context of a specific model to be trained.In particular, the unexplored dataset can be generated taking into account each of the constraints for the intended application of the model. Based on the unexplored dataset defining the application scenario, the explored dataset can then be utilized to generate a training dataset for a specific model. For example, from the explored dataset, explored molecules and / or molecular feature parameters can then be selected to form a training dataset. The inventors have thus been able to provide a training dataset for training a specific model tailored to the application scenario of the technical application characteristics without a cumbersome trial-and-error process and without consuming high computational resources. Such a trained specific model then enables a particularly accurate identification of the technical application characteristics of molecules falling within the application scenario of the specific model.

[0019] In a further step, the identified technical application characteristics of the potential target molecule are then compared with the target technical application characteristics. Based on the comparison, it is then determined whether a) the potential target molecule is identified as the target molecule or b) the potential target molecule is provided and the identification of the technical application characteristics is repeated using a new potential target molecule. Preferably, step a) is carried out when the comparison determines that the identified technical application characteristics of the potential target molecule meet the target technical application characteristics, and step b) is carried out when the comparison determines that the identified technical application characteristics of the potential target molecule do not meet the target technical application characteristics. In particular, the comparison of the identified technical application characteristics with the target technical application characteristics makes it possible to determine whether the identified technical application characteristics meet a predetermined criterion, for example, whether the identified technical application characteristics meet the target technical application characteristics within a predetermined limit. If such a criterion is met, the potential target molecule is identified as the target molecule and the method proceeds to the next step. However, if the comparison indicates that the identified technical application characteristics do not meet the target technical application characteristics within a predetermined limit, the next iterative step of using a new potential target molecule must be processed. In particular, for each iterative step of the iteration, preferably based on the previous potential target molecule, for example, by modifying one or more characteristics of the previous synthesis specification and / or the molecular structure of the previous molecule, a new potential target molecule is identified. However, the new potential target molecule can also be generated by arbitrarily selecting a new potential target synthesis specification and / or molecular structure from a large number of previously generated potential target synthesis specifications and / or molecular structures. Moreover, a more sophisticated method can also be used to select a new potential target molecule from a plurality of potential target molecules that have already been generated previously. Based on the new potential target molecule, in each iterative step, again, the trained identification model is used to identify the technical application characteristics, and such identified technical application characteristics are compared again with the target technical application characteristics so that the comparison can lead to a further iterative step again, or so that each new potential target molecule can be selected as the target molecule if each criterion is met.In addition, additional stopping criteria for the iteration can also be selected. For example, the number of iteration steps can be specified before notifying the user that the iteration is stopped because the target molecule could not be found for each technical application characteristic. However, alternatively, after a predetermined number of iteration steps, the method can further include modifying the target technical application characteristics, for example, by increasing a predetermined limit around the technical application characteristics and repeating the iteration while utilizing the increased limit during the comparison. This can make it possible to find a target molecule that satisfies the technical application characteristics as much as possible even if the original goal cannot be met. After the target molecule is identified as described above, the target molecule and optionally additional information such as a molecular structure or synthesis specification for generating the molecule can be provided to the user, for example, via an output unit. In the case of small molecules, a data-driven retrosynthesis planner that provides potential starting materials for the synthesis of the target molecule can be used. Subsequently, a database containing many published synthetic routes can be used. This gives possible instructions for synthesizing the target molecule from the proposed starting materials. Alternatively, data-driven tools can be used to provide synthesis instructions as well as an estimated likelihood of synthesis success.

[0020] In one embodiment, providing a specified target molecule includes generating a control signal based on the target molecule and providing the control signal, where the control signal is configured to control a production system for producing the target molecule according to a target synthesis specification. Preferably, the control signal is provided to each production system configured to produce the target molecule based on the control signal. Generally, the control signal can be provided in any form that enables directly or indirectly controlling a production system for producing the target molecule according to a target synthesis specification. For example, the control signal can be provided in a form that enables a production system management application to interpret the control signal and then control a production system, such as a synthesis robot, and accordingly produce the target molecule. However, the control signal can also be provided in a form that directly enables controlling each production system for producing the target molecule, for example, directly enabling controlling components, such as starting or stopping a heating unit, opening or closing a valve, starting or stopping a mixer, etc. Generally, the production system can refer to any fully or partially automated production system provided in a laboratory or industrial environment that is generally configured to produce molecules based on a synthesis specification.

[0021] In one embodiment, the parameterization of the machine learning-based identification model is based on identifying a similarity measure between members of the explored dataset and members of the unexplored dataset. Generally, in this regard, the similarity measure can be defined as a measure that identifies the distance between one or more members of the explored dataset and one or more members of the unexplored dataset in a given space. Preferably, the given space is defined by one or more molecular feature parameters. Thus, the similarity measure in this case indicates how close one or more members of the explored dataset are to one or more members of the unexplored dataset with respect to the one or more molecular feature parameters. In this regard, members of the explored dataset refer to the molecules provided by each explored dataset and all quantities associated with those molecules, and members of the unexplored dataset refer to the molecules provided by the unexplored dataset and all quantities associated with those molecules. In particular, the similarity measure between members of the explored dataset and members of the unexplored dataset is preferably specified with respect to the molecular feature parameters of the members of the explored dataset and the unexplored dataset. In particular, the similarity measure indicates the distance between the explored dataset and members of the unexplored dataset with respect to a subset of the molecular feature parameters. By utilizing the similarity measure between members of the explored dataset and members of the unexplored dataset, it becomes possible to identify which member of the explored dataset is most similar to a member of the unexplored dataset for different molecular feature parameters. This also makes it possible to select the most suitable combination of molecular feature parameters not only for members of the explored dataset but also for a part of the training dataset for training an identification model for the technical application scenarios defined by the unexplored dataset.

[0022] Preferably, the parameterization of the machine learning-based identification model includes selecting, based on a similarity measure, a subset of the molecular feature parameters of the explored data set as the training molecular feature parameters. The parameterization of the machine learning-based identification model utilizes a training data set that includes, from the explored data set: a) the technical application characteristics for at least two of the plurality of explored molecules; and b) the molecular feature parameter values for the training molecular feature parameters associated with the synthesis specifications and / or molecular structures corresponding to each of the at least two explored molecules. Generally, the selected subset of the molecular feature parameters of the explored data set can refer to a combination of one or more molecular feature parameters whose values are provided by the explored data set. In particular, the subset of molecular feature parameters is selected by comparing similarity measures for different combinations of the molecular feature parameters of the members of the explored data set and the members of the unexplored data set, and it is preferable to identify the combination in which the members of the explored data set are most similar to the members of the unexplored data set from these combinations, and then each combination of molecular feature parameters can be selected as a subset. Therefore, based on the similarity measure, it is possible to identify which combination of molecular feature parameters enables the explored data set to optimally cover the technical application situation of the identification model defined by the unexplored data set.

[0023] Preferably, a subset of the molecular feature parameters is selected based on the optimization of a similarity measure for the subset of molecular feature parameters between members of the unexplored dataset and members of the explored dataset. In particular, the optimization preferably refers to identifying a subset of molecular feature parameters for which members of the unexplored dataset and members of the explored dataset are most similar. In other words, the similarity measure is optimized such that members of the explored dataset cover the portion of the space defined by the unexplored dataset as well as possible in the space defined by the subset of molecular feature parameters. Thus, by training a specific model with such a specified training dataset, it becomes possible to generate a specific model that is particularly well-suited to the intended technical application situation without the need to perform trial and error or conduct additional experiments to create additional explored molecules.

[0024] Preferably, the similarity measure is mathematically defined as follows for optimization.

[0025]

Number

[0026] Here,

[0027]

Number

[0028] are molecular feature parameters from the unexplored (scr) and explored (tr) datasets, where f refers to a specific feature parameter, i and j refer to the respective molecules of the unexplored and explored datasets, the coefficient cf can be 0 or 1 depending on whether each feature parameter is selected, the sum over f is equal to the number K of feature parameters in the selected subset of feature parameters, D is a general distance function between any two given data points in the space defined by the selected subset of feature parameters, and S is the number of data points in the unexplored dataset. The number of feature parameters in the selected subset of feature parameters K can refer to a predetermined number. For example, fewer than 3 feature parameters may result in less accurate results when specifying the technical application characteristics by such a trained specific model, while it is known from experience that it may not be advantageous to have to provide a specific model with more than 20 feature parameters to specify the technical application characteristics. For example, a large number of feature parameters that have to be provided to a specific model to specify the technical application characteristics may lead to a decrease in the predictability of the model obtained as a result for a dataset different from the dataset used for training the model, which may lead to overfitting of the model to be avoided. Thus, for example, the number of feature parameters for a subset of feature parameters can be specified in advance based on considerations such as those outlined above. However, it is preferred that the number of feature parameters in the selected subset of feature parameters itself undergoes an optimization process and can thus be regarded as a hyperparameter. In particular, the optimization of the similarity measure can also include an optimization with respect to the number of feature parameters in the selected subset of feature parameters such that the number does not need to be specified or fixed in advance during the optimization process.

[0029] In a preferred embodiment, the subset is selected by further optimizing a specific accuracy of a specific model with respect to a subset of molecular feature parameters. By further selecting a subset of feature parameters based on the optimization of the specific accuracy of a specific model with respect to a subset of molecular feature parameters, it is possible to already take into account the potential accuracy of the specific model during the identification of the training dataset. This makes it possible to optimize the selection of feature parameters within the training dataset already during the identification of the training dataset with respect to the specific accuracy. Preferably, the specific accuracy is mathematically defined as follows for optimization.

[0030]

Number

[0031] Here,

[0032]

Number

[0033] refers to the technical application characteristics specified by a specific model trained by each training dataset based on each selected subset,

[0034]

Number

[0035] refers to the technical application characteristics for each molecule, T is the number of data points in the explored dataset, D* is the general distance measured between the technical application characteristics specified for the molecules of the explored dataset by the specific model to be trained and the technical application characteristics of the molecules of the explored dataset, and the parameter λ refers to the weighting of the specific accuracy with respect to the other terms of the optimization.

[0036] In a preferred embodiment, the subset is selected by further optimizing the applicability of a particular model in the unexplored dataset. In particular, the applicability of a particular model in the unexplored dataset can be measured by using criteria that specify the coverage of the space defined by the unexplored dataset by the explored dataset with respect to the selected subset of molecular feature parameters. For example, if the explored dataset covers only a very small portion of the space defined by the unexplored dataset, i.e., the coverage is incomplete, the particular model may not be applicable to the complete intended technical application scenarios defined by the unexplored dataset. Therefore, considering the applicability of a particular model in the unexplored dataset, for example, by using a coverage metric, the training dataset can also be directly optimized with respect to applicability and the intended applicability of the particular model can be ensured as much as possible. Preferably, the applicability is mathematically defined as follows for optimization.

[0037]

Number

[0038] Here,

[0039]

Number

[0040] refers to the technical application characteristics specified by the particular model being trained,

[0041]

Number

[0042] refers to the technical application characteristics for each molecule, G is a function for evaluating the applicability of a specific model in an unexplored dataset, for example, the Euclidean distance, S is the number of data points in the unexplored dataset, and the parameter μ refers to the weighting of the applicability measure relative to other terms of the optimization. Generally, the weighting parameters μ and λ for weighting additional measures for the similarity measure during optimization can refer to predetermined values specified, for example, based on user experience, theoretical considerations, or resampling methods. For example, an increase in μ increases the application space for a specific model around the explored dataset, and a decrease also decreases the application space of the specific model. λ strongly specifies whether the decision model follows the explored dataset. However, these parameters can also be subject to optimization, that is, they can also be regarded as hyperparameters that can be changed during optimization.

[0043] In one embodiment, selecting a subset of molecular feature parameters includes: a) identifying a parameter training space based on an unexplored dataset, where the parameter training space is defined by the molecular feature parameter values of the molecular feature parameters of a plurality of unexplored molecules; b) identifying a subspace of the parameter training space such that the molecular feature parameter values of the molecular feature parameters of the explored dataset cover the subspace according to a predetermined criterion based on a similarity measure; and c) selecting a subset of the molecular feature parameters based on the identified subspace. Generally, the predetermined criterion can also be based on other measures already defined above as an addition. Therefore, optimizing the similarity measure and optionally also one or more of the measures defined above for the molecular feature parameters also refers to optimizing the subspace of the parameter training space such that the explored dataset covers the subspace as well as possible.

[0044] In one embodiment, providing an unexplored dataset includes a) receiving information indicating constraints for generating molecules, and b) generating an unexplored dataset based on the received information. The constraints can be any constraints for the generation of molecules, such as constraints provided by the technical constraints of the production system for generating molecules, such as production constraints referring to constraints in the availability of starting materials or catalysts or other necessary substances for generating molecules, or physical or chemical constraints that prevent the generation of certain molecules. Such constraints can generally be applicable to all molecules, for example, when they refer to physical or chemical constraints, and thus can generally be considered when generating an unexplored dataset, or can refer to specific constraints, such as constraints of the technical facilities of potential users that must be considered when generating an unexplored dataset. For example, if the technical constraints indicate that only the production facilities of potential users can enable the production of molecules within a specific temperature range and based on a specific starting material, the synthesis specifications of the unexplored dataset can be easily generated taking these constraints into account so that only molecules can be considered in the unexplored dataset that can be generated by each production facility of potential users. In a preferred embodiment, the constraints restrict the potential target molecules to molecules whose respective synthesis specifications are known or specifiable. Generally, the above-mentioned constraints also make it possible to specify a training dataset for a specific model so that the specific model is specifically trained for this intended application, i.e., trained only for molecules that can be produced by the production facility. Moreover, during the identification of the target molecules, it is preferable that new potential target molecules are generated taking into account each constraint so that only potential target molecules are considered during the iteration that can actually be generated by each target facility of potential users.

[0045] In one embodiment, parameterization of a machine learning-based identification model includes selecting, based on an explored dataset and an unexplored dataset, a subset of molecular feature parameters from the molecular feature parameters of the explored dataset as training molecular feature parameters. The parameterization of the machine learning-based identification model utilizes a training dataset comprising, from the explored dataset, a) technical application characteristics for at least two of the plurality of explored molecules and b) molecular feature parameter values for the training molecular feature parameters associated with the synthesis specifications and / or molecular structures corresponding to each of the at least two explored molecules. Preferably, selecting the subset includes a) specifying a parameter training space based on the unexplored dataset, the parameter training space being defined by the molecular feature parameter values of the molecular feature parameters of a plurality of unexplored molecules, b) specifying a subspace of the parameter training space such that the molecular feature parameter values of the molecular feature parameters of the explored dataset cover the subspace according to a predetermined criterion, and c) selecting the training molecular feature parameters based on the specified subspace. Further, the predetermined criterion preferably refers to the similarity between members of the explored dataset and members of the unexplored dataset. In particular, the similarity can be specified using a similarity measure as already described above. Moreover, the predetermined criterion can additionally refer to other measures already defined above, and specifying each subspace can refer to the optimization of the similarity measure and optionally the additional optimization of any of the other measures already described above.

[0046] In a further aspect of the present invention, there is provided a method executed by a computer for generating a machine learning-based identification model such that the trained identification model is adapted to identify the technical application characteristics of a molecule based on one or more molecular feature parameters, the molecular feature parameters being associated with the synthesis specification and / or molecular structure of the molecule, the method comprising: i) providing a searched data set comprising: a) the technical application characteristics for a plurality of searched molecules; and b) a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with the synthesis specification and / or molecular structure corresponding to each of the plurality of searched molecules; ii) providing an unsearched data set comprising a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with the synthesis specification and / or molecular structure corresponding to a plurality of unsearched molecules; iii) generating a trained identification model by parameterizing a machine learning-based identification model based on the searched data set and the unsearched data set such that the trained identification model is adapted to identify the technical application characteristics of a molecule based on one or more molecular feature parameters of the molecule; and iv) providing the generated trained identification model. Preferably, providing the unsearched data set comprises generating the unsearched data set based on the intended application of the identification model. In particular, as already explained above, the plurality of synthesis specifications and / or molecular structures can be generated based on the intended application, in particular based on constraints.

[0047] In a further aspect, a method performed by a computer for identifying one or more molecular feature parameters for training a machine learning-based identification model such that the trained identification model is adapted to identify the technical application characteristics of a molecule based on the one or more molecular feature parameters is presented, the molecular feature parameters being associated with a synthesis specification and / or molecular structure corresponding to the molecule, the method comprising: i) providing a searched dataset comprising: a) technical application characteristics for a plurality of searched molecules; and b) a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with a synthesis specification corresponding to each of the plurality of searched molecules; ii) providing an unsearched dataset comprising a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with a synthesis specification and / or molecular structure corresponding to a plurality of unsearched molecules; iii) selecting, based on the searched dataset and the unsearched dataset, one or more molecular feature parameters from the feature parameters of the searched dataset; and iv) providing the selected molecular feature parameters. In one embodiment, providing the unsearched dataset preferably includes generating the unsearched dataset based on the intended application of the identification model. In particular, as already explained above, the plurality of synthesis specifications and / or molecular structures can be generated based on the intended application, in particular based on constraints.

[0048] Generally, the same embodiments and definitions described with respect to the method for identifying a target molecule can also be applied to the respective features of the method for identifying an identification model and the method for identifying molecular feature parameters for training an identification model.

[0049] In a further aspect, an interface method for providing a specific model is presented. The interface: i) receives, via an input unit, a searched dataset comprising, for a plurality of searched molecules, a) technical application characteristics and b) a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with the synthesis specifications and / or molecular structures corresponding to each of the plurality of searched molecules; ii) receives, via the input unit, information indicating an unsearched dataset comprising a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with the synthesis specifications and / or molecular structures corresponding to a plurality of unsearched molecules; iii) is interfaced with a processor that executes a method as described above for providing the searched dataset and the information via an interface unit, and receives a trained specific model; and iv) provides the trained specific model via an output unit. In one embodiment, the interface method can include steps executed by a processor according to a method as described above for providing the searched dataset and the information.

[0050] In a further aspect, an interface method for providing one or more molecular feature parameters for training a machine learning-based identification model is presented. The interface includes: i) receiving, via an input unit, a searched dataset comprising a) technical application characteristics for a plurality of searched molecules and b) a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with the synthesis specifications and / or molecular structures corresponding to each of the plurality of searched molecules; ii) receiving, via the input unit, information indicating an unsearched dataset comprising a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with the synthesis specifications and / or molecular structures corresponding to a plurality of unsearched molecules; iii) interfacing, via an interface unit, with a processor that executes a method as described above for providing the searched dataset and the information, and receiving identified molecular feature parameters for training a machine learning-based identification model; and iv) providing, via an output unit, the identified molecular feature parameters. In one embodiment, the interface method can include steps executed by a processor according to a method as described above for providing the searched dataset and the information and receiving the identified molecular feature parameters for training a machine learning-based identification model.

[0051] In a further aspect, an interface method for providing a target synthesis specification indicating a target molecule having target technical application characteristics is presented. The interface includes: i) receiving, via an input unit, the target technical application characteristics; ii) interfacing, via an interface unit, with a processor that executes a method as described above for providing the target technical application characteristics and receiving the identified target synthesis specification; and iii) providing, via an output unit, the identified target synthesis specification. In one embodiment, the interface method can include steps executed by a processor according to a method as described above for providing the target technical application characteristics and receiving the identified target synthesis specification.

[0052] In a further aspect, an apparatus for generating a machine learning-based identification model is presented such that the trained identification model is adapted to identify the technical application characteristics of a molecule based on one or more molecular feature parameters, the molecular feature parameters being associated with a synthesis specification corresponding to the molecule and / or a molecular structure, the apparatus comprising: i) an input interface configured to provide a searched dataset comprising a) the technical application characteristics for a plurality of searched molecules and b) a plurality of molecular feature parameter values for a plurality of molecular feature parameters for each of the plurality of searched molecules, and to provide an unsearched dataset comprising a plurality of molecular feature parameter values for a plurality of molecular feature parameters for a plurality of unsearched molecules; ii) a processor configured to generate a trained identification model by parameterizing a machine learning-based identification model based on the searched dataset and the unsearched dataset such that the trained identification model is adapted to identify the technical application characteristics of a molecule based on one or more molecular feature parameters; and iii) an output interface configured to provide the generated trained identification model.

[0053] In a further aspect, there is provided an apparatus for identifying one or more molecular feature parameters for training a machine learning-based identification model such that the trained identification model is adapted to identify the technical application characteristics of a molecule based on the one or more molecular feature parameters, the molecular feature parameters indicating the characterized properties of the molecule and / or being derivable from one or more characterized properties of the molecule, the apparatus comprising: i) an input interface configured to provide a searched dataset comprising a) the technical application characteristics for a plurality of searched molecules and b) a plurality of molecular feature parameter values for a plurality of molecular feature parameters for each of the plurality of searched molecules, and to provide an unsearched dataset comprising a plurality of molecular feature parameter values for a plurality of molecular feature parameters for a plurality of unsearched molecules; ii) a processor configured to select one or more molecular feature parameters from the feature parameters of the searched dataset based on the searched dataset and the unsearched dataset; and iii) an output interface configured to provide the selected molecular feature parameters.

[0054] In a further aspect, an apparatus for specifying a target synthesis specification indicating a target molecule having target technical application characteristics is presented. The apparatus includes: i) an input interface configured to provide a digital representation of potential target molecules that provides target technical application characteristics and indicates or is associated with characteristic parameters of the potential target molecules; ii) a processor configured to utilize a trained identification model for identifying the technical application characteristics of molecules, where the trained identification model is parameterized based on a searched dataset and an unsearched dataset such that the trained identification model is adapted to identify the technical application characteristics of molecules based on one or more molecular characteristic parameters of the molecules. The searched dataset includes: a) the technical application characteristics for a plurality of searched molecules; and b) a plurality of molecular characteristic parameter values for a plurality of molecular characteristic parameters associated with the synthesis specifications and / or molecular structures corresponding to each of the plurality of searched molecules. The unsearched dataset includes a plurality of molecular characteristic parameter values for a plurality of molecular characteristic parameters associated with the synthesis specifications and / or molecular structures corresponding to a plurality of unsearched molecules. The processor is configured to compare the identified technical application characteristics of the potential target molecules with the target technical application characteristics and, based on the comparison, perform either a) identifying the potential target molecule as the target molecule, or b) providing a new potential target molecule and repeating the identification of the technical application characteristics using the new potential target molecule; iii) an output interface configured to provide the identified target molecule and the corresponding identified target technical application characteristics.

[0055] In a further aspect, an interface device for providing a specific model is presented, the interface device comprising: i) an input unit configured to receive a searched dataset comprising a) technical application characteristics for a plurality of searched molecules and b) a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with a synthesis specification and / or a molecular structure corresponding to each of the plurality of searched molecules, and to receive information indicating an unsearched dataset comprising a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with a synthesis specification and / or a molecular structure corresponding to a plurality of unsearched molecules; ii) an interface unit configured to interface with a processor that provides the searched dataset and the information and executes the method as described above for receiving a trained specific model; and iii) an output unit configured to provide the trained specific model. In one embodiment, the interface device can comprise a processor that executes the method as described above for providing the searched dataset and the information.

[0056] In a further aspect, an interface device for providing one or more molecular feature parameters for training a machine learning-based identification model is presented. The interface device includes: i) an input unit configured to receive a searched dataset including information showing a searched dataset including: a) technical application characteristics for a plurality of searched molecules; b) a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with the synthesis specifications and / or molecular structures corresponding to each of the plurality of searched molecules, and a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with the synthesis specifications and / or molecular structures corresponding to a plurality of unexplored molecules; ii) an interface unit configured to provide the searched dataset and the information and to interface with a processor that executes the method as described above for receiving identified molecular feature parameters for training a machine learning-based identification model; and iii) an output unit configured to provide the identified molecular feature parameters. In one embodiment, the interface device can include a processor that provides the searched dataset and the information and executes the method as described above for receiving identified molecular feature parameters for training a machine learning-based identification model.

[0057] In a further aspect, an interface device for providing a target molecule having target technical application characteristics is presented. The interface device includes: i) an input unit configured to receive the target technical application characteristics; ii) an interface unit configured to provide the target technical application characteristics and to interface with a processor that executes the method as described above for receiving a specified target synthesis specification; and iii) an output unit configured to provide the specified target molecule. In one embodiment, the interface device can include a processor that provides the target technical application characteristics and executes the method as described above for receiving a specified target synthesis specification.

[0058] In a further aspect, the use of specified training molecular feature parameters for training a machine learning-based identification model is presented, and the training molecular feature parameters are generated using the apparatus and method as described above.

[0059] In a further aspect, a computer program product for identifying a target molecule having target technical application characteristics is presented, and the computer program product comprises program code means for causing a computing system to execute the method as described above.

[0060] In a further aspect, a control signal generated using the method as described above is presented.

[0061] In a further aspect, the use of a control signal generated using the method as described above for controlling a production system for producing a target molecule, in particular laboratory equipment.

[0062] In a further aspect, a trained identification model generated using the method as described above is presented.

[0063] In a further aspect, a data set comprising feature parameters is presented, and the feature parameters are identified using the method as described above.

[0064] In a further aspect of the present invention, there is provided an apparatus for optimizing a training dataset for training a machine learning-based identification model such that the trained identification model is adapted to identify the technical application characteristics of molecules based on one or more molecular physicochemical parameters, the molecular physicochemical parameters indicating the physicochemical characteristics of the molecules and / or being derivable from one or more physicochemical characteristics of the molecules, the apparatus comprising: i) a searched data providing unit for providing a searched dataset comprising a) characteristics of a plurality of molecules and b) a plurality of molecular physicochemical parameter values for a plurality of molecular physicochemical parameters for the plurality of molecules; ii) a parameter training space providing unit for providing a parameter training space, the parameter training space being defined in a molecular physicochemical parameter space with respect to the intended application of the identification model; iii) a subspace identifying unit for identifying a subspace of the parameter training space defined by one or more of the molecular physicochemical parameters defining the parameter training space, the subspace being identified such that the training dataset covers the subspace according to a predetermined criterion; iv) a training molecule identifying unit for identifying training molecular physicochemical parameters for training the identification model based on the identified molecular physicochemical parameters defining the identified subspace; and v) a training data selection unit for selecting data from the searched dataset such that the training dataset comprises characteristics of a plurality of molecules and molecular physicochemical parameter values of the identified training molecular physicochemical parameters for the plurality of molecules. In particular, it is preferable that the parameter training space providing unit is adapted to identify the parameter training space based on an unexplored dataset as defined above. Preferably, the predetermined criterion refers to the similarity between the searched dataset and the unexplored dataset, and the subspace is identified by optimizing a similarity measure between the searched dataset and the unexplored molecular dataset with respect to the molecular physicochemical parameters. Generally, for this aspect as well, the same embodiments and definitions as already described above can be applied.

[0065] In a preferred embodiment, the molecule is a polymer. In this embodiment, a method executed by a computer for specifying a target synthesis specification indicating a target polymer having target technical application characteristics is presented. The method includes: i) providing the target technical application characteristics; ii) providing a digital representation of a potential target synthesis specification indicating or associated with the characteristic parameters of the potential target polymer; iii) using a trained identification model for identifying the technical application characteristics of the potential target polymer, where the trained identification model is adapted to identify the technical application characteristics of the polymer based on one or more polymer characteristic parameters of the polymer, and the trained identification model is parameterized based on an explored dataset and an unexplored dataset. The explored dataset includes: a) the technical application characteristics of a plurality of explored polymers; and b) a plurality of polymer characteristic parameter values for a plurality of polymer characteristic parameters associated with the synthesis specification corresponding to each of the plurality of explored polymers. The unexplored dataset includes a plurality of polymer characteristic parameter values for a plurality of polymer characteristic parameters associated with the synthesis specification corresponding to a plurality of unexplored polymers; iv) comparing the identified technical application characteristics of the potential target polymer with the target technical application characteristics, and based on the comparison, either a) identifying the potential target polymer as the target polymer and identifying the potential target synthesis specification as the target synthesis specification, or b) providing a new potential target synthesis specification for a new potential target polymer and repeating the identification of the technical application characteristics using the new potential target synthesis specification for the new potential target polymer; and v) providing the identified target polymer and the target synthesis specification. Preferably, providing the identified target polymer and the target synthesis specification includes generating a control signal based on the target synthesis specification and providing the control signal, where the control signal is configured to control a production system for producing the target polymer according to the target synthesis specification.Moreover, a method executed by a computer for generating a machine learning-based identification model is presented such that the trained identification model is adapted to identify the technical application characteristics of a polymer based on one or more polymer characteristic parameters, the polymer characteristic parameters being associated with a synthesis specification corresponding to the polymer, the method comprising: i) providing a searched dataset comprising: a) technical application characteristics for a plurality of searched polymers; and b) a plurality of polymer characteristic parameter values for a plurality of polymer characteristic parameters associated with a synthesis specification corresponding to each of the plurality of searched polymers; ii) providing an unsearched dataset comprising a plurality of polymer characteristic parameter values for a plurality of polymer characteristic parameters associated with a synthesis specification corresponding to a plurality of unsearched polymers; iii) generating a trained identification model by parameterizing a machine learning-based identification model based on the searched dataset and the unsearched dataset such that the trained identification model is adapted to identify the technical application characteristics of a polymer based on one or more polymer characteristic parameters of the polymer; and iv) providing the generated trained identification model.Furthermore, a method executed by a computer for identifying one or more polymer characteristic parameters for training a machine learning-based identification model is presented such that the trained identification model is adapted to identify the technical application characteristics of a polymer based on the one or more polymer characteristic parameters. The polymer characteristic parameters are associated with a synthesis specification corresponding to the polymer. The method includes: i) providing a searched dataset comprising: a) technical application characteristics for a plurality of searched polymers; and b) a plurality of polymer characteristic parameter values for a plurality of polymer characteristic parameters associated with a synthesis specification corresponding to each of the plurality of searched polymers; ii) providing an unsearched dataset comprising a plurality of polymer characteristic parameter values for a plurality of polymer characteristic parameters associated with a synthesis specification corresponding to a plurality of unsearched polymers; iii) selecting, based on the searched dataset and the unsearched dataset, one or more polymer characteristic parameters from the characteristic parameters of the searched dataset; and iv) providing the selected polymer characteristic parameters. Additionally, an interface method for providing a target synthesis specification indicating a target polymer having target technical application characteristics is presented. The interface includes: i) receiving, via an input unit, the target technical application characteristics; ii) interfacing, via an interface unit, with a processor that executes the method according to any one of claims 1 to 8 for providing the target technical application characteristics and receiving the identified target synthesis specification; and ii) providing, via an output unit, the identified target synthesis specification.Also provided is an apparatus for specifying a target synthesis specification indicating a target polymer having target technical application characteristics, the apparatus comprising: i) an input interface configured to a) provide the target technical application characteristics and b) provide a digital representation of a potential target synthesis specification indicating or associated with the characteristic parameters of a potential target polymer; ii) a processor configured to a) utilize a trained identification model for identifying the technical application characteristics of a polymer, the trained identification model being adapted to identify the technical application characteristics of a polymer based on one or more polymer characteristic parameters of the polymer, the trained identification model being parameterized based on a searched data set and an unsearched data set, the searched data set comprising: I) the technical application characteristics for a plurality of searched polymers and II) a plurality of polymer characteristic parameter values for a plurality of polymer characteristic parameters associated with the synthesis specification corresponding to each of the plurality of searched polymers, the unsearched data set comprising a plurality of polymer characteristic parameter values for a plurality of polymer characteristic parameters associated with the synthesis specification corresponding to a plurality of unsearched polymers, and b) compare the identified technical application characteristics of the potential target polymer with the target technical application characteristics and, based on the comparison, either I) identify the potential target polymer as the target polymer and identify the potential target synthesis specification as the target synthesis specification or II) provide a new potential target synthesis specification for the potential target polymer and repeat the identification of the technical application characteristics using the new potential target synthesis specification for the potential target polymer; and iii) an output interface configured to provide the identified target polymer and the corresponding identified target technical application characteristics. Further provided is a computer program product for specifying a target synthesis specification indicating a target polymer having target technical application characteristics, the computer program product comprising program code means for causing a computing system to execute the method as described above. Additionally, provided is the use of a control signal generated using the method as described above for controlling a production system, particularly laboratory equipment, for producing a target polymer according to the target synthesis specification.The methods, apparatuses, and computer program products described above are to be understood to have similar and / or identical preferred embodiments, particularly as defined in the dependent claims.

[0066] It should be understood that the preferred embodiments of the present invention can also be any combination of the dependent claims or of each of the above-described embodiments with the respective independent claims.

[0067] These and other aspects of the present invention will become apparent with reference to the embodiments described hereinafter.

Brief Description of the Drawings

[0068]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Modes for Carrying Out the Invention

[0069] Detailed Description of the Drawings FIG. 1 schematically and exemplarily shows a method executed by a computer for specifying a target synthesis specification that indicates a target molecule having target technical application characteristics. The method includes providing target technical application characteristics that can refer to any technical application characteristics to be satisfied by each target molecule. For example, the target technical application characteristics can be provided in the form of a target value with a predetermined limit that the target molecule should satisfy. However, the target technical application characteristics can also be provided in the form of a target value range, in which case the target molecule should include technical application characteristic values that fall within the target range. Moreover, the target technical application characteristics can also be provided in the form of an upper limit or a lower limit such that the target molecule should include technical application characteristic values that are below or above each respective upper or lower limit.

[0070] Furthermore, the method includes providing a digital representation of a potential target molecule that indicates or is associated with characteristic parameters of the potential target molecule. The digital representation can refer to any representation that enables derivation of the potential target molecule from the digital representation, and in particular, the digital representation can indicate a chemical structure, i.e., a molecular structure, or a synthesis specification of the molecule. The characteristic parameters of the potential target molecule can in particular refer to any kind of parameters that enable characterization of the properties of the potential target molecule. For example, the characteristic parameters can comprise process parameters that characterize at least part of a process by which the potential target molecule can be manufactured. Furthermore, the characteristic parameters can also comprise recipe parameters that characterize a recipe by which the potential target molecule can be manufactured, for example, characterizing substances used during the manufacturing process such as starting materials or catalysts. Preferably, the characteristic parameters comprise physicochemical parameters of the potential target molecule itself, for example, molecular descriptors such as quantum mechanical descriptors. Preferably, the characteristic parameters indicate molecular connectivity or other properties of the molecule. The digital representation of the potential target molecule indicates or is associated with the characteristic parameters of the potential target molecule. For example, the characteristic parameters of the potential target molecule can be derived from the molecular structure and / or synthesis specification provided as the digital representation. However, if the potential target molecule is known, at least some of the characteristic parameters can already be part of the digital representation of the potential target molecule or can be available on respective databases. If necessary, well-known methods for deriving characteristic parameters from the target molecule can be utilized, for example, based on the molecular structure and / or synthesis specification, such as 3D simulations, quantum mechanical simulations, electronic structure calculations, etc. However, for example, in the case of characteristic parameters that generally refer to molecular connectivity or the number of atoms provided as part of molecule identification, the characteristic parameters can also be directly read from the potential target molecule. Moreover, suitable characteristic parameters can already be stored for a plurality of potential target molecules and can then be provided by accessing each storage and selecting the characteristic parameters associated with a particular potential molecule.

[0071] In the next step, the specific model is utilized to identify the technical application characteristics of potential target molecules, and the trained specific model is parameterized based on the explored dataset and the unexplored dataset such that the trained specific model is adapted to identify the technical application characteristics of a molecule based on one or more molecular feature parameters of the molecule. Further details referring to the training, and in particular the parameterization of the specific model, are described with respect to FIG. 2. Generally, a specific model that is suitable for the intended use, i.e., trained to identify technical application characteristics with reference to the target technical application characteristics based on the derived or provided feature parameters and is generally applicable to potential target molecules, is utilized. Thus, generally, different specific models can be stored, for example, on a storage unit for different molecular types, different technical application characteristics to be specified, and / or different feature parameters and can then be utilized to identify the technical application characteristics. The specific model is a data-driven model that is parameterized to be able to identify the technical application characteristics associated with a molecule based on a digital representation, for example, based on molecular descriptors representing the physicochemical characteristics of a subgroup. In a preferred embodiment, the data-driven model refers to a machine learning model, for example, an algorithm based on a regression model or an algorithm based on a classifier model. The algorithm based on a regression model can be based on any of a neural network algorithm, a LASSO algorithm, a ridge regression algorithm, a MARS algorithm, and a random forest algorithm. The model algorithm based on a classifier can be based on either a random forest algorithm or an SVM algorithm. The inventors have found that for most applications, in particular, the random forest-based algorithm and the MARS-based algorithm are suitable. Based on the feature parameters and the specific model associated with or indicated by the potential target polymer, the respective technical application characteristics of the potential target molecule are identified.

[0072] In a further step, the identified technical application characteristics and the target technical application characteristics are compared. If the identified technical application characteristics of the potential target molecule do not meet the target technical application characteristics, new potential target molecules are provided, in particular those with a new molecular structure and / or synthesis specification. For example, the new potential target molecule can be provided from a storage in which a plurality of potential target molecules are already stored. The new potential target molecule can then be optionally selected from the plurality of already stored potential target molecules or based on a predetermined rule. However, the new potential target molecule can also be generated, in particular in the form of a new potential target synthesis specification and / or molecular structure, for example by modifying one or more of the parameters of the previously used potential target synthesis specification and / or molecular structure based on a predetermined rule. The identification of the technical application characteristics using the specific model is then repeated with the new potential target molecule. If the identified technical application characteristics meet the target technical application characteristics within a predetermined limit, the potential target molecule is identified as a target molecule, and such an identified target molecule is provided, for example, in the form of an executable control file that can be used to control a synthesis process, for example, in a laboratory, to synthesize the target molecule.

[0073] FIG. 2 schematically and illustratively shows a method executed by a computer for providing a specific model that can be used in the method executed by the computer described above. In particular, the provided specific model is parameterized based on a searched data set and an unsearched data set. The method includes providing a searched data set comprising: a) technical application characteristics for a plurality of searched molecules, and b) a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with the synthesis specifications and / or molecular structures corresponding to each of the plurality of searched molecules. Further, the unsearched data set is provided with a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with the synthesis specifications and / or molecular structures corresponding to a plurality of unsearched molecules. In particular, at least some of the molecular feature parameters for which the feature parameter values are provided in the unsearched data set and the searched data set refer to the same molecular feature parameters. The provided searched data set and unsearched data set are then used to generate a trained specific model by parameterizing a machine learning-based specific model based on the searched data set and the unsearched data set such that the trained specific model is adapted to identify the technical application characteristics of a molecule based on one or more molecular feature parameters of the molecule. Generally, parameterizing a specific model based on a searched data set and an unsearched data set not only enables a more effective training process for the specific model, i.e., one that consumes fewer computational resources and less time, but also enables better adaptation of the trained specific model to the intended application scenario, and thus enables more accurate results when using a specific model trained to identify the technical application characteristics of a molecule, as has been found by the inventors.

[0074] In particular, the explored dataset and the unexplored dataset are preferably used to identify a training dataset for training a specific model that is particularly adapted to the application situation in which the specific model is to be used. FIG. 3 illustratively shows in a schematic diagram the background regarding the identification of a training dataset based on the unexplored dataset and the explored dataset. In particular, in FIG. 3, two different cross-sections of the space spanned by all the feature parameters provided by the explored and unexplored datasets, i.e., subspaces, are shown. In this example of the schematic diagram, a cross-section 310 of this space spanned by feature parameter 1 and feature parameter 2 is shown. Further, a cross-section 320 spanned by feature parameter 3 and feature parameter 4 is shown. Based on the values provided for each feature parameter within the unexplored dataset and the explored dataset, the members of the unexplored and explored datasets can be recorded in their respective cross-sections 310, 320. In particular, each point or cross mark shown in each of the cross-sections 310, 320 can be regarded as being associated with one molecule of the unexplored or explored dataset. In the example shown here, the cross mark refers to the explored dataset and the point refers to the unexplored dataset. Generally, the explored dataset refers to molecules for which at least one technical application characteristic is already known, e.g., measured, and thus defines a dataset that can generally be used to train a specific model due to the already known technical application characteristics. As illustratively shown in FIG. 3, in cross-section 310 based on feature parameters 1 and 2, the explored molecules, i.e., the members of the explored dataset, can be found only in a specific region 311 of cross-section 310, i.e., a subspace, for this combination of feature parameters. The unexplored dataset generally refers to molecules that are artificially generated in large quantities, particularly molecular structures and / or synthesis specifications. For example, more than 1000 unexplored molecules, particularly molecular structures and / or synthesis specifications, can be generated to form the unexplored dataset. The unexplored dataset is then used to define the technical application situation targeted by the specific model.For example, an unexplored dataset can be generated to comprise a molecular structure and / or synthesis specifications and related molecules in consideration of specific production constraints or application constraints. Thus, a training dataset for training a specific model preferably covers a space defined by members of the unexplored dataset according to a predetermined criterion. For the first cross-section 310, i.e., for the first subspace of the feature parameter space defined by the feature parameters, as exemplarily shown in FIG. 3, the members of the explored dataset and the members of the unexplored dataset do not mainly cover the same region, i.e., the explored dataset covers region 311 and the unexplored dataset covers region 312. Generally, the predetermined criterion specifies that in such a case, the coverage is not sufficient. In contrast, in cross-section 320, i.e., in the subspace of the feature parameter space spanned by feature parameters 3 and 4, it is shown that the area 322 defined by the unexplored dataset is also sufficiently covered by the members of the explored dataset. Thus, changing the perspective, i.e., selecting each subspace of the feature parameter space, makes it possible to find feature parameters, in particular combinations of feature parameters, for which the explored dataset is similar to the unexplored dataset and thus highly suitable for the intended technical use of the specific model. Thus, for the embodiment shown in FIG. 3, a training dataset can be selected that comprises at least two technical application characteristics for two of the explored molecules of the explored dataset and further the values of feature parameters 3 and 4 of each explored molecule.

[0075] According to the principle described above with respect to FIG. 3, the unexplored dataset and the explored dataset are utilized to select a training dataset for parameterizing a specific model. In particular, selecting refers to the selection of feature parameters utilized in the training dataset for training a specific model. Preferably, the feature parameters, i.e., each subset of the feature parameters, are selected based on a similarity measure indicating the similarity between members of the explored dataset and members of the unexplored dataset with respect to each subset of the feature parameters. One possibility for mathematically finding the optimal subset of feature parameters such that members of the explored dataset are as similar as possible to members of the unexplored dataset is to define an optimization function. For example, a possibility for the optimization function can be defined as

[0076] [Number]

[0077] as follows, where

[0078] [Number]

[0079] are molecular feature parameters from the unexplored (scr) and explored (tr) datasets, f refers to a specific feature parameter, i, j refer to each molecule of the unexplored and explored datasets respectively, the coefficient cf can be 0 or 1 depending on whether each feature parameter is selected, the sum over f is equal to the number K of feature parameters in the selected subset of feature parameters, D is a general distance function between any given data points in the space defined by the selected subset of feature parameters, and S is the number of data points in the unexplored dataset. Each function can then be optimized with respect to the selected subset of feature parameters such that members of the explored dataset are as similar as possible to members of the unexplored dataset.

[0080] Moreover, the optimization function can be modified to take into account further advantageous features. In particular, the subset can be selected by further optimizing the specific accuracy of a specific model with respect to a subset of the molecular feature parameters. For example, for this case, the optimization function is

[0081]

Number

[0082] defined as follows, where

[0083]

Number

[0084] refers to the technical application characteristics specified by a specific model trained by each respective training data set based on each selected subset,

[0085]

Number

[0086] refers to the technical application characteristics for each molecule, T is the number of data points in the explored data set, D* is a general distance measure between the technical application characteristics specified for the molecules of the explored data set by the specific model to be trained and the technical application characteristics of the molecules of the explored data set, and the parameter λ refers to the weighting of the similarity between the technical application characteristics specified by the specific model and the technical application characteristics of the explored data set. Moreover, the subset can be selected by further optimizing the applicability of the specific model in the unexplored data set. In this case, the optimization function is

[0087]

Number

[0088] can be defined as follows, where

[0089]

Number

[0090] refers to the technical application characteristics specified by a specific model to be trained,

[0091]

Number

[0092] refers to the technical application characteristics for each molecule, G is a function for evaluating the applicability of a specific model in an unexplored dataset, S is the number of data points in the unexplored dataset, and the parameter μ refers to the weighting of the scale of the applicability of a specific model in the unexplored dataset. However, other mathematical formulations may also be suitable as an optimization function or for deriving each subset of feature parameters for training a specific model.

[0093] Each selected subset of feature parameters then makes it possible to define a training dataset for parameterizing the specific model, and the training dataset comprises at least two technical application characteristics of explored molecules and the feature parameter values associated with at least two explored molecules that refer to the selected subset of feature parameters. Known machine learning and training techniques can then be utilized to parameterize the specific model based on the training dataset. Such a trained specific model can then be provided and utilized, for example, by the target molecule identification method as described above with respect to FIG. 1.

[0094] In the following, preferred embodiments of the methods and corresponding apparatuses described above will be explained in more detail. In particular, for example, as described above, the method can be used to identify molecules for target technical application characteristics belonging to at least the following groups: a) mechanical characteristics, such as adhesion, tensile strength, rigidity, hardness, shrinkage, elongation, tear, tear strength, resilience, compressibility, abrasion, flow, morphology, tactile characteristics, stress at break, elongation at break, particle size distribution, degree of filling, b) optical characteristics, such as hue, turbidity, opacity, brightness, reflection, appearance, absorbance, scattering, color density, cloud point, haze, optical density, spectrum, refractive index, c) physicochemical characteristics, such as density, viscosity, K value, molar weight, dispersity / molar mass distribution, particle size distribution, solubility, partition coefficient, interfacial characteristics, surface tension, dispersibility, storage stability, odor, separation, electrical conductivity, capacitance, surface area, flow time, vapor pressure, VOC, solids content, hygroscopicity, magnetism, miscibility, thixotropy, phase transition characteristics, glass transition temperature, corrosion inhibition, solvent separation, aggregation, self-heating, shock sensitivity, loss on drying, reaction angle, electrostatic charge, minimum film-forming temperature, charge density, drop weight, melt volume rate, fluidity, tear propagation resistance, seal strength, permeability, d) chemical characteristics, such as functional group count, atom type count, functional group density, atom type density, chemical resistance, reaction timing, release time, growth, hard / soft segment content, crystallinity, reaction temperature, reaction pressure, decomposition, thermal decomposition, photodegradation, acidity, pKa, pH, carbon footprint, production cost, waste formation, moisture / water content, flammability, burn rate, self-ignition, flash point, generation of flammable gas, reaction to fire, deflagration rate, residual monomer count, generation of by-products, degree of polymerization, salt content, temperature resistance, oxidation characteristics, reduction characteristics, reactivity, ash content, non-volatile matter content, stability, chelating ability, calorific value, saponification value, and e) biological characteristics, such as biodegradability, biological resistance, toxicity, in vivo conversion, ecotoxicity, sensitization, bacterial count, enzyme activity, distribution in the environment, in vivo accumulation.

[0095] Therefore, the method can be applied in a plurality of technical fields, such as agricultural molecules, coatings, dispersions, structural molecules, such as polymer foams for thermal and acoustic insulation, shoes, and automotive applications.

[0096] An exemplary embodiment of a method of using a specific model trained as described above to identify the technical application characteristics can be composed of the steps described below. In particular, such a method can be used to identify the technical application characteristics in a method as described with respect to FIG. 1, in one embodiment where the molecule is a polymer. A schematic and exemplary flow diagram of an exemplary embodiment of method 400 is shown in FIG. 4. In a first step 410, a digital representation of the polymer is provided. The digital representation can directly comprise polymer feature parameters, in particular polymer descriptors, in which case the steps up to step 450 shown in FIG. 4 can be omitted. However, in many cases, the polymer feature parameters first have to be identified based on the provided digital representation, in which case, for example, a synthesis specification for the synthesis of the polymer is referred to. Generally, the digital representation can comprise any one or more of the following information, or enable the derivation thereof: the amounts of monomer components; the amounts of non-monomer components such as initiators, fillers, additives; reaction conditions such as temperature, vessel, pressure, stirring speed; condition profiles such as temperature profiles, pH values, solvents; feed profiles; types of polymerization such as radical, cationic, anionic, polycondensation, polyaddition, polyether formation; post-treatment such as amounts of components, conditions, and temperature and feed profiles; types of post-treatment such as radical, cationic, anionic, polycondensation, polyaddition, polyether formation; chemical information regarding compounds such as mixtures, the connectivity of non-polymerizable pure compounds, the composition of polymerizable pure compounds based on subgroups, the connectivity of monomers associated with subgroups in polymerizable pure components; for block copolymers, information regarding the blocks in which each monomer and reactive prepolymer is incorporated; for structural / layered materials and compounds, information regarding the phases / layers in which each component is included. If such information is not directly provided by the digital representation, in any step 420, reactive components and subgroups can be derived from the digital representation, for example from a synthesis specification.

[0097] If the information provided indicates the presence of a mixture, in the next step, the mixture is decomposed into its pure components and each polymer component is treated as an input polymer. Additionally, the polymer components can also be converted to mol% if necessary.

[0098] In the next step 430, the polymerizable components can be converted into subgroups, such as repeating units, and the subgroups are identified as different types. For example, the polymerizable subgroups can be identified based on the connectivity information of non-polymerizable pure compounds, for example, by using SMARTS via a KNIME workflow. Also, the connectivity information of all the expected subgroups can be derived from the connectivity information of non-polymerizable pure compounds, for example, by using reaction SMARTS via a KNIME workflow as well.

[0099] After the subgroups and their types are identified, at step 440, the types of characteristic parameters to be utilized may be provided. However, the characteristic parameters can also be identified without first selecting the types of subgroups. To reduce the computational resources for this method, at step 441, it is preferably determined whether the subgroup characteristic parameters associated with each type of subgroup are already stored in the database, for example, it can be determined whether there are already entries of subgroups having the same connectivity information in the database. If so, for example, at step 444, each associated subgroup characteristic parameter can be directly downloaded. If the identified types of subgroups are not stored in the database, the subgroup characteristic parameters associated with each type of subgroup can be identified, for example, at step 442. For example, the 3D structure of each type of subgroup can be derived based on the connectivity information, and the automatic calculation of subgroup characteristic parameters can be initiated, for example, using a computer cluster, or an existing machine learning identification can be utilized as the subgroup characteristic parameters. Generally, when calculations for new subgroups are required, after the calculations are completed, it is preferable that the results are stored in the database at step 443. Optionally, additional subgroup characteristic parameters can be provided from the topological geometric analysis, quantum chemical calculations, molecular dynamics calculations, coarse-grained methods, finite element calculations, and kinetic simulations of the subgroups. In particular, polymer reaction techniques can be used to derive subgroup characteristic parameters that enable consideration of the fine structure of the polymer.

[0100] In step 431, the amount of the subgroup, i.e., the amount of each type of subgroup, is specified based on, for example, the synthesis specification provided for the polymer. For example, the amount can be specified by counting the amount of polymerizable groups per polymerizable component, optionally including the prepolymer. In this case, information regarding the polymerizable groups can be derived from the non-polymerizable components, and such specified amount is optionally added to the counting of the number of unpolymerized polymerizable groups of the subgroup regarding the polymerizable components based on the composition of the polymerizable components, and the resulting amount is specified. Further, it is preferable that the amount of polymerizable groups derived from the agents used for the post-treatment after polymerization is removed from the resulting amount. In step 432, such specified amount of the subgroup is provided and can be stored, for example, in a database. Before further processing the subgroup of the specified amount, subgroups that are completely represented by other subgroups can be removed. Moreover, subgroups having the same connectivity can be combined.

[0101] Optionally, a subgroup of the derived quantities can be used for further interpretation of the polymer composition. For example, the total number of polymerizable functional groups such as double bonds, amine groups, alcohol groups, thiol groups, carboxylic acid groups, isocyanate groups, epoxide groups, and forming functional groups such as amide groups, ester groups, thioester groups, urea groups, urethane groups, thiourethane groups, ether groups can be specified. Also, the total molar weighted number of polymerizable functional groups, the total mass weighted number of polymerizable functional groups, the total number of residual functional groups such as double bonds, amine groups, alcohol groups, thiol groups, carboxylic acid groups, isocyanate groups, epoxide groups, the total molar weighted number of residual functional groups, the total mass weighted number of residual functional groups, the sum of all residual functional groups, the ratio between functional groups after polymerization, the number of crosslinks in the polymer, optionally mass weighted, the molar fraction of crosslinks in the polymer, the average number of atoms per subgroup, optionally per weight, the average number of non-H atoms per subgroup, optionally per weight, the average number of bonds per subgroup, optionally per weight, the average number of bonds between non-H atoms per subgroup, optionally per weight, the average number of rotors per subgroup, optionally per weight, the average number of rotors between non-H atoms per subgroup, optionally per weight, the average number of rings per subgroup, optionally per weight, the average polar surface area per subgroup, optionally per weight, the average refractive index per subgroup, optionally per weight, the total number of blocks, the molar size of the first block, the molar size of the last block, the HLB value of the polymer, optionally using the area weighted HLB value, the HLB value of the block with the minimum HLB value, optionally using the area weighted HLB value, the HLB value of the block with the maximum HLB value, optionally using the area weighted HLB value, the HLB value of the first block, optionally using the area weighted HLB value, the HLB value of the last block, optionally using the area weighted HLB value, the mass of the first block, the mass of the last block, the area of the block with the minimum HLB value, the area of the block with the maximum HLB value, the difference in HLB values of the blocks, optionally using the area weighted HLB value, the hydrophilic area of the polymer, the lipophilic area of the polymer, the number of arms of ring-opening polymerization, or the arm length of ring-opening polymerization, can be specified.

[0102] In step 450, the amount and type of the identified subgroup, as well as the associated subgroup characteristic parameters, can be utilized to calculate the polymer characteristic parameters. For example, the polymer characteristic parameters can be specified by one or more of the molar weighted averages of the associated descriptors of the subgroup, such as arithmetic mean, harmonic mean, logarithmic mean, mass weighted average, such as arithmetic mean, harmonic mean, logarithmic mean, volume weighted average, such as arithmetic mean, harmonic mean, logarithmic mean, surface area weighted average, such as arithmetic mean, harmonic mean, logarithmic mean. Moreover, the polymer characteristic parameters can be identified by identifying one or more of the molar weighted standard deviation, mass weighted standard deviation, volume weighted standard deviation, surface area weighted standard deviation, molar weighted maximum value, mass weighted maximum value, volume weighted maximum value, surface area weighted maximum value, molar weighted minimum value, mass weighted minimum value, volume weighted minimum value, surface area weighted minimum value, molar weighted sum, mass weighted sum, volume weighted sum, surface area weighted sum, and maximum difference from the associated subgroup characteristic parameters.

[0103] In step 460, the derived or provided polymer characteristic parameters can then be provided to a trained specific model to identify application characteristics. Generally, different trained specific models can also be utilized, as already explained above, for identifying different polymer characteristics. The specific model is trained using a training dataset identified based on an unexplored dataset and an explored dataset, as detailed above with respect to FIGS. 2 and 3, for example. The specific model can generally refer to sparse, e.g., spline, LASSO regression, PLS, and non-sparse, e.g., ridge regression, tree methods, kernel-based methods, statistical learning models for relating the polymer characteristic parameters to the application characteristics of interest. Moreover, the specific model can further provide a reliability estimate of the identified technical application characteristics, depending on the specific model used. In step 470, the identified technical application characteristics can then be provided to the user, e.g., via a user interface, or further utilized to be compared with the target technical application characteristics to identify a target polymer as a target molecule, as explained with respect to FIG. 1. In further applications, embodiments of the methods described above can be used for virtual screening, e.g., using one of configuration optimization, Pareto optimization, low-dimensional visualization of screened recipes, feature selection for the maximum applicability domain of the model, and an applicability domain checker for new recipes at the descriptor level.

[0104] Preferably, the molecular characteristic parameters used in the methods described above, the polymer characteristic parameters in the above example, are derived from quantum chemical calculations involving solvation treatment. Quantum chemical calculations have extremely unfavorable scaling with respect to system size, and calculations for polymers or shorter monomer sequences are impractical. This barrier is resolved by the above methods, including preferably cutting the polymer into subgroups at non-polarized bonds. The resulting subgroups have a size similar to that of the monomers, and the characteristic parameters can be calculated using quantum chemical methods.

[0105] The following provides more detailed preferred examples of preferred embodiments of a method of using a specific model trained as described above. A schematic and illustrative flow diagram of an exemplary preferred embodiment of the method is provided by FIG. 5. In this exemplary embodiment, the method begins, for example, via a user interface, by requesting a target value for a target application. Moreover, in the next step, the optimization is initialized by providing potential target molecules, for example, by providing a synthesis specification and / or a molecular structure, i.e., an initial synthesis specification. Optionally, in this process, constraints on the chemical structure of the potential target molecules can be considered, for example, if the user provides such constraints. The constraints can refer to, for example, constraints in the production of the molecule, in the starting materials to be used for synthesizing the molecule, etc. Moreover, additional application conditions representing particularly further possible information can be required for the target molecule to be realized. Based on the above steps, an optimization for identifying the target molecule, i.e., the target synthesis specification and / or the molecular structure, can be initialized. In the first step of the optimization, the molecular feature parameter values can be derived from the provided initial potential target molecules, for example, as described in detail with respect to FIG. 4 for polymers. However, the derivation of the molecular feature parameters can also refer to accessing a storage in which the respective molecular feature parameters for each of the potential target molecules are already stored. Moreover, if the provided digital representation of the potential target molecule already comprises the molecular feature parameters, this step can also be omitted. Based on the required additional application conditions, respective specific models can be provided. Based on the provided specific models and the digital representation of the potential target molecules, values of the target technical application characteristics of the potential target molecules can be provided. In the next step, it is determined whether the specified performance value, i.e., the specified technical application characteristic, meets the target value within a predetermined limit. If not, i.e., if this condition is not met, the formulation of the potential target molecule is modified and a new target molecule is identified taking into account the previously provided constraints. Then, a new iteration can start anew for the new potential target molecule.At a certain point, when the specified performance value meets the target value within its limit, i.e., when each condition is realized, the potential target molecule is specified as the target molecule and provided to the user or the control unit, for example, to produce each specified target molecule.

[0106] Figure 6 schematically and illustratively shows a further preferred embodiment of the above method for identifying a target molecule using a predetermined first target technical application characteristic. In this embodiment, in addition to the first target application characteristic, it is desirable for the target molecule to also achieve a further second target value, i.e., a target technical application characteristic. The additional target technical application characteristic can refer to any technical application characteristic. In particular, the method follows the same principle as described above with respect to Figure 5. However, due to the additional target value, additional conditions must be satisfied during optimization. Therefore, only the main differences from the method described above will be pointed out below. In particular, in this preferred embodiment, the optimizer module performs the optimization not only over the first target value, i.e., over the first target application characteristic, but also over the second target value. Preferably, for the second target value as well, a specific model adapted to identify the value of the technical application characteristic based on the physicochemical parameters of the molecule is utilized. Therefore, in addition to the method described above for the second target application, a second specific model is provided that enables the identification of the application characteristic value based on the physicochemical parameters of the molecule for the second target application. The second specific model can be based on, for example, the same algorithm as the first specific model and is trained according to the same principle as described above only with respect to the identification of other technical application characteristics. The comparison then refers not only to identifying whether the identified first application characteristic satisfies the target application characteristic within the limits, but also to identifying whether the identified second application characteristic satisfies the second target application characteristic within the limits. A predetermined rule can be utilized to identify in which case to continue the iteration, i.e., a new potential target molecule is provided, and for those conditions, the potential target molecule is identified as the target molecule. For example, the user can pre-identify a weighting value for weighting how much each condition must be satisfied. For example, for the user, it may be more important for the first technical application characteristic to be satisfied, while the other target application characteristics are not as important. In this case, the limit for the second target application characteristic to be satisfied can be set wider, or the weighting for satisfying this condition can be reduced.Next, at repeated points, if a condition is met and it is identified that a predetermined rule is satisfied, each potential target molecule can be identified as a target molecule and provided as an output to the user, or can be used to generate a control file for producing each target molecule.

[0107] Figure 7 shows a block diagram of an exemplary system architecture of an automated laboratory system 1000 for synthesizing molecules, which system has a laboratory equipment control device 1102, a network 1150, and synthesis specifications, i.e., recipes, modules 1100 / 1110, and a client device 1108. The automated laboratory system includes a laboratory equipment control device layer 1152 as part of the laboratory equipment control device 1102, and a synthesis specification module layer 1154 associated with the synthesis specification module, and a remote control or client layer 1156 associated with the client device 1108. The laboratory equipment control device layer can be divided into several hierarchical layers, namely, a hardware layer, a middleware layer, and an interface layer. The hardware layer relates to hardware resources such as sensors and actuators, particularly for controlling the synthesis of molecules. The middleware relates to any of the known middleware for synthesis operations in a laboratory or a plant. An example thereof is LABS / QM, which provides various abstractions for hardware, network, and operating systems, such as low-level device control and message passing. The communication layer relates to a communication protocol, and the protocol is REST, which may be implemented over different transport protocols (i.e., UDP, TCP, telemetry), enabling the exchange of messages between the laboratory equipment control device and the laboratory equipment device. Such a software architecture enables the control and monitoring of laboratory equipment without interacting with the hardware.

[0108] The synthesis specification module layer 1154 may include a mass storage layer, a computing layer, and an interface layer. The mass storage layer is configured to provide mass storage for a data-driven identification model for providing a molecular synthesis specification based on technical application characteristics, as described in detail above. In particular, the functions implemented by the devices as described above may be provided as program code means stored in the mass storage. Further, synthesis specifications for a plurality of molecules may be stored in the mass storage. Such data may be stored in a structured database such as an SQL database, or a distributed file system such as HDFS, or a NoSQL database such as HBase or MongoDB. The computing layer may include an application layer that enables customizing the functions provided by a standard cloud service to execute a computing process based on target characteristics. Such functions may include identifying a digital representation of a target molecule based on the target technical application characteristics and the identification model, generating a synthesis specification from the digital representation of the target molecule, and providing the synthesis specification as control data, i.e., a control signal, to a laboratory equipment control device.

[0109] The interface layer may implement a web service, a network interface as UDP or TCP, or a web socket interface. For communication with the laboratory equipment control device, a REST API is implemented.

[0110] The client layer 1156 provides an interface for the end user. For the end user, the client layer 1156 can launch a client-side web application that provides an interface to the synthesis specification module layer 1154 or the laboratory equipment control device layer 1152. A UI for selecting the target technical application characteristics and the target value of this characteristic may be provided to the user, and the target value may also include a range of technical application characteristics. In other embodiments, a UI for selecting a plurality of technical application characteristics and their respective values may be provided to the user. The application may be configured such that the user can remotely monitor and control laboratory equipment and its operations. In other embodiments, the client device layer and the synthesis specification module layer may be integrated into one device. The alternative forms described herein are for illustrative purposes only and should not be construed as limiting.

[0111] FIG. 8 shows an exemplary system architecture block diagram of a system and apparatus for generating a specific model for identifying technical application characteristics, network 2150, a model generation module 2100 / 2110 that can be regarded as or includes a training model device, a synthesis specification module 1100 / 1110, and a client device 2108. The system for generating a specific model includes a model generation module layer 2154 as part of the model generation module and a client layer 2156 associated with the client device 2108.

[0112] The model generation module layer 2154 may include a mass storage layer, a computing layer, and an interface layer. The storage layer is configured to provide mass storage for the data-driven specific model as described above. Further, the mass storage is configured to store the synthesis specifications for the molecules and the technical application characteristics. Such data may be stored in a structured database such as an SQL database, or a distributed file system such as HDFS, or a NoSQL database such as HBase or MongoDB. The computing layer may include an application layer that customizes the functionality provided by standard cloud services to enable a computing process for generating a specific model for identifying the characteristics of molecules. Such functionality includes receiving, for at least two previously explored molecules, their respective digital representations associated with at least one technical application characteristic of each of the at least two previously explored molecules and measurement data, receiving the digital representation of at least one unexplored molecule in the model generation module, training the model according to the training principle described above based on a similarity measure between the digital representations of at least two previously explored molecules, at least one technical application characteristic for each of the at least two previously measured molecules, and preferably the digital representations associated with at least two previously explored molecules and the digital representations associated with at least one unmeasured molecule, and providing a specific model for the technical application characteristics via an output interface. The model generation module layer may be configured to deploy the generated model and the synthesis specification database to the synthesis specification module layer. This may include storing the generated model and the synthesis specification database in a mass storage device associated with the synthesis specification module.

[0113] The model generation module layer may further be configured to identify from the synthesis specification a digital representation of the molecule associated with the synthesis specification. The digital representation may include a set of molecular feature parameters and molecular feature parameter values associated with the synthesis specification of each explored molecule. One way to derive these molecular feature parameters may be to apply the SMILES algorithm or any other principle already described above. When a model is generated based on the digital representation derived from the synthesis specification, the relationship between the synthesis specification and the feature parameters may be stored in a mass storage device associated with the model generation module. In such a case, deploying the model includes providing that relationship.

[0114] The interface layer may implement a web service, a network interface as UDP or TCP, or a web socket interface. In this embodiment, a REST API is implemented for communication with the client device. The client layer 2156 provides access to a mass storage device, which contains synthesis specifications of molecules and includes at least one technical application characteristic for at least two molecules. The client layer further provides an interface to the end user. To the end user, the client layer 2156 may execute a client-side web application that provides an interface to the model generation module layer 2154 or to the mass storage device associated with the client layer. The user may be provided with a UI for selecting technical application characteristics. The user may further be provided with a UI for selection of molecular data and technical application characteristic data associated with the molecular data. The user interface may also provide an option to upload the selected data to the model generation module layer and optionally an option to initiate model generation.

[0115] Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, by reviewing the drawings, the disclosure, and the appended claims.

[0116] For the processes and methods disclosed in this specification, the operations performed in the processes and methods may be carried out in a different order. Further, the operations outlined are provided as examples only, and some of the operations are optional and may be combined into fewer steps and operations, supplemented with additional operations, or extended with additional operations without impairing the essence of the embodiments of the present disclosure.

[0117] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite articles "a" or "an" do not exclude a plurality.

[0118] A single unit or device may perform the functions of a plurality of items recited in the claims. The fact that certain means are recited in mutually different dependent claims does not indicate that a combination of these means cannot be used advantageously.

[0119] Procedures performed by one or more units or devices, such as providing digital representations and target technical application characteristics, identifying technical application characteristics, providing technical application characteristics, etc., may be executed by any other number of units or devices. These procedures can be implemented as program code means of a computer program and / or as dedicated hardware.

[0120] A computer program product may be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium, supplied together with or as part of other hardware, but may also be distributed in other forms, for example via the Internet or other wired or wireless telecommunications systems.

[0121] Any unit described herein may be a processing unit that is part of a classical computing system. The processing unit may include a general-purpose processor, or may include a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or any other dedicated circuit. Any of the memories may be physical system memory, and may be volatile, non-volatile, or a combination of both. The term "memory" may include computer-readable storage media such as non-volatile mass storage. If the computing system is distributed, the processing power and / or storage capacity may also be distributed. The computing system may include multiple structures as "executable components". The term "executable component" is a structure that may be software, hardware, or a combination thereof, and is well understood in the computing field. For example, when implemented in software, those skilled in the art will understand that the structure of an executable component may include software objects, routines, methods, etc. that can be executed on a computing system. This may include both executable components within the heap of a computing system and executable components on a computer-readable storage media. The structure of an executable component may be present on a computer-readable medium such that when interpreted by one or more processors of the computing system, e.g., by a processor thread, it causes the computing system to perform a function. Such a structure may be directly computer-readable by a processor, e.g., when the executable component is binary, or may be structured to be interpretable and / or compiled so as to generate such a binary that is directly interpretable by a processor, whether in one or multiple steps. In other examples, the structure may be hard-coded logic gates or hard-wired logic gates implemented exclusively or almost exclusively in hardware, e.g., within a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or other dedicated circuit.Accordingly, the term "executable component" refers to a structure well understood by those skilled in the art of computing, regardless of whether it is implemented in software, hardware, or a combination thereof. Embodiments herein are all described with reference to operations performed by one or more processing units of a computing system. When such operations are implemented in software, one or more processors direct the operation of the computing system in response to the execution of computer-executable instructions that make up the executable component. The computing system may also include a communication channel, for example via a network, that enables the computing system to communicate with other computing systems. A "network" is defined as one or more data links that enable the transmission of electronic data between computing systems and / or modules and / or other electronic devices. When information is transferred or provided to a computing system via a network or another communication connection, such as either hardwired, wireless, or a combination of hardwired and wireless, the computing system appropriately considers the connection to be a carrier medium. A carrier medium can be used to carry the desired program code means in the form of computer-executable instructions or data structures and can include a network and / or data link that can be accessed by a general-purpose computing system or a dedicated computing system or a combination thereof. Not all computing systems require a user interface, but in some embodiments, the computing system includes a user interface system for interfacing with a user. The user interface functions as an input or output mechanism to the user, for example via a display.

[0122] Those skilled in the art will understand that at least a part of the present invention may be implemented in a network computing environment having many types of computing system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable household appliances, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, pagers, routers, switches, data centers, wearable devices such as glasses, etc. The present invention may also be implemented, for example, in a distributed system environment where local and remote computing systems linked together via a network by either a hardwired data link, a wireless data link, or a combination of a hardwired data link and a wireless data link both execute tasks. In a distributed system environment, program modules may be located in both local and remote memory storage devices.

[0123] Those skilled in the art will also understand that at least a portion of the present invention may be implemented in a cloud computing environment. The cloud computing environment may be distributed, although this is not essential. When distributed, the cloud computing environment may be distributed internationally within an organization and / or may have components held across multiple organizations. In this specification and the following claims, "cloud computing" is defined as a model that enables on-demand network access to a shared pool of configurable computing resources, such as networks, servers, storage, applications, and services. The definition of "cloud computing" is not limited to any of the many other advantages that can be obtained when such a model is deployed. The computing systems in the drawings, as described, include various components or functional blocks that may implement the various embodiments disclosed herein. The various components or functional blocks may be implemented on a local computing system, or may include elements resident in the cloud or be implemented on a distributed computing system that implements aspects of cloud computing. The various components or functional blocks may be implemented as software, hardware, or a combination of software and hardware. The computing systems shown in the drawings may include more or fewer components than those shown in the drawings, and some of the components may be combined as circumstances permit.

[0124] No reference signs in the claims shall be construed as limiting the scope.

[0125] The present invention refers to a method for identifying a synthetic specification. A digital representation of a target property and a potential synthetic specification is provided. Then, a model is utilized to identify the properties of potential target molecules. The model is parameterized based on a searched dataset and an unsearched dataset. The searched dataset comprises properties for a plurality of searched molecules and molecular feature parameter values for the plurality of searched molecules. The unsearched dataset comprises feature parameter values for a plurality of unsearched molecules. The identified properties of the potential target molecules are compared to the target property. Based on the comparison, either i) the potential target molecule is identified as a target molecule, or ii) a new potential target synthetic specification is identified and the identification of the properties is repeated. The identified target molecule is then provided.

Claims

Claim 1 A method executed by a computer for specifying a synthesis specification and / or molecular structure of a target molecule, particularly a target polymer, having application characteristics of a target technology, comprising: providing application characteristics of the target technology; providing a digital representation of potential target molecules indicating or associated with characteristic parameters; utilizing a trained identification model for identifying application characteristics of the potential target molecules, the trained identification model being parameterized based on a searched dataset and an unsearched dataset such that the trained identification model is adapted to identify application characteristics of a molecule based on one or more molecular characteristic parameters of the molecule, the searched dataset comprising: a) application characteristics for a plurality of searched molecules; and b) a plurality of molecular characteristic parameter values for a plurality of molecular characteristic parameters associated with a synthesis specification and / or molecular structure corresponding to each of the plurality of searched molecules, the unsearched dataset comprising a plurality of molecular characteristic parameter values for a plurality of molecular characteristic parameters associated with a synthesis specification and / or molecular structure corresponding to a plurality of unsearched molecules; comparing the identified application characteristics of the potential target molecules with the application characteristics of the target technology, and based on the comparison, either i) identifying the potential target molecule as the target molecule, or ii) providing a new potential target molecule and repeating the identification of the application characteristics using the new potential target molecule; providing the identified target molecule; A method comprising the above steps. Claim 2 The method according to claim 1, wherein providing the identified target molecule comprises generating a control signal based on the target molecule and providing the control signal, the control signal being configured to control a production system for producing the target molecule. Claim 3 The method according to claim 1 or 2, wherein the parameterization of the machine learning-based identification model is based on identifying a similarity measure between members of the searched dataset and members of the unsearched dataset. Claim 4 The parameterization of the machine learning-based identification model includes selecting, based on the similarity measure, a subset of the molecular feature parameters of the explored data set as training molecular feature parameters. The parameterization of the machine learning-based identification model uses, from the explored data set, a training data set comprising: a) the technical application characteristics for at least two of the plurality of explored molecules; and b) the molecular feature parameter values for the training molecular feature parameters associated with the synthesis specifications and / or molecular structures corresponding to each of the at least two explored molecules. The method according to claim 3.

5. The subset of the molecular feature parameters is selected based on the optimization of the similarity measure for the subset of the molecular feature parameters between the members of the unexplored data set and the members of the explored data set. The method according to claim 4.

6. The similarity measure indicates the distance between the explored data set for the subset of the molecular feature parameters and the members of the unexplored data set. The method according to any one of claims 3 to 5.

7. The subset is selected by further optimizing over the identification accuracy of the identification model for the subset of the molecular feature parameters. The method according to claim 5 or 6.

8. The subset is selected by further optimizing over the applicability of the identification model in the unexplored data set. The method according to any one of claims 5 to 7.

9. A method executed by a computer for generating a machine learning-based identification model such that the trained identification model is adapted to identify the technical application characteristics of a molecule based on one or more molecular feature parameters, wherein the molecular feature parameters are associated with the synthesis specification and / or molecular structure of the molecule, - providing an explored data set, comprising providing an explored data set comprising: a) the technical application characteristics for a plurality of explored molecules; and b) the plurality of molecular feature parameter values for the plurality of molecular feature parameters associated with the synthesis specification and / or molecular structure corresponding to each of the plurality of explored molecules. - Providing an unexplored dataset comprising a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with synthesis specifications and / or molecular structures corresponding to a plurality of unexplored molecules; - Generating the trained specific model by parameterizing the machine learning-based specific model based on the explored dataset and the unexplored dataset such that the trained specific model is adapted to identify the technical application characteristics of a molecule based on one or more molecular feature parameters of the molecule; - Providing the generated trained specific model; A method comprising.

10. A method executed by a computer for identifying one or more molecular feature parameters for training a machine learning-based specific model such that the trained specific model is adapted to identify the technical application characteristics of a molecule based on one or more molecular feature parameters, wherein the molecular feature parameters are associated with a synthesis specification and / or molecular structure corresponding to the molecule, - Providing an explored dataset, comprising: a) technical application characteristics for a plurality of explored molecules; and b) a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with a synthesis specification corresponding to each of the plurality of explored molecules; - Providing an unexplored dataset comprising a plurality of molecular feature parameter values for a plurality of molecular feature parameters associated with synthesis specifications and / or molecular structures corresponding to a plurality of unexplored molecules; - Selecting one or more molecular feature parameters from the feature parameters of the explored dataset based on the explored dataset and the unexplored dataset; - Providing the selected molecular feature parameters; A method comprising.

11. An interface method for providing a target synthesis specification indicating a target molecule having target technical application characteristics, the interface comprising: - Receiving target technical application characteristics via an input unit; - Connecting, via an interface unit, the target technical application characteristics to a processor that executes the method according to any one of claims 1 to 8 for providing the target technical application characteristics and receiving the identified target synthesis specification; - providing the specified target synthesis specification via an output unit; A method comprising. **Claim 12** An apparatus for specifying a target synthesis specification indicating a target molecule having target technical application characteristics, - an input interface, i) providing target technical application characteristics, ii) an input interface configured to provide a digital representation of a potential target molecule that indicates or is associated with characteristic parameters; - a processor, i) using a trained identification model for identifying the technical application characteristics of the molecule, the trained identification model being parameterized based on a searched dataset and an unsearched dataset such that the trained identification model is adapted to identify the technical application characteristics of a molecule based on one or more molecular characteristic parameters of the molecule, the searched dataset comprising a) technical application characteristics for a plurality of searched molecules and b) a plurality of molecular characteristic parameter values for a plurality of molecular characteristic parameters associated with the synthesis specification and / or molecular structure corresponding to each of the plurality of searched molecules, the unsearched dataset comprising a plurality of molecular characteristic parameter values for a plurality of molecular characteristic parameters associated with the synthesis specification and / or molecular structure corresponding to a plurality of unsearched molecules; ii) comparing the identified technical application characteristics of the potential target molecule with the target technical application characteristics, and based on the comparison, either i) identifying the potential target molecule as the target molecule or ii) providing a new potential target molecule and repeating the identification of the technical application characteristics using the new potential target molecule; - an output interface configured to provide the identified target molecule and the corresponding identified target technical application characteristics; An apparatus comprising. **Claim 13** A computer program product for specifying a target synthesis specification indicating a target molecule having target technical application characteristics, the computer program product comprising program code means for causing a computing system to execute the method according to any one of claims 1 to 8. **Claim 14** A control signal generated using the method according to any one of claims 2 to 8. **Claim 15** Use of a control signal generated using the method according to any one of claims 2 to 8 for a production system for producing said target molecule, in particular for controlling laboratory equipment.