Method for generating surfactant properties

By employing data-driven property models to predict surfactant properties from molecular structures, the method addresses the challenges of costly and time-consuming experimental methods, enabling efficient and targeted surfactant selection and production.

WO2025131351A1PCT designated stage expired Publication Date: 2025-06-26BASF SE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/074989
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-09-06
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The existing methods for determining surfactant properties are costly, time-consuming, and lack standardization, making it challenging to reliably select surfactants for various industrial applications.

Method used

A computer-implemented method using data-driven property models to generate surfactant properties by mapping molecular structure specifications to numeric molecular representations and predicting properties such as critical micelle concentration (CMC) at different temperatures.

Benefits of technology

This approach enables more efficient and targeted surfactant selection and production by providing consistent and comparable surfactant properties, reducing experimental uncertainties and costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000021_0001
    Figure IMGF000021_0001
  • Figure IMGF000021_0002
    Figure IMGF000021_0002
  • Figure IMGF000021_0003
    Figure IMGF000021_0003
Patent Text Reader

Abstract

The disclosure relates to the field of using machine learning for producing surfactants with surfactant properties generated by machine learning. Disclosed are methods, apparatuses, computer elements, surfactants or systems for selecting and producing surfactants based on generated surfactant properties.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD FOR GENERATING SURFACTANT PROPERTIES

[0002] TECHNICAL FIELD

[0003] The disclosure relates to the field of using machine learning for producing surfactants with surfactant properties generated by machine learning. Disclosed are methods, apparatuses, computer elements, surfactants or systems for selecting and producing surfactants based on generated surfactant properties.

[0004] TECHNICAL BACKGROUND

[0005] Surface-active agents (surfactants) are used in multiple industrial applications such as detergents, de-emulsifiers, wetting agents, oil recovery enhancers, pour-point depressants, pharmaceutical formulations, and drug delivery. Due to the numerous applications of surfactants, knowledge about the values of specific properties is essential under different conditions. Experimental measurements are one way to access to accurate values. However, conducting experiments in laboratories are expensive and / or time-consuming and may involve uncertainties about influencing factor like experimental setup, impurities, possible decompositions, etc. Moreover, no common standard is available in the market and measured properties may differ hampering reliable selection of surfactants.

[0006] SUMMARY OF THE INVENTION

[0007] In one aspect disclosed is a method, in particular a computer-implemented method, for selecting a surfactant based on at least one surfactant property, the method comprising the steps of: providing one or more molecular structure specification(s) associated with one or more candidate molecular structure(s) of one or more surfactant(s) or associated with one or more surfactant segment(s) for determining one or more candidate molecular structure(s) of one or more surfactant(s), optionally wherein the one or more candidate molecular structure(s) of one or more surfactant(s) are determined based on one or more surfactant segment(s); mapping the one or more molecular structure specification(s) to corresponding numeric molecular representations) associated with the atoms and bonds of the molecular structure; providing at least one data-driven property model trained on numeric molecular representation(s) and related property data associated with at least one surfactant property; generating at least one surfactant property by providing the numeric molecular representation(s) to the at least one data-driven property model; providing the at least one generated surfactant property for selecting a surfactant.

[0008] In another aspect disclosed is an apparatus for selecting a surfactant based on at least one surfactant property, the apparatus comprising: an input interface configured to provide one or more molecular structure specification (s) associated with one or more candidate molecular structure(s) of one or more surfactant(s) or associated with one or more surfactant segment(s) for determining one or more candidate molecular structure(s) of one or more surfactants), optionally wherein the one or more candidate molecular structure(s) of one or more surfactant(s) are determined based on one or more surfactant segment(s); a mapping interface configured to map the one or more molecular structure specification(s) to corresponding numeric molecular representation(s) associated with the atoms and bonds of the molecular structure; a model providing interface configured to provide at least one data-driven property model trained on numeric molecular representation(s) and related property data associated with at least one surfactant property; a generator configured to generate at least one surfactant property by providing the numeric molecular representations) to the at least one data-driven property model; an output interface configured to provide the at least one generated surfactant property for selecting a surfactant.

[0009] In one aspect disclosed is a method, in particular a computer-implemented method, for generating at least one surfactant property, the method comprising the steps of: providing one or more molecular structure specification(s) associated with one or more candidate molecular structure(s) of one or more surfactant(s) or associated with one or more surfactant segment(s) for determining one or more candidate molecular structure(s) of one or more surfactant(s), optionally wherein the one or more candidate molecular structure(s) of one or more surfactant(s) are determined based on one or more surfactant segment(s); mapping the one or more molecular structure specification (s) to corresponding numeric molecular representations) associated with the atoms and bonds of the molecular structure; providing at least one data-driven property model trained on numeric molecular representation(s) and related property data associated with at least one surfactant property; generating at least one surfactant property by providing the numeric molecular representation(s) to the at least one data-driven property model; providing the at least one generated surfactant property.

[0010] In another aspect disclosed is an apparatus for generating at least one surfactant property, the apparatus comprising: an input interface configured to provide one or more molecular structure specification(s) associated with one or more candidate molecular structure(s) of one or more surfactant(s) or associated with one or more surfactant segment(s) for determining one or more candidate molecular structure(s) of one or more surfactants), optionally wherein the one or more candidate molecular structure(s) of one or more surfactant(s) are determined based on one or more surfactant segment(s); a mapping interface configured to map the one or more molecular structure specification(s) to corresponding numeric molecular representation(s) associated with the atoms and bonds of the molecular structure; a model providing interface configured to provide at least one data-driven property model trained on numeric molecular representation(s) and related property data associated with at least one surfactant property; a generator configured to generate at least one surfactant property by providing the numeric molecular representation(s) to the at least one data-driven property model; an output interface configured to provide the at least one generated surfactant property.

[0011] In another aspect disclosed is a method, in particular a computer-implemented method, for selecting a surfactant based on at least two surfactant properties, the method comprising the steps of: providing one or more molecular structure specification(s) associated with one or more candidate molecular structure(s) of one or more surfactant(s) or associated with one or more surfactant segment(s) for determining one or more candidate molecular structure(s) of one or more surfactant(s); mapping the one or more molecular structure specification(s) to corresponding numeric molecular representations) associated with the atoms and bonds of the molecular structure; providing at least one data-driven property model trained on numeric molecular representation(s) and related property data associated with at least two surfactant properties; generating at least two surfactant properties by providing the numeric molecular representation(s) to the at least one data-driven property model; providing the at least two generated surfactant properties for selecting a surfactant.

[0012] In another aspect disclosed is a method, in particular a computer-implemented method, for selecting a surfactant based on at least one surfactant property, the method comprising the steps of: providing one or more molecular structure specification(s) associated with one or more candidate molecular structure(s) of one or more surfactant(s) or associated with one or more surfactant segment(s) for determining one or more candidate molecular structure(s) of one or more surfactant(s); mapping the one or more molecular structure specification(s) to corresponding numeric molecular representations) associated with the atoms and bonds of the molecular structure; providing at least one data-driven property model trained on numeric molecular representation(s) and related property data associated with at least one surfactant property, wherein at least one of the at least one surfactant properties is a critical micelle concentration at at least two different temperatures; generating at least one surfactant property by providing the numeric molecular representation(s) to the at least one data-driven property model; providing the at least one generated surfactant property for selecting a surfactant.

[0013] In another aspect disclosed is a method, in particular a computer-implemented method, for selecting a surfactant based on at least one surfactant property, the method comprising the steps of: providing one or more molecular structure specification(s) associated with one or more candidate molecular structure(s) of one or more surfactant(s) or associated with one or more surfactant segment(s) for determining one or more candidate molecular structure(s) of one or more surfactant(s), wherein the one or more surfactant(s) are at least two surfactants (e.g. a mixture comprising at least two surfactants); mapping the one or more molecular structure specification(s) to corresponding numeric molecular representations) associated with the atoms and bonds of the molecular structure; providing at least one data-driven property model trained on numeric molecular representation(s) and related property data associated with at least one surfactant property; generating at least one surfactant property by providing the numeric molecular representation(s) to the at least one data-driven property model; providing the at least one generated surfactant property for selecting a surfactant.

[0014] In another aspect disclosed is a use of the generated surfactant property according to the methods or by the apparatuses disclosed herein to produce a surfactant with surfactant structure selected based on a surfactant property generated.

[0015] In another aspect disclosed is a use of the generated surfactant property and associated surfactant structure selected based on the generated surfactant property according to the methods or by the apparatuses disclosed herein to produce a surfactant with the surfactant structure.

[0016] In another aspect disclosed is a method for producing a surfactant with surfactant structure selected based on a surfactant property generated according to the methods or by the apparatuses disclosed herein.

[0017] In another aspect disclosed is an apparatus configured to produce a surfactant with surfactant structure selected based on a surfactant property generated according to the methods or by the apparatuses disclosed herein.

[0018] In another aspect disclosed is a surfactant with surfactant structure selected based on a surfactant property generated according to the methods or by the apparatuses disclosed herein.

[0019] In another aspect disclosed is a surfactant with surfactant structure produced based on the generated surfactant property and associated surfactant structure selected based on the generated surfactant property according to the methods or by the apparatuses disclosed herein.

[0020] In another aspect disclosed is a surfactant with the surfactant structure produced based on the generated surfactant proper-ty and associated surfactant structure generated and selected according to the methods disclosed herein.

[0021] In another aspect disclosed is a use of the surfactant property and associated surfactant structure generated and selected according to the methods disclosed herein to produce a surfactant with the selected surfactant structure. In another aspect disclosed is a computer element, such as a (e.g. tangible and / or non-transitory) computer readable storage medium, a computer program or a computer program product, comprising instructions, which when executed by a computing node or a computing system, direct the computing node or computing system to perform the methods disclosed herein.

[0022] In another aspect disclosed is an apparatus comprising respective means for carrying out or performing the steps of the method disclosed herein or comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to carry out the steps of the method disclosed herein.

[0023] Any disclosure and embodiments described herein relate to the methods, the apparatuses, the surfactants, the uses and the computer elements. Advantageously, the benefits provided by any of the embodiments and examples equally apply to all other embodiments and examples.

[0024] EMBODIMENTS

[0025] In the following, embodiments of the present disclosure will be outlined by ways of embodiments and / or example. It is to be understood that the present disclosure is not limited to said embodiments and / or examples.

[0026] The methods, the methods, apparatuses, computer elements, surfactants or uses disclosed herein allow for more reliable and targeted surfactant selection and production. In particular, the surfactant may be selected based on surfactant properties as derived by a data-driven model. This is of high relevance in the area of surfactants due to their broad range of applications and diversity of properties required by such applications. In addition, standardized experimental setups are not existent for surfactants and comparability of different surfactants and related properties is challenging in the marketplace. To improve surfactant selection based on a common standard, the data-driven model provides a consistent and comparable result that speeds up surfactant usage in application. Lastly, the use of biosurfactants is increasing and the switch from synthetic surfactants to biosurfactants can be further spurred by the consistent transparency the approach provided herein can provide.

[0027] Surfactants or surface-active agents are organic compounds that contain hydrophilic groups e.g. as heads and hydrophobic groups e.g. as tails. The organic compounds are composed of two chemical parts with different polarities, e.g. a head group with affinity for polar phases, and a tail group that is attracted to nonpolar phases.

[0028] Surfactant properties may characterize physio-chemical properties and / or functional activities of the surfactant. Examples for characterizing surfactants include properties such as surface tension of water, interfacial tension for reference interface such as oil / water interface, foamability, foam rate, critical micelle concentration (CMC), the adsorption effectiveness given by the surface excess concentration (F m) , physio-chemical activity in dependence of temperature or contaminants, surface tension, foaming capacity, emulsification, stabilizing ability, solubility, detergency, activity with respect to a host reference solution or the like. The high diversity in molecular structure of different surfactants and specifically biosurfactants may result in functional activities or properties that are depended on the specific molecular structure. The surfactant properties may depend on the chemical or molecular structure of the surfactant, the application and / or a reference such as a host solution or reference experiment used to determine the properties. The surfactant properties may be derived from measurement based on one or more measurement techniques. Measurement techniques are known in the art and include optical, electrical, techniques such as tensiometry, refrac- tiv index, calorimetry, viscosity, chromatography, spectroscopy like IR spectroscopy and conductivity measurements.

[0029] Molecular structure specification (s) may relate to digital representation of the molecular structure relating to atoms and bonds forming the molecule. The molecular structure may specify atoms of the molecule, the atom positions, and the bonds among the atoms. For example, the type of atoms, the number of atoms, the atom numbers, the bond types, like single, double, aromatic or triple, the bond stereo like up, down, cis or trans, bond topology like chain or ring, or other structural characteristics may be specified.

[0030] The molecular structure specification (s) may be associated with one or more candidate molecular structure(s) of one or more surfactant(s). The molecular specifications may relate to surfactant type such as non-ionic, ionic, cationic, anionic, zwitterionic, or any combination of surfactant types. The molecular specifications may relate to surfactant function such as Water / Oil emulsifier, wetting agent, Oil / Water emulsifier, detergent, or solubilizer. The surfactant function may be provided by the hydrophile-lipophile balance (HLB) or other suitable specification.

[0031] The molecular structure specification (s) may be associated with one or more surfactant segment(s) for determining one or more candidate molecular structure(s) of one or more surfactant(s). The segments may relate to hydrophilic segments, hydrophobic segments and / or intermediate segments. The segments may be provided to find new surfactant structures with specific properties.

[0032] The numeric molecular representation (s) associated with the atoms and bonds of the molecular structure may include numeric representations of the molecular structure specification (s). The numeric representations may be generated based on feature vectors or feature representations in a numerical space. The numeric representations may be derived based on pre-defined feature vectors that may be referred to as vocabulary. The numeric representations may be derived from an embedding model configured to map the molecular structure specification(s) to a numerical space.

[0033] The data-driven property model may relate to data-driven models that are configured to map the numeric molecular representation to one or more properties. The data-driven property model may comprise at least one structural component that is configured to map numeric molecular representation (s) to specific molecular fingerprints or specific molecular representations depending on the specific chemical structure of the molecule and / or correlations between atoms of the molecule, such as first order correlations between nearest neighbor atoms, second order correlations between next nearest neighbor atoms or even higher. The data-driven property model may comprise at least one property component that is configured to map specific molecular fingerprints to properties. The data-driven may be trained based on an end-to-end training process training both components or in other words both components may be trained on the training data set in a single training process. The data driven model may be trained or configured to map numeric molecular representation to specific molecular representations to property of the specific molecular representation. The data-driven model may include a graph neural network, a transformer-based network, a recurrent neural network or any other suitable network architecture.

[0034] The data-driven model may be trained on numeric molecular representation(s) and related property data associated with at least one surfactant property. The training may include mapping molecular structure specifications to numeric representations that may than be used in relation to the property data as training data set. The property data may relate to one or more surfactant properties as measured for the surfactant associated with the molecular structure specification or the related numeric molecular representation. The property data may include measurement values and / or measurement conditions. Measurement conditions may relate to one or more measurement techniques, measurement settings and / or reference systems used for measurement of the property values. Measurement conditions may for example specify the temperature or pressure settings used during measurement, measurement techniques like tensiometry or calorimetry, reference systems like saline concentrations of test solutions the surfactant was contained in upon measurement, interfaces like water / oil, oil / water or the like the measurement was conducted at or surface contaminants used for measurement.

[0035] The surfactant property may relate to at least one application of the surfactant. The surfactant property may relate to an application performance of the surfactant based on the molecular structure of the surfactant. The application performance may relate to the surfactant property that may be measured under measurement conditions equivalent to the desired application of the surfactant. The application performance may relate to the surfactant property that may be measured in relation to a measurement system such as a reference solution or a reference interface equivalent to the reference solution or interface of the desired application of the surfactant.

[0036] In one embodiment the one of the at least two surfactant properties is a critical micelle concentration, in particular a critical micelle concentration at at least two different temperatures. For instance, at least one surfactant property may be a temperature dependent critical micelle concentration, e.g. comprising critical micelle concentration values for at least two temperature values, e.g. a number of critical micelle concentration values suitable for (accurate) interpolation between the values and extrapolation to obtain the critical micelle concentration at any given temperature relevant for the application of the surfactant or mixture of surfactants, e ,g. between 10°C and 90°C. For example, a temperature-dependent critical micelle concentration may comprise values of the critical micelle concentration at at least two of the following temperatures: 10°C, 20°C, 25°C, 30°C, 35°C, 40°C, 45°C, 55°C, 60°C, 65°C, 75°C, 85°C, 90°C or 95°C.

[0037] The data-driven property model may be configured with a specific element trainable to generate a functional relationship of a surfactant property, e.g. temperature-dependent surfactant property. For instance, in case the data-driven property model is or comprises a graph neural network or an ensemble of graph neural networks, at least one node of the graph neural network may be specific for the functional relationship, e.g. specific for temperature dependency of a surfactant property.

[0038] In one embodiment the one or more surfactant(s) are at least two surfactants (e.g. a mixture comprising the at least two surfactants). For example, at least two molecular structure specification (s) or molecular fingerprint may be used. In particular, more than one molecular fingerprint or molecular structure specification(s) may be generated. For binary mixtures, two molecular fingerprints or molecular structure specification (s) may be generated, for ternary mixtures three molecular fingerprints or molecular structure specification (s) may be generated etc. Afterwards, each molecular fingerprint may be multiplied with the molar fraction of each mixture component (from 0 to 1). As an example, at least two (two / three / four, etc.) molecular fingerprints or molecular structure specification (s) may be summed up to one for use with the data-driven property model. Alternatively or additionally, a new graph may be constructed for the data-driven property model, the new graph considering hydrogen bonding information between the at least two surfactants.

[0039] In one embodiment the surfactant property relates to at least one surfactant structure associated with a bio-based surfactant. The surfactant property may relate to at least one application of the bio-based surfactant. The surfactant property may relate to at least one application performance of the bio-based surfactant based on the molecular structure of the bio-based surfactant. Bio-based surfactants are often chemically more complex than many of the synthetically produced surfactants. Bio-based surfactants may include sugar-based and / or plat-based segments that may form the hydrophilic and / or hydrophobic segments of the bio-based surfactant. Bio-based surfactants are due to the higher structural complexity more expensive to produce. Bio-based surfactants and their properties are thus more difficult to tailor compared to synthetically produced, well-known surfactants. In view of the structural complexity, the more complex structural-property relationships and the resulting more difficult tailoring of bio-based surfactant properties, there is a need to simplify the tailoring of bio-based surfactants structurally with regard to their properties e.g. for applications like personal care or cosmetics.

[0040] In another embodiment the one or more molecular structure specification(s) relate to at least one surfactant type and / or at least one surfactant segment type. In particular the one or more molecular structure specification (s) may relate to at least one surfactant type and / or at least one surfactant segment type associated with hydrophilic and / or hydrophobic characteristics of the surfactant. For example, the surfactant type may relate to one or more surfactant classes related to the polarity of the surfactant structure, such as non-ionic, cationic, anionic, or zwitterionic. For example, the segment type may relate to hydrophilic or head segments, hydrophobic or tail segments and / or intermediate segments. One or more candidate molecular structure(s) of one or more surfactant(s) may be determined based on one or more surfactant segment(s). One or more candidate molecular structures may be generated randomly by mutation of one or more surfactant segment(s) per surfactant type. One or more candidate molecular structures may be generated based on a chemical rule set relating to possible bonds and / or reaction pathways for the one or more surfactant segment(s). One or more candidate molecular structures may be generated by machine learning models. Examples may include evolutionary or genetic processes or reinforcement learning processes where a recurrent neural network seeks to learn a policy that maximizes a reward as a function of chemical stability. In some implementations, an iterative search can be performed to identify candidate molecular structure(s) predicted to exhibit one or more target properties. For instance, an iterative search can propose several candidate molecule structures that can be evaluated by the data-driven property model(s). The distance of the resulting properties generated by the data- driven property model(s) per candidate may be evaluated and the one or more candidate structures with surfactant properties closest in distance may be selected.

[0041] In another embodiment the mapping of the one or more molecular structure specifications) to the corresponding or respective one or more numeric molecular representation(s) encodes at least the atom types and the bond types between the atoms. The numeric molecular representation may include a graph numerical representation with nodes encoding the atoms and / or edges encoding the bonds. Through the graph representation the chemical structure of the surfactant molecules can be encoded. Moreover, the graph numeric molecular representation(s) can be mapped by the at least one structural component of the data-driven model to specific molecular fingerprints or specific molecular representations depending on the specific chemical structure of the molecule and / or correlations between atoms of the molecule, such as first order correlations between nearest neighbor atoms, second order correlations between next nearest neighbor atoms or even higher. This way the numeric molecular representation can be refined in a data- driven manner to take account of the chemical specificities of different molecular structures.

[0042] In another embodiment one or more data-driven property model(s) are provided dependent on one or more surfactant type(s), property type(s) and / or measurement condition type(s). E.g. the data-driven property model may be trained to generate one or more surfactant properties per surfactant type and / or measurement condition type.

[0043] The surfactant type, property type and / or measurement condition type may be provided e.g. to select one or more data-driven property model (s) based on surfactant type, property type and / or measurement condition type. The data- driven model may be selected based on surfactant type, property type and / or measurement condition type. The data- driven model may be trained on training data relating to one or more surfactant type(s), property type(s) and / or measurement condition type(s). One or more data-driven property model(s) may be trained per surfactant type, property type and / or measurement condition type. The training data may be provided and split by the respective surfactant type, property type and / or measurement condition type. In some embodiments, transfer learning may be employed where at least one data-driven property model is trained on at least one property prediction task for example dependent on surfactant type(s), property type(s) and / or measurement condition type(s). The thus trained model may be employed as pre-trained model and may be used to train at least one other data-driven property model on at least one other property prediction task for example dependent on another set of surfactant type(s), property type(s) and / or measurement condition type(s). The trained data-driven property models may be stored in a model store in association with or in relation to metadata specifying surfactant type(s), property type(s) and / or measurement condition type(s). This way the model can be selected based on surfactant type, property type and / or measurement condition type. The surfactant type, property type and / or measurement condition type and related property data may be provided e.g. to the data-driven property model(s) for property generation. The property data specifying the surfactant type, property type and / or measurement condition type may be provided to the property component of the data-driven property model mapping the generated specific molecular representation to the one or more surfactant properties. The property data specifying the surfactant type, property type and / or measurement condition type may be mapped to a numeric representation that may be fused with the generated specific molecular representation.

[0044] In another embodiment at least one data-driven property model trained to map numeric molecular representation(s) to multiple property types or surfactant properties and / or multiple data-driven property models trained to map numeric molecular representation(s) to one or more property type(s) or surfactant properties are provided to generate one or more surfactant properties.

[0045] At least one data-driven property model trained to map numeric molecular representation (s) to multiple surfactant properties is provided. The data-driven may be trained based on an end-to-end process to perform multitasking with respect to the surfactant properties. The data-driven property model may include a multitask model configured to generate multiple different surfactant properties per numeric molecular representation. The data-driven property model may comprise at least one structural component configured to map numeric molecular representation to specific molecular fingerprint or specific molecular representation depending on the specific chemical structure of the molecule and / or correlations between atoms of the molecule, such as first order correlations between nearest neighbor atoms, second order correlations between next nearest neighbor atoms or even higher. The data-driven property model may comprise multiple property component(s) configured to map specific molecular fingerprint to one or more properties. Property component(s) may be configured to map from a single specific molecular fingerprint to one or more properties. Property component(s) may be configured to map from a single specific molecular fingerprint to one or more properties per surfactant type, property type and / or measurement condition type. The data-driven property model may comprise multiple property component(s) configured to map specific molecular fingerprint per property or property type. The data-driven property model may comprise at least one structural component e.g. per one or more surfactant type(s) and multiple property component(s) per one or more property type(s).

[0046] In another embodiment multiple data-driven property models trained to map numeric molecular representation(s) to one or more property type(s) or surfactant properties, preferably one or more surfactant properties of one type or a pre-defined set of surfactant properties, is provided. The data-driven models may be trained based on end-to-end processes to perform single tasking per model per one or more surfactant properties, preferably one or more surfactant properties of one type or a pre-defined set of surfactant properties. For single task generation, the surfactant properties to be generated may relate to a single property type. The surfactant properties to be generated may for example be related to each other or characterize related chemical structure or functional properties.

[0047] In another embodiment the at least one generated surfactant property is provided for selecting at least one surfactant and for providing the at least one generated surfactant property in relation to at least one other or another surfactant and associated surfactant properties. In particular, the at least one generated surfactant property may be provided for selecting at least one bio-based surfactant and for providing the at least one generated bio-based surfactant property in relation to at least one other or another bio-based and / or synthetically produced surfactant and associated surfactant properties. The methods disclosed herein may further include the steps selecting at least one surfactant and providing the at least one generated surfactant property in relation to another surfactant, such as a produced surfactant, and associated surfactant properties. The methods disclosed herein may further comprise the steps of selecting at least one bio-based surfactant and / or providing the at least one generated surfactant property in relation to another bio-based or a synthetically produced surfactant and associated surfactant properties.

[0048] In another embodiment property data associated with at least one surfactant property includes one or more measurement condition(s) that influence the one or more surfactant properties. The one or more measurement condition(s) may relate to measurement techniques, measurement settings, measurement reference systems. For example, the measurement conditions may include measurement setting like temperature or pressure, and / or surface contaminants. The data-driven model may be trained to include the one or more measurement condition(s). The one or more measurement condition(s) may be mapped to corresponding numeric representation(s). The one or more measurement condition(s) may be provided to the data-driven property model. The one or more measurement condition(s) may be combined or fused such as concatenated with numeric representations used by the data-driven model. One or more measurement condition(s) may be fused or combined with specific molecular fingerprints or specific molecular representations generated by at least one component of the data-driven property model.

[0049] In another embodiment one or more measurement condition(s) are provided in relation to one or more surfactant properties. The at least one data-driven property model may be trained on numeric molecular representation (s) and related one or more surfactant properties dependent on one or more measurement condition(s).

[0050] In another embodiment one or more measurement condition(s) are provided in relation to one or more surfactant properties, wherein the one or more surfactant properties are provided to the data-driven property model to generate at least one molecular fingerprint and the one or more measurement condition(s) are provided in relation to the mapping of the generated molecular fingerprint to the one or more surfactant properties.

[0051] In another embodiment the data-driven property model is configured to generate at least one molecular fingerprint, wherein data-driven property model is configured to generate at least one surfactant property, wherein the data- driven model is trained to map numeric molecular representation (s) to at least one surfactant property optionally dependent on one or more measurement condition(s).

[0052] In another embodiment one or more text input instruction(s) related at least one target property and / or at least one surfactant class are provided, wherein one or more molecular structure specification(s) and / or numeric structure representations) are generated from text input instructions related to surfactant properties. The text input instructions may be provided through a pre-defined selection of text elements or natural language text instructions. The at least one target property and / or at least one surfactant class may be derived from the text instruction(s). Based on the target property and / or the surfactant class one or more molecular structure specification (s) and / or numeric structure representation(s) may be generated. The one or more molecular structure specification(s) sand / or numeric structure representation(s) may be provided to one or more data-driven property model(s) as disclosed herein.

[0053] In another embodiment at least one target property and / or at least one surfactant class are provided, wherein multiple molecular structure specification (s) and / or numeric structure representation (s) are generated based on the at least one target property and / or at least one surfactant class, wherein one or more surfactant properties are generated per molecular structure specification and / or numeric structure representation, wherein at least one surfactant is selected based on the generated surfactant properties.

[0054] Further possible implementations or alternative solutions of the invention also encompass combinations - that are not explicitly mentioned herein - of features described above or below in regard to the embodiments. The person skilled in the art may also add individual or isolated aspects and features to the most basic form of this disclosure. Other features will become apparent from the following detailed description considered in conjunction with the accompanying drawings. It is to be understood, however, that the drawings are designed solely for purposes of illustration and not as a definition of the limits, for which reference should be made to the appended claims. It should be further understood that the drawings are not drawn to scale and that they are merely intended to conceptually illustrate the structures and procedures described herein.

[0055] BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In the following, the present disclosure is further described with reference to the enclosed figures:

[0057] Fig. 1 illustrates examples of surfactants.

[0058] Figs. 2-5 illustrate an example method for generating surfactant properties based on a graph molecular representation of a surfactant molecular structure generated by a graph neural network in connection with a mapping network configured to map the molecular representation to the surfactant property.

[0059] Fig. 6 illustrates an example of an apparatus for providing surfactant properties based on given surfactants and / or providing surfactants based on surfactant target properties.

[0060] Fig. 7 illustrates an example of an user interface, where a SMILES string is provided as one possible molecular structure specification, temperature as measurement condition and optionally a surfactant target property to be generated.

[0061] Fig. 8 illustrates an example of an user interface, where target properties are specified in string format for a surfactant type and the temperature is specified as measurement condition. Fig. 9 illustrates an example of a method for generating on one or more surfactant properties dependent on at least one measurement condition for selecting one or more surfactant(s).

[0062] Fig. 10 illustrates an example of a method or apparatus for selecting one or more surfactant(s) based on one or more surfactant properties dependent on at least one measurement condition.

[0063] Figs. 11-15 illustrate results of an example training of GNNs.

[0064] Figs. 16-19 illustrate results of another example training of GNNs.

[0065] Figs. 20a - 22b illustrate performance of an example GNN in particular for mixtures.

[0066] Fig. 23 illustrate a parity plot of an example GNN for surfactant mixtures.

[0067] DETAILED DESCRIPTION

[0068] The following embodiments are mere examples for implementing the method, the system, the apparatus or application device disclosed herein and shall not be considered limiting. The following description serves to deepen the understanding and shall be understood to complement and be read together with the description as provided in the above summary and embodiment sections of this specification. Some aspects may have a different terminology than e.g. provided in the description above. The skilled person will nevertheless understand that those terms refer to the same subject-matter, e.g. by being more specific.

[0069] Fig. 1 illustrates examples of surfactants. Surfactants or surface-active agents are organic compounds that contain hydrophilic groups e.g. as heads and hydrophobic groups e.g. as tails. The organic compounds are hence composed of two chemical parts with different polarities, e.g. a head group with affinity for polar phases, and a tail group that is attracted to nonpolar phases. These properties increase their tendency to generate self-assembled structures in solution leading to the formation of micelles with diameters ranging from nanometers to microns.

[0070] Due to their structural uniqueness, surfactants can be widely used. Applications include medicines, corrosion inhibitors for protecting steel and other corrosive metals, detergents, cosmetics, personal care formulations, de-emulsifi- ers, wetting agents, oil recovery enhancers, pour-point depressants, pharmaceutical formulations, or drug delivery formulations. For example, in cosmetics and personal care formulations, surfactants with at least one hydrophilic and a hydrophobic molecular moiety, may be used for lowering of the surface tension of the water, the wetting of the skin, the facilitation of soil removal and dissolution, easy rinse-off and - if desired - foam regulation. Surfactants are typically understood to mean surface-active substances which have an HLB value of greater than 20. Different types of surfactants include e.g. anionic, cationic, non-ionic or zwitterionic surfactants as illustrated in Fig. 1. Different surfactants exist per surfactant type and the chemical structure of the surfactant determines the surfactant properties desired for specific applications. For example, surfactants may be used as synthesis agent in the synthesis of nanoparticles or quantum dots with well-controlled geometries and / or properties to tune the surface / interface tensions between solid / liquid interfaces and for improving the dispersion stability. For example, surfactants may be used in home care or personal care products such as detergents, washing liquids or the like to tune the foamability, foam rate and / or foam stability.

[0071] The chemical structure of surfactants influences their properties targeted for specific applications. Examples of such properties include the maximum concentration of a surfactant at which micelles do not form or the minimum surfactant concentration at which such micelles are formed in a solution is called critical micelle concentration (CMC), the adsorption effectiveness given by the surface excess concentration (F m) as a measure of surfactant concentration at interfaces such as at air / water and oil / water interfaces, the foamability, the foam rate, the foam stability, the dispersion stability and may more. Moreover, the properties of surfactants are not standardized and depend e.g. on measurement conditions or the setup of the test used to measure the properties.

[0072] Due to the numerous applications of surfactants, knowledge about the values of specific properties such as CMC or F m under different conditions can guide surfactant selection. For example, surfactant concentrations larger than CMC, the solution is considered micellar and exhibits different behavior from a dilute solution (e.g., a solution with concentration less than the CMC). In certain situations, it is desirable for surfactants to have a low CMC, such as when they are used to dissolve hydrophobic drugs in micellar cores with minimal surfactant quantities. In applications like foaming, wetting, and hard surface cleaning, where a low product surface tension is often desired, micelles act as surfactant reservoirs above the CMC, allowing for product dilution without significant changes in surface tension. In cases like membrane protein extraction, a high CMC is preferred since the extraction efficiency typically plateaus at around four times the CMC of the surfactant due to self-association.

[0073] The CMC is influenced by multiple factors, like temperature, solvent, pH, chemical structure, pressure conditions and size of the tail and head groups. Determination of CMC is time-consuming and expensive, and several methods can be used like tensiometry, refractiv index, calorimetry, viscosity and conductivity measurements. In most methods, a breakpoint in the measured property (e.g., surface tension or conductivity) vs. surfactant concentration curve is observed, and the CMC is defined to be at that point.

[0074] Surfactant-containing compositions or formulations may comprise one or more surfactant(s). The surfactants used may be anionic, nonionic, cationic, and / or amphoteric or zwitterionic surfactants. In surfactant-containing compositions, for example shower gels, foam baths, shampoos, cleaner, washing liquid etc. at least one anionic surfactant may be present. The anionic surfactant may be combined with another type of surfactant, typically except for cationic surfactants, in particular nonionic or zwitterionic / ampholytic surfactants. In all embodiments disclosed herein, more than one surfactant of any type may be included, for example 2, 3, 4, 5 or more different (types of) surfactants. "Alkyl (ether) sulfates” and "fatty alcohol (ether) sulfates”, as used herein, relate to the well-known class of anionic surfactants of sulfated fatty alcohols and sulfated fatty alcohol ethers, in particular the ethoxylated, propoxylated or mixed ethoxylated / propoxylated ethers of fatty alcohols. Examples of such surfactants thus include sodium lauryl ether sulfate (SLES) and sodium lauryl sulfate (SLS or SDS). As disclosed herein below, it is preferred that these are only used in low amounts, or the compositions are free of such surfactants. Different surfactants may be considered depending on the application and target properties desired for the respective applications.

[0075] Typical examples of usable nonionic surfactants are fatty alcohol polyglycol ethers, alkylphenol polyglycol ethers, fatty acid polyglycol esters, fatty acid amide polyglycol ethers, fatty amine polyglycol ethers, alkoxylated triglycerides, mixed ethers and mixed formals, optionally partially oxidized alk(en)yl (poly)glycosides and glucuronic acid derivatives, fatty acid N-alkylglucamides, protein hydrolysates (especially wheat-based vegetable products), polyol fatty acid esters, sugar esters, sorbitan esters, polysorbates and amine oxides. If the nonionic surfactants contain polyglycol ether chains, they may have a conventional homolog distribution, but preferably have a narrow homolog distribution.

[0076] Zwitterionic surfactants refer to those surface-active compounds which bear at least one quaternary ammonium group and at least one -COO(-) or -SO3(-) group in the molecule. Particularly suitable zwitterionic surfactants are the betaines, such as the N-alkyl-N,N-dimethylammonium glycinates, for example cocoalkyl dimethylammonium glycinate, N-acylaminopropyl-N,N-dimethylammonium glycinates, for example cocoacylamino-propyldimethylammo- nium glycinate, and 2-alkyl-3-carboxymethyl-3-hydroxyethylimidazoline having in each case 8 to 18 carbon atoms in the alkyl or acyl group, and also cocoacylaminoethyl hydroxyethylcarboxymethyl glycinate.

[0077] Also suitable, especially as cosurfactants, are ampholytic surfactants. Ampholytic surfactants are understood to mean those surface-active compounds which, apart from a C8-C18-alkyl or acyl group in the molecule, contain at least one free amino group and at least one -COOH or SO3H group and are capable of forming internal salts. Examples of suitable ampholytic surfactants are N-alkylglycines, N-alkylpropionic acids, N-alkylaminobutyric acids, N-al- kyliminodipropionic acids, N-hydroxyethyl-N-alkylamidopropyhglycines, N-alkyltaurines, N-alkylsarcosines, 2-alkyla- minopropionic acids and alkylaminoacetic acids having in each case about 8 to 18 carbon atoms in the alkyl group. Particularly preferred ampholytic surfactants are N-cocoalkylaminopropionate, cocoacylaminoethyl-aminopropionate and C12-18-acylsarcosine.

[0078] Typical examples of amphoteric or zwitterionic surfactants are alkylbetaines, alkylamidobetaines, aminopropionates, aminoglycinates, imidazolinium betaines and sulfobetaines.

[0079] Typical examples of anionic surfactants are soaps, alkylbenzenesulfonates, alkanesulfonates, olefin-sulfonates, alkyl ether sulfonates, glycerol ether sulfonates, o-methyl ester sulfonates, sulfo fatty acids, mono- and dialkyl sulfosuccinates, mono- and dialkyl sulfosuccinamates, sulfotriglycerides, amide soaps, ethercarboxylic acids and salts thereof, fatty acid isethionates, fatty acid sarcosinates, fatty acid taurides, N-acylamino acids, for example acyl lactylates, acyl tartrates, acyl glutamates and acyl aspartates, alkyl oligoglucoside sulfates, protein fatty acid condensates (especially vegetable products based on wheat) and alkyl (ether) phosphates. If the anionic surfactants comprise polyglycol ether chains, these may have a conventional homolog distribution, but preferably have a narrow homolog distribution.

[0080] Cationic surfactants which can be used are especially quaternary ammonium compounds. Preference is given to ammonium halides, especially chlorides and bromides, such as alkyltrimethylammonium chlorides, dialkyldimethylammonium chlorides and trialkylmethyl-ammonium chlorides, e.g. cetyltrimethylammonium chloride, stearyltrime- thylammonium chloride, distearyldimethylammonium chloride, lauryldimethylammonium chloride, lauryldimethylbenzylammonium chloride and tricetylmethylammonium chloride. In addition, the very readily biodegradable quaternary ester compounds, for example the dialkylammonium methosulfates and methylhydroxyalkyldialkyloxyalkylammonium methosulfates sold under the trade name Stepantex® and the corresponding products of the Dehyquart® series can also be used as cationic surfactants. The term "ester quats” are generally understood to mean quaternized fatty acid triethanolamine ester salts. These are known substances which are prepared by the relevant methods of organic chemistry. Further cationic surfactants which can be used in accordance with the invention are the quaternized protein hydrolysates.

[0081] Typical examples of particularly suitable mild, i.e. particularly skin-friendly, surfactants are mono- and / or dialkyl sulfosuccinates, fatty acid isethionates, fatty acid sarcosinates, fatty acid taurides, fatty acid glutamates, a-olefinsul- fonates, ether carboxylic acids, 2-sulfonated fatty acids, alkyl (poly)glycosidesZ-glucosides and / or mixtures thereof with alkyl oligoglucoside carboxylates, fatty acid glucamides, alkylamidobetaines, amphoacetals, protein hydrolysates, and / or protein fatty acid condensates, the latter preferably based on wheat proteins or salts thereof. These surfactants are preferred surfactants to be used in the compositions of the invention. In the compositions of the invention, these may be used individually or in combination.

[0082] Alk(en)yl (poly)glycosides (APGs) may be compounds of formula (I), R1 O-[G]p (I) in which R1 is an alkyl and / or alkenyl radical with 4 to 18 carbon atoms, G is a sugar radical with 5 or 6 carbon atoms and p is numbers between 1 and 10. While such compounds are commonly referred to as alkyl glycosides or alkyl poly glycosides or alkyl oligo glycosides it is also possible that the alkyl moiety is an alkenyl radical. The term APG, as used herein, thus is intended to cover both alkyl and alkenyl (poly)glycosides.

[0083] APGs can be obtained by the relevant methods of preparative organic chemistry. The APGs can be derived from aldoses or ketoses with 5 or 6 carbon atoms. While various sugar units may be used, in various embodiments, the sugar units of the APGs are derived from glucose.

[0084] The index number p in the general formula (I) indicates the degree of oligomerization (DP degree = degree of polymerization). The degree of oligomerization of the APGs is between 1 and 10 and preferably between 1 and 6. Whereas p in an individual APG molecule must always be an integer and here in particular assumes the values in the range from 1 to 6, the value p for an APG which is a mixture of different APG molecules, which differ in their individual p values, is an analytically determined calculated parameter which in most cases is a fraction. Preferably, APGs are used with an average degree of oligomerization p in the range from 1 .1 to 3.0 or from 1 .1 to 1 .8 or from 1 .2 to 1.7.

[0085] The average degree of oligomerization here is to be understood in the sense of how it is defined in the monograph K. Hill, W. von Rybinski, G. Stoll "Alkyl Polyglycosides. Technology, Properties and Applications” (VCH-Verlagsgesell- schft, 1996) in the section "Degree of polymerization” (compare pages 11-12 of the book): there it reads "The average number of glycose units linked to an alcohol group is described as the (average) degree of polymerization (DP).” In explanatory figure 2, which describes a typical distribution of dodecyl glycoside oligomers of an AOPG with a degree of DP of 1 .3, the average degree of DP is also described by a corresponding mathematical formula.

[0086] The radical R1 is preferably derived from primary alcohols with 4 to 12 carbon atoms and preferably 8 to 10 carbon atoms. Typical examples of suitable radicals R1 are butyl, hexyl, octyl, decyl, undecyl, dodecyl and myristyl. They are derived from the saturated fatty alcohols butanol-1, caproic alcohol (hexanol-1), caprylic alcohol (octanol-1), capric alcohol (decanol-1), undecanol-1, lauryl alcohol (dodecanol-1) and myristyl alcohol (tetradecanol-1), as are obtained for example in the hydrogenation of technical-grade fatty acid methyl esters or in the course of the hydrogenation of aldehydes during Roelen oxo synthesis.

[0087] Preference may be given to APGs which are derived from glucose and in which the radical R1 is a saturated alkyl radical with 8 to 12 carbon atoms and which have an average degree of oligomerization in the range from 1.1 to 3 and in particular in the range from 1 .2 to 1 .8 and particularly preferably in the range from 1 .2 to 1 .7. These APGs can for example be prepared by reacting a sugar, in particular glucose, under acid catalysis with a fatty alcohol mixture, the fatty acid mixture used preferably being a forerunning produced during the distillative separation of technical- grade C8-18-coconut fatty alcohol, which comprises predominantly octanol-1 and decanol-1 and also small amounts of dodecanol-1.

[0088] Suitable 2-sulfonated fatty acids include, without limitation, 2-sulfolaurate and salts thereof, in particular disodium 2- sulfolaurate.

[0089] The one or more surfactant(s) may be present in the surfactant containing composition in amounts of from 1 to 30 wt.-%, preferably 2 to 25 wt.-% or 5 to 25 wt.-% or 8 to 30 wt.-% or 10 to 25 wt.-% or 15 to 25 wt.-%, relative to the total weight of the composition. In various embodiments, the one or more surfactant(s) may be present in amounts of 5 to 20 or 10 to 15 wt.-%. Such surfactant-containing compositions may be rinse-off compositions.

[0090] In such surfactant-containing compositions, for example shower gels, foam baths, shampoos etc. at least one anionic surfactant is preferably present, preferably in combination with at least one nonionic or zwitterionic surfactant. In such embodiments, the at least one anionic surfactant and the at least one nonionic surfactant may be selected from the surfactants indicated as being preferred above. For example, the anionic surfactant may be a 2-sulfonated fatty acid, such as 2-sulfolaurate, and the nonionic surfactant may be an APG. The anionic surfactant may alternative be a sulfosuccinate or fatty acid glutamate and the nonionic / zwitterionic surfactant may be an alkylamidobetaine.

[0091] Surfactants as commonly used for e.g. detergents or cosmetic and personal care applications may include biological surfactants. Biosurfactants, also known as biological surfactants, are a class of surface-active compounds derived from living organisms. Unlike traditional surfactants, which are predominantly synthesized from petroleum-based resources, biosurfactants are produced by microorganisms such as bacteria, yeasts, and fungi. These microorganisms produce biosurfactants as a natural defense mechanism, allowing them to interact with their environment. The production of biosurfactants can occur through various mechanisms, including extracellular secretion or cell-associated synthesis. Biosurfactants offer several advantages over their synthetic counterparts. They have lower toxicity, higher biodegradability, and are derived from renewable resources. Due to their natural origin, biosurfactants are considered more environmentally friendly and sustainable. Additionally, some biosurfactants exhibit unique properties such as stability over a wide range of temperatures, pH levels, and salinity, making them highly versatile in different applications. The use of biosurfactants in cosmetic compositions aligns with the growing consumer demand for natural and sustainable products. They contribute to reducing the environmental impact associated with conventional surfactants while maintaining or even improving the performance of cosmetic formulations.

[0092] Biosurfactants can be classified into several different types based on their chemical structure and the microorganisms that produce them. Here are some commonly encountered types of biosurfactants:

[0093] 1. Glycolipids: These biosurfactants consist of a hydrophilic carbohydrate moiety linked to a hydrophobic fatty acid chain. Glycolipids are produced by various microorganisms, including bacteria, yeasts, and fungi. Examples of glycolipids include rhamnolipids, sophorolipids, and trehalolipids.

[0094] 2. Lipopeptides: Lipopeptides are biosurfactants composed of a peptide chain attached to a lipid group. They are predominantly produced by certain bacteria, such as Bacillus species. Examples of lipopeptides include surfactin, iturin, and fengycin.

[0095] 3. Phospholipids: These biosurfactants contain a phosphate group in their molecular structure. Phospholipids are commonly found in cell membranes and are produced by various microorganisms, including bacteria and yeasts. Examples of phospholipids include phosphatidylcholine, phosphatidylglycerol, and phosphatidylethanolamine.

[0096] 4. Saponins: Saponins are biosurfactants that are naturally found in plants. They are glycosides composed of a hydrophilic sugar moiety attached to a hydrophobic triterpenoid or steroid structure. Saponins have been used for centuries in traditional medicine and have surfactant properties, making them suitable for cosmetic applications. 5. Polymeric biosurfactants: These biosurfactants are composed of repeating units of hydrophilic and hydrophobic moieties, forming long polymer chains. They are produced by certain bacteria and yeasts. Examples of polymeric biosurfactants include lipopolysaccharides (LPS) and lipoproteins.

[0097] Biosurfactants of the group of glycolipids are the most investigated for cosmetic and personal care use today. Given the wealth of possibilities, surfactants may be selected by their target properties for the specific application. However, the lack of standards in surfactant property measurements and the dependence on measurement conditions make such selection difficult. For example, conducting experiments in laboratories is not always simple, especially at high temperatures and pressures. In some cases, experimental measurements are expensive and / or timeconsuming and may involve uncertainties about impurities, possible decompositions, etc. In addition, the specific properties may by influenced by multiple factors, like temperature, solvent, pH, chemical structure, pressure conditions and size of the tail and head groups. Determination of specific properties is hence only reliable in connection with the measurement conditions, time-consuming and expensive.

[0098] To simplify and harmonize the generation or measurement of surfactant properties and to reliably provide surfactant properties also dependent on measurement conditions, data-driven models such as graph neural networks may be utilized for property prediction and surfactant selection.

[0099] Figs. 2-5 illustrate an example method for generating surfactant properties based on a molecular representation of a surfactant molecular structure generated by a graph neural network in connection with a mapping network configured to map the molecular representation to the surfactant property.

[0100] Fig. 2 illustrates the starting point of the generation process. The molecular structure specification of the surfactant may be mapped to a numerical graph representation. Such mapping may include the determination of feature vectors based on a pre-defined feature specification. Pre-defined feature specifications for molecular structures like surfactants are illustrated in Figs. 3 and 4. The features may be implemented as one-hot encoding or in other words based on a pre-defined feature specification. The pre-defined feature specification may relate to atom features corresponding to atoms of the molecular structure. The atom features may be represented as nodes or vertices of the numerical graph representation. The pre-defined feature specification may relate to bond features corresponding to bonds of the molecular structure. The bond features may be represented as edges or arcs of the numerical graph representation. In other words, the graph representation may include nodes corresponding to atoms and vertices corresponding to bonds between two atoms. The feature vector may be assigned per node and vertex that represent atom types such as C atom or orbital hybridization, and bond type such as double bond or ring structures.

[0101] In the data representation, illustrated in Figs. 3 and 4 atoms may be represented as nodes and bonds as edges. Each node may encode atomic information such as the atom type, aromaticity, hybridization, number of bonds the atoms is connected with and the number of bonded hydrogen atoms to treat hydrogen implicitly. In the example, the atom type can be one-hot encoded into 13 categorical features based on the predefined list of chemical elements. Edge features (e.g., bond type) may be explicitly included through bond type, bond being part of a ring, conjungation, and stereo. This type of data representation may result in 30 features per atom and 12 features per bond.

[0102] Based on the feature vectors the molecular structure may be represented in a matrix as numerical representation. For example, a feature matrix and / or adjacency matrix may be generated from the feature vectors. This way the molecule may be represented as a molecular graph where nodes correspond to atoms and edges correspond to bonds between two atoms. Be assigning the feature vector to each node and each edge that includes information about atom types and bond types, the molecular structure can be mapped from molecular structure representation e.g. via SMILES to the numerical representation e.g. via feature and / or adjacency matrix. The feature and / or adjacency matrix may represent the molecular structure specification of the surfactant in the numerical graph representation that is processable by the graph neural network (GNN).

[0103] The GNN architecture configured to generate surfactant properties based on numerical graph representation may include a graph convolutional neural network with multiple convolutional layers as illustrated in Fig. 5. The goal of graph convolutional layers may be seen in combining node information of a considered node with node information of its neighboring nodes and bonds. In other words, the node state vectors, and edge state vectors may be combined through message and update functions to an updated the state vector of the considered node. This functionality is commonly referred to as message passing. To illustrate message passing on an example: Within the graph convolutional layer, information about neighboring nodes N(v) = {w | evwe E, w + v} may be passed to the considered node via message and update functions. The message function may relate to where hwl~1denotes the hidden state of a neighbor in layer I - 1 denotes the respective edges.

[0104] I e {1 , 2, .... L} relates to the number of neighboring node layers L to be included. The message function Mtdepends on the previous hidden states of the neighbors hwl~1and the features of the respective edges fE(evw). The hidden state of the considered node hvl~1is updated to hvlby an update function hwl= [ / ( ( / iJi,- 1, mvl) that combines its hidden state from the previous layer with the received message containing information of its neighbors. The message passing and updating is repeated for a fixed number of iterations L which results in multiple graph convolutional layers. This example is not to be considered limiting and other message passing algorithms may be employed.

[0105] Once the GNN layers are passed, the results may be pooled though a pooling function such as mean, sum, or max function. The molecular fingerprint or lower-level molecular representation may for example be generated according to the following GNN architecture: where v e G represent the nodes in the graph, h^1the hidden state (or feature vector) of node v in layer L- 1 , 9Vis learnable parameter matrix, w e JV(v) the neighbor nodes of every node v, h^1the hidden state of neighbor nodes in layer L-1 and fEthe hidden state of edge evw. This example is not to be considered limiting and other architectures may be employed.

[0106] The molecular fingerprint vector or lower-level molecular representation may be enriched with measurement conditions for which the property is to be generated. For example, the molecular fingerprint vector or lower- level molecular representation may be concatenated with one or more representations of the measurement conditions. Such representation may include normalized values of one or more measurement condition(s). For example a normalized temperature value may be concatenated to the molecular fingerprint vector or lower- level molecular representation as generated by the graph convolutional networks.

[0107] The enriched molecular fingerprint vector or lower-level molecular representation may be provided to a data driven property model mapping the enriched molecular fingerprint vector or lower-level molecular representation to the surfactant properties. The data driven property model may include a regression model such as a multilayer percepton (MLP). The MLP may include input and output layers with multiple hidden layers in between. They may utilize activation functions at each of their calculated layers. Hence, the inputs may be pushed forward through the MLP by taking the dot product of the input with the weights that exist between the input layer and the first hidden layer. This dot product may yield a value at the hidden layer. Then the calculated output at the current hidden layer may be transformed through one or more activation function(s) such as rectified linear units (ReLU), sigmoid function, tanh. Once the calculated output at the hidden layer has been pushed through the activation function, it maybe pushed to the next layer in the MLP by taking the dot product with the corresponding weights. These steps may be repeated until the output layer is reached. At the output layer, the calculations will either be used for a backpropagation algorithm that corresponds to the activation function that was selected for the MLP (in the case of training) or the surfactant properties may be generated (in the case of inference).

[0108] The example architecture illustrated in Fig. 5 may enable an end-to-end learning process from molecular structure matrix to property. The architecture of Fig. 5 allows all functions from the molecular graph to the property to be explicit and differentiable allowing for supervised training using backpropagation. In addition, it relies on only few atomic and bond, and thus provides a flexible model structure that can possibly learn a broad variety of properties.

[0109] Fig. 3 illustrates an example of a graph neural network GNN trained to generate surfactant properties.

[0110] The architecture illustrated in Fig. 5 may be trained based on known molecular structures and corresponding surfactant properties measured for such structures. Depending on the behavior of such properties multiple models based on the illustrated model architecture may be trained per surfactant propriety and / or one model based on the illustrated model architecture may be trained for multiple surfactant proprieties. In the latter case the weights of the model may be shared across properties. Further hidden layers may be introduced per property to include shared and property specific weights. For example, the GNN for generating the molecular fingerprint vector or lower-level molecular representation may be shared while separate MLPs may be used for property generation. The combination of shared and specific layers in one model may be particularly relevant for small data sets and may reduce overfitting. In addition or alternatively, ensemble training may be employed where multiple models are trained and utilized for a single property generation. In such embodiment multiple models may be trained independently on a randomly drawn subset of the training data. Then, the predictions of the individual models may be averaged to receive a more accurate property generation. This way, property generation can be improved as random model errors are averaged out by the bootstrap aggregating or bagging.

[0111] For training hyperparameters or pre-defined parameters not adjusted on training may be pre-set. For the model architecture illustrated in Fig. 5 hyperparameters may include the hidden state size of the GNN, the number of convolutional layers, the message passing function and / or the number of MLP layers. On training the data set including known molecular structures mapped to the numerical molecular representation and corresponding surfactant properties measured for such structures may be provided to the model. The weights of the model may be adjusted during training such that the measured surfactant properties may be generated. For weight adjustment the loss function may be defined by the difference between measured surfactant properties and generated surfactant properties. In addition, backpropagation may be used for weight adjustment. The performance of the trained model may be validated through test data sets not included in the training or cross validation.

[0112] On inference, the thus trained model may be provided with the numerical molecular representations associated with molecular structures not used for training. The thus trained model may hence generate surfactant properties for molecular structures not seen on training. It may be used to generate the surfactant properties of new surfactants with new molecular structures. It may be used to validate properties of new surfactants with new molecular structures. It may be used to select surfactants based on target properties.

[0113] Fig. 6 illustrates an example of an apparatus for providing surfactant properties based on given surfactants and / or surfactants based on target properties.

[0114] The apparatus may include a user interface configured to provide a request to provide surfactant properties and / or surfactants. For providing surfactant properties, the request may include the molecular structure specification associated with the surfactant and optionally at least one measurement condition and further optionally a property to be generated. Fig. 7 illustrates an example of the user interface, where a SMILES string is provided as one possible molecular structure specification and temperature as measurement condition and optionally a property to be generated. The request may be provided to a processing layer configured to map the molecular structure specification to the numerical molecular representation, to provide the numerical molecular representation to the data-driven model and to generate the surfactant property based on the provided numerical molecular representation. To execute such steps, the processing layer may include a mapping agent configured to map the molecular structure specification to the numerical molecular representation and a model execution engine configured to generate the surfactant property based on the provided numerical molecular representation. The mapping agent may be communicatively coupled to a structure store configured to store pre-defined feature vectors for molecular structure. The mapping agent may be configured to access the structure store based on the molecular structure specification. The model execution engine may be communicatively coupled to a model store configured to store trained models. The trained models may be trained to provide multiple properties per model and / or one property per model. The model execution engine may be configured to access the model store and access the model based on the property to be generated as provided by the request. The model execution engine may be configured to provide the numerical molecular representation and optionally the at least one measurement condition to the accessed model. As illustrated by the example user interface of Fig. 7, the numerical molecular representation may be displayed in molecular structure format. As further illustrated by the example user interface of Fig. 7, the one or more surfactant properties - in this example CMC, Fm, foam rate and foam stability - may be displayed for the optionally provided at least one measurement condition. As further illustrated by the example user interface of Fig. 7, the one or more surfactant properties may be displayed in a spider diagram to compare with one or more reference surfactant(s).

[0115] For providing surfactant properties, the request may include one or more target properties and optionally at least one measurement condition and further optionally one surfactant type. Fig. 8 illustrates an example of the user interface, where target properties are specified in string format for one surfactant type and the temperature is specified as measurement condition. The request may be provided to a processing layer configured to generate one or more molecular structure specifications and / or numerical molecular representations, to map the molecular structure specification to the numerical molecular representation, to provide the numerical molecular representation to the data-driven model, to generate the surfactant property based on the provided numerical molecular representation and / or to select the surfactant(s) with the desired target properties. To execute such steps, the processing layer may include an instruction agent configured to generate based on the provided surfactant type, measurement condition and the required target property one or more molecular structure specifications and / or numerical molecular representations. The one or more molecular structure specifications and / or numerical molecular representations may be generated based on surfactant structure data stored and provided by the structure store. For example, the structure store may include segments of surfactant molecules that may be concatenated to molecular structure specification. Such segments may be pre-defined per surfactant type and may include head segments and / or tail segments. If one or more molecular structure specifications are generated the, the processing layer may include the mapping agent configured to map the molecular structure specification to the numerical molecular representation. The processing layer may include a model execution engine configured to generate the surfactant properties based on the provided numerical molecular representations. The mapping agent may be communicatively coupled to a structure store as described above. The model execution engine may be communicatively coupled to a model store as described above. The model execution engine may be configured to access the model store and access the model based on the one or more target properties to be generated as provided by the request. The model execution engine may be configured to provide the numerical molecular representations and the at least one measurement condition to the accessed model. As illustrated by the example user interface of Fig. 8, the numerical molecular representation may be displayed in molecular structure format. As further illustrated by the example user interface of Fig. 8, the one or more surfactant properties - in this example CMC, F m, foam rate and foam stability - may be displayed for different surfactants. As further illustrated by the example user interface of Fig. 8, the one or more surfactant properties may be displayed in a spider diagram to compare the proposed surfactant(s).

[0116] Fig. 9 illustrates an example of a method for generating on one or more surfactant properties dependent on at least one measurement condition for selecting one or more surfactant(s).

[0117] One or more molecular structure specification (s) associated with surfactant(s) may be provided. The molecular structure specification (s) may be provided in string data format such as SMILES of the molecular structure or in graphical representation of the molecular structure.

[0118] At least one data-driven model parametrized on structure representations and corresponding property data including one or more measurement condition(s) may be provided. The data-driven model may be specific to the surfactant type, at least one property to be generated and / or the measurement condition(s). The data-driven model may be specific to the surfactant type and independent of the at least one property to be generated and / or the measurement condition(s). The data-driven model may be based on the model architectures described in the context of Figs. 2-5. The data-driven model may be parametrized based on training data including molecular structure specification(s) mapped to respective numerical molecular representation(s) and measured property data including one or more measurement condition(s).

[0119] The one or more molecular structure specification (s) may be mapped to one or more numeric molecule representations). The one or more numeric molecule representation(s) may include graph representation(s) with nodes corresponding to atoms of the molecular structure and edges corresponding to bonds between atoms of the molecular structure.

[0120] One or more surfactant properties may be generated dependent on the one or more measurement condition(s) for one or more numeric graph representation(s). The surfactant properties may include the properties the data driven model is trained to generate. The surfactant properties to be generated may be provided for the provided one or more molecular structure specification (s) associated with surfactant(s). The surfactant properties dependent on the one or more measurement condition(s) may include the measurement condition(s) the data driven model is trained to generate. The surfactant properties and the related measurement conditions to be generated may be provided for the provided one or more molecular structure specification(s) associated with surfactant(s).

[0121] One or more surfactant properties dependent on one or more measurement condition(s) may be provided for selecting one or more surfactant(s). The selection of one or more surfactant(s) may be based on a manual selection by a user or an automatic selection as will be described in more detail in the context of Fig. 10.

[0122] Fig. 10 illustrates an example of a method for selecting one or more surfactant(s) based on one or more surfactant properties dependent on at least one measurement condition.

[0123] As described in the context of Fig. 9 surfactant properties dependent on measurement conditions may be generated based on the numeric molecular representations provided to the data driven model. For selection the generated surfactant properties and their corresponding numeric molecular representations may be provided to control loop. The control loop may include the target surfactant properties as objective function and compare the generated surfactant properties with the target surfactant properties. Based on such comparison and the target properties new numeric molecular representations may be generated. If the generated surfactant properties with the target surfactant properties are within a pre-defined range to the target surfactant properties, the underlying numeric molecular representations and respective molecular structure specifications may be provided. One or more surfactants corresponding to the provided numeric molecular representations and respective molecular structure specifications may be selected. For example, the surfactant with generated surfactant properties closer to the target surfactant properties than other surfactants may be selected as target surfactant. The target surfactant and its molecular structure specification may be provided for producing the target surfactant.

[0124] Example 1

[0125] Figs. 11-15 illustrate results of an example training of GNNs as e.g. described in the context of Figs. 2-5.

[0126] Data set

[0127] For model training literature data (CMC and F m) was extracted from multiple sources for all available molecules at temperatures between 20-28°C. This procedure resulted in an extended data set of 429 distinct substances for CMC and 164 different Fm values from multiple sources. During data collection, multiple CMC values for the same surfactant were found, differing from source to source, due to factors such as purity levels, measuring method of choice and mathematical evaluation of experimental data. The CMC variations remain an issue in surfactant science that the disclosure aims to solve. The total data set included Nonionics 220 points for CMC, 86 points I’m, 19 duplicate values CMC, Anionics 130 points for CMC, 44 points I’m, 44 duplicate values CMC, Cationics 55 points for CMC, 13 points F m, 27 duplicate values CMC, Zwitterionics 24 points for CMC, 21 points Fm, 9 duplicate values CMC, Total substances 429 CMC, 164 Fm and 99 duplicate CMC.

[0128] To handle duplicate values, a ranking according to the measurement method was performed. CMC values obtained via tensiometry were preferred to evaluate model on industrial grade surfactants using CMC values measured through tensiometry. If tensiometry data was not available, data from refractometry measurements taken since it was found reliable by experiment. If data only from other methods was available, i.e., neither tensiometry nor refractometry, such data was included into main data set. All remaining values for a surfactant, i.e., duplicates, were not included in the main data set but rather collected in a separate data set, which we utilize for a transfer learning approach. We note that for most surfactants where duplicate values exist, the values tend to be very similar to each other and in some cases even equal.

[0129] The described sampling process led to construct the three data sets related to nonionics, anionics, cationics, and zwitter ionics together with a detailed surfactant class distribution. Nonionic surfactants were dominant class followed by anionics. This class distribution matches with consumption data of surfactants, where anionic and nonionic surfactants are the most used in industrialized areas. In other words, the research focus is matching the industrial output. Afterwards, a statistical overview of the target properties, i.e. CMC (left) and F m (right), is presented in Fig. 11 . Fig. 11 illustrates statistical overview of CMC and F m databases, assembled from literature. Normal distribution of the data set would ensure an even train-validation-test split without artifact. F m shows a natural normal distribution without applying the logarithm. Both data sets have similar mean, median, 5th and 95th percentile values although no comparison between them should be made, as CMC values are scaled. The smallest and biggest values in both data sets are similar too. Finally, a correlation plot between log CMC and Fm is given in Fig. 12, containing surfactants for which both CMC and Fm values are collected. Fig. 12 shows correlation plot between log CMC and Fm for all surfactant classes. In the plot, 141 surfactants are presented, for which both CMC and F m data was collected from the literature.

[0130] The main species in surfactants included in data set are (S1) Texapon 842 UP (Sodium Caprylyl Sulfate), (S2) Tex- apon EHS (Sodium 2-Ethylhexyl Sulfate) and (S3) Texapon K 12 G (Sodium Dodecyl Sulfate).

[0131] Methods

[0132] In this section, the fundamentals of a GNN model, the general training settings of current works and the hyperparameter selection are presented. Afterwards, the learning techniques applied will be presented. The CMC of the three industrial surfactants was determined by plotting the surface tension as a function of the logarithm of the surfactant concentration. From this plot, two linear regions were determined, which correspond to the linear concentration-dependent and the linear concentration-independent region, respectively. The CMC value is then obtained from the intersection of the straight lines. Finally, for the surface tension measurement a Force Tensiometer - K100 (Kriiss, Germany), at 23°C was used.

[0133] Graph Neural Networks

[0134] In GNN models, every molecule is treated as an undirected graph, where atoms correspond to nodes and bonds to edges. A feature vector, containing chemical information, is assigned to each atom and each edge. The node and edge features of choice are shown in Figs. 2 - 4. Hydrogen atoms are not considered as individual nodes but are implicitly represented in the node feature vector. A surfactant example is given in Fig. 2. The molecular graph then passes through graph convolutions, where neighbor information, i.e., neighboring node and edge features, is aggregated for each node in the graph accordingly. The network depth L, i.e., the number of graph convolutional layers, defines the neighborhood pool from which structural information will be aggregated. Edge-conditioned graph convolutional layers and a gated recurrent unit (GRU) are used. Thus explicitly included is bond type information in learning the molecular structure which potentially facilitates distinguishing molecules with similar heavy atoms but different bonds, e.g., alkanes versus alkenes. After the last graph convolutional layer, the final updated atom feature vectors are pooled into a final molecular fingerprint vector hFP through a permutation invariant function, i.e., summation of all node vectors. The hFP contains all the necessary structure-related information of a specific molecule required for molecular property prediction, thereby replacing the selected descriptors in classical QSPR techniques.

[0135] The model is implemented in the Pytorch Geometric (PyG) framework. For the attributed molecular graph generation, we use the SMILES string of each molecule and RD Ki t (version 2022.3.5), an open-source toolkit for cheminformatics.

[0136] The high-quality data set CMC is used to define the hyperparameters of our GNN models. For the target property CMC, the log CMC is calculated and then standardized to a zero mean and a standard deviation of one. The train and test sets are separated randomly in a 85% - 15% ratio respectively. For the hyperparameter selection an internal validation set is used, which is a subset of the training set with 20 substances each. The loss function is the mean squared error (MSE) and the optimizer is Adam. For every modeling approach, models were trained on 40 individual training subsets and the results are averaged.

[0137] Hyperparameter selection

[0138] With the hyperparameter selection procedure, the aim is to find the suitable hyperparameters of our GNN model. For a robust hyperparameter selection, each model was tested in different internal validation sets. Aa grid search for the following hyperparameters of the GNN model was performed, varying them within the respective ranges: Graph convolutional type tNNConv, GINEConvu, number of graph convolutional layers 1, 2 or 3 , usage of GRU = True or False, the batch size 4, 8 or 16, the initial learning rate 0.005, 0.01 or 0.05, dimensions of molecular fingerprint and of MLP 64, 128, or 256. In other words, the hidden layers of the graph convolution part and the hidden layers of the MLP are chosen to be the same size. The optimum combination is a GNN architecture with an initial learning rate of 0.005, a hidden state size of 64, a batch size of 16, total graph convolutional layers of 1, the NNConv graph convolutional type and the usage of GRU for the message passing scheme. Our edge feature network includes three layers with the following number of neurons: #1 12, #2: 64, and #3: 4096. The architecture exhibits 306,561 learnable parameters in total.

[0139] Single- and Multitask Learning

[0140] A surfactant molecule usually has multiple target properties. The classical learning approach, single-task learning, is to train individual models for every property of interest. In single-task learning, model parameters are directly optimized based exclusively on a single target property only, and not transferred to another property prediction task. In that sense, all available QSPR methods for surfactants are single-task learning. On the other hand, in multi-task learning multiple target properties are simultaneously predicted. The simultaneous prediction has been shown to improve the modeling accuracy of GNNs in molecular property prediction. Normally during multi-task learning, the graph convolutional layers are shared and individual MLPs are constructed for each target property. The benefits of this approach, are mainly models' ability to generalize, learn faster, reduce overfitting and data efficiency. In the present work, we investigate the prediction of CMC and of I’m with both single- and multi-task learning. Specifically, we develop a single-task learning model for each property individually and multi-task learning models for simultaneous prediction of the properties. Since both of the properties come from the same measurement procedure and therefore are correlated, we expect an improved model accuracy during multi-task learning.

[0141] Transfer Learning

[0142] Another technique for improving machine learning models is transfer learning. During transfer learning, a model is usually pre-trained on a data set, for example a synthetic one, and then the model parameters are used to initialize the training on a new unseen data set. This technique is very useful when only small data sets are available. Duplicate values to apply transfer learning to single-task CMC prediction are collected, with the scope of utilizing bigger portions of experimental data from the literature. We use the data set Duplicate Value-CMC to pre-train the model, i.e., learn the graph convolutions and MLP parameters, and afterwards we initiate our single-task CMC model with them. All the initialized parameters are optimized based on the CMC data set.

[0143] Ensemble Learning

[0144] Training and using single models can lead to under- or / and over-predictions. A well-known technique to mitigate this phenomenon in machine learning is ensemble learning. In ensemble learning, multiple models are trained on different subsets of training data set and their final predictions are averaged, resulting in more robust and generalized predictions. Ensemble learning is used for both for single- and multi-task models, by training 40 different models in 40 different subsets of our training data set, in each case. Afterwards, we use the 40 different models to perform predictions in our test set are averaged to obtain the final scores.

[0145] Results

[0146] Predictive performance Fig. 13 illustrates the performance summary for different models and prediction tasks over 40 different runs. In each case the standard deviation is also given, except ensemble learning. In the above table we use the following abbreviations: STL = single-task learning, MTL= multi-task learning, TL = transfer learning, EL = ensemble learning, MAE = mean absolute error, RMSE = root mean squared error.

[0147] The single-task GNN model for CMC exhibits an average RMSE of 0.27 on validation set and 0.33 on test set, while the variance is bigger in the test set than in the validation set. For F m the average RMSE in test set is lower than the one in the validation set, with the former equal to 0.85 and the later equal to 1 .02. Using the logarithm of Fm did not improve the performance. The model exhibits great performance in predicting the log CMC and slightly lower performance in Fm prediction. The reason for the model's under-performance may be the small size of the data set used (140 molecules) for the training and the ambiguous measuring procedures.

[0148] In multi-task learning, the GNN model for CMC prediction exhibits an average RMSE of 0.26 on validation set and 0.36 on test set. For Fm prediction, the average RMSE in test set is again lower than the one in the validation set, with the former equal to 0.59 and the later equal to 0.43. In the CMC task, the model perform identical with the one in single-task learning, both for validation and test sets, while in the F m task the multitask model exhibits significantly better performance on the validation set, with the RMSE reducing by 60%, and improved performance on the test set, with the RMSE reducing by 20%. Therefore, we conclude that the data limitations of the F m database as single target property can be overcome by applying multi-task learning. On the other hand, the CMC model did not benefit from the additional data and showed identical results with a slightly higher variance but an overall similar accuracy compared to the single-task learning. The transfer learning approach, i.e., pre-training the model using the 99 collected duplicate values, is applied only on the single-task CMC model. The RMSE on the validation set remains the same as in single-task learning, equal to 0.27 and on the test set equal to 0.33. Interestingly, the standard deviation increases in both sets. The increase may be due to the broader range of target values for the same property, which leads the model to deviate more from the true value. We observed that transfer learning slightly reduced the final model training time, i.e., the model reached its optimum sooner. Besides the slight reduction of final model training time, using duplicate values for transfer learning led to similar performance.

[0149] Ensemble learning, i.e., averaging the predictions of the 40 trained models, slightly reduces the RMSE on test set for single-task learning to 0.28 and for multitask learning to 0.31 in the CMC case. A similar RMSE reduction is observed on test set for F m accordingly, to 0.76 for single-task learning and to 0.56 for multitask learning. We use the ensem- bled results in multitask CMC and Fm learning to draw the parity plots, shown in Figure 4 on the independent test sets. The parity plots (measured vs predicted values) for CMC and F m show a high determination coefficient for the former, R2 CMC=0.94, and moderate one for the latter R2 Fm354 = 0.74. We demonstrate that the GNN approach in the present work, can predict CMC and F m across all surfactant classes.

[0150] Fig. 14 illustrates the results for multi-task GNN ensemble models for (left) CMC and (right) Fm for all surfactant classes. The light dashed lines represent the 10% error and the dark dashed lines the 20% error. On the left the CMC test set is shown by predicted versus experimental value of log(CMC) in pM. On the right the F m test set is shown by predicted versus experimental value of Fm in mol / cm2.

[0151] Fig. 15 illustrates the distribution plot of RMSE on validation set for all the learning tasks. The single and multi task results for CMC and F m are illustrated as example. The boxplots are the results of 40 runs in different validation sets. The dark point represents the outliers.

[0152] Industrial Surfactants

[0153] The developed GNN model was used to predict on the three pure component industrial surfactants. According to the above discussed results, the best learning approach for CMC is the combination of single-task with ensemble learning was used. We use the 40 trained models to perform ensembled predictions on the three surfactants. The predicted log CMC values, as well the experimental measured ones, are given below for comparison. For all of the three, the predicted values are very close to the measured ones. Overall, the data indicates that the developed GNN model trained on literature data can accurately predict the CMCs for all three single-molecule industrial unpurified surfactants. Below the comparison of predicted values from single-task GNN with ensemble versus experimentally calculated values for three selected industrial grade surfactants is given. For the predicted values, the standard deviation over 40 individual runs is also given. The values are the logarithmic ones.

[0154] Predicted CMC (pM) vs Measured CMC (pM)

[0155] 51 4.87 ±0.11 vs 4.88

[0156] 52 4.95 ±0.17 vs. 4.98

[0157] 53 3.91 ±0.08 vs. 3.86

[0158] Example 2

[0159] Figs. 16-19 illustrate results of an example training of GNNs as e.g. described in the context of Figs. 2-5 including temperature dependence.

[0160] Data set

[0161] The data set of example 1 was extended to include multiple temperatures for each molecule (based on data availability). In addition, CMC values measured through tensiometry were prioritized and duplicates excluded. In total our new data set consists of 1387 data points, with 493 unique molecules and 228 molecules measured at least in two different temperatures. A detailed class distribution of the collected data set is given below in Table 1 . The isomeric SMILES strings are used for each molecule, in order to distinguish between different anomers, e.g., octyl-a / p-D-gly- coside and chiral centers in the sugar head, e.g., glucoside with galactoside. iii i lii v irk Im i In - lull d,! t . i , m i both test sets.

[0162] Two types of data splitting are implemented, one to test the prediction accuracy of the model for new temperatures, and one to test the ability of our model to generalize to new, unseen surfactant components.

[0163] In the first split, all unique molecules with experiments in at least two temperatures are identified and out of them, randomly select one. Note that from the shuffled data set is selected, and therefore ensure data selection in all possible temperatures. A total of 228 molecules in various temperatures are selected, which accounts for about 16% of the whole data set size. We report the class distribution of this test set in Table 1. We denote this split as new temperature.

[0164] The second split aims to evaluate models' predictive performance in completely unseen surfactant molecules in various temperatures. For consistency in the performance comparison, we choose the same test size as before, namely 228 molecules. Afterwards, we ensure that the test set contains randomly selected molecules in one or more temperatures, previously completely unseen during training. Specifically, about 70% of test set are molecules measured in different temperatures. For these molecules, we include all available temperatures on the test set. The rest 30% are molecules measured only in one temperature. The new randomly selected test set has a similar surfactant class distribution with one above although in was not enforced. We denote this split as new component and we report it is class distribution in Table 1.

[0165] Methods

[0166] The methods as in example 1 are used with the atom feature vectors including atom type, is in a ring, is aromatic, hybridization, chirality, charge, number of bonds and number of bonded hydrogen atoms and bond feature vectors including bond type, is in ring, conjugated and stereo. After the sum pooling layer, the molecular fingerprint is concatenated with the min-max normalized temperature. The temperature-informed molecular fingerprint is run through an MLP to obtain CMC predictions on the logarithmic scale.

[0167] Hyperparameter selection and GNN training

[0168] To determine the optimized hyperparameters, we train our GNN model on 40 different seeded validation sets, e.g. subsets of the training set, similar to example 1 . The size of the validation set is kept constant at 200 molecules, which represents about 14% of the whole data set size and thus a general training-validation-test split of 70:14:16. New temperatures split for hyperparameter tuning are used and the ones leading to the minimum root mean squared error (RMSE) are selected. A grid search to investigate the following hyperparameters of the GNN model is used, varying them within the respective ranges: Graph convolutional type tNNConv or GINEConvu, number of graph convolutional layers 1, or 2, usage of GRU = True, Falseu, the batch size 16, 32, or 64, the initial learning rate 0.005, 0.01, 0.05, dimensions of molecular fingerprint and of MLP P 64, 128. The number of MLP layers is kept constant at 3 layers throughout. In total 144 different unique parameterized models are developed and result in a model architecture with GINEConv, 1 graph convolutional layer, a batch size of 32, an initial learning rate of 0.005 and 128 neurons in the graph convolutional layer, as well as in the MLP. Note that the first layer of the MLP has a size of 129 neurons, because we are concatenating the normalized temperature. Furthermore, we apply the following general training settings: a maximum of 300 epochs, an early stopping patience of 60 epochs, a learning rate decay of 0.8 with a patience of 3 epochs and Adam as optimizer.

[0169] Ensemble Learning

[0170] A common technique in machine learning to reduce the noise of randomly chosen training and validation sets, is ensemble learning. Multiple models are trained, i.e., on different seeded subsets of non test data, and their predictions are averaged, thus leading to more robust and generalized predictions. We train 40 different models on both split types mentioned above, and then we average out the predictions to report the ensembled prediction accuracy.

[0171] Predictive performance on new temperatures and on new components

[0172] The performance of our GNN models on new temperature data split by reporting the mean absolute error (MAE), the RMSE and the mean absolute percentage error (MAPE) on the validation and test sets in Fig. 15. The validation set exhibits an average RMSE of 0.299, an average MAE of 0.20 and an average MAPE of 6.149. The test set has a smaller average RMSE, MAE and MAPE of 0.193, 0.132 and 4.144 respectively. The lowest RMSE, MAE and MAPE are found on the ensembled test set prediction, with a RMSE of 0.171, a MAE of 0.114 and an MAPE of 3.611. The model performs significantly better on the test set, as it contains molecules already seen during training but at different temperatures, while the validation subsets contain molecules completely unseen during training in a single temperature entry.

[0173] In the new components split the generalization capabilities of the developed GNN models are tested. The model performs predictions on previously unseen surfactant molecules in single and multiple temperatures. For the non-en- sembled model, the average RMSE, MAE and MAPE on the validation set are 0.309, 0.225 and 6.512 respectively. The test set has a similar average RMSE, MAE and MAPE, namely 0.286 and 0.199 and 6.554 respectively. Again, the ensembled predictions have the lowest errors, 0.186, 0.123 and 4.045 for RMSE, MAE and MAPE respectively on the test set. Contrast to split for new temperatures, the averaged errors on the validation and test sets of the non- ensembled model are similar. This behavior is explained from the fact that both sets, on the components split, contain identical amount of unseen surfactant molecules. Overall in both splits, the ensembled GNN model outperforms the single-task ones. As expected, the harder prediction task for new components shows a slightly higher error than the prediction task for new temperature. Furthermore, we use the ensembled model predictions to construct parity plots for both split types in Fig. 17. Overall, both new temperatures and new components splits show a high R2 score and very good agreement between measured and predicted data, as most of the points lie close to the diagonal. We notice the existence of outliers in both cases and we report the 4 with the highest absolute percentage error (APE) in Figs. 16 and 17 respectively. The exact model predictions are directly compared with the experimental measurements in Tables S1 and S2 and we also provide the data source for the interested user.

[0174] Analysis of temperature impact on model accuracy

[0175] Besides the overall model performance on both splits, we are interested on the model performance for both new temperatures and new components in the whole temperature spectrum. Therefore, we separate the individual temperatures into temperature bins (ranges) and we calculate the MAPE and RMSE in each one. The calculations are made with the ensembled GNN models, as they were found above to perform better. The results are plotted in Fig. 18. In almost every temperature bin, the MAPE is slightly higher on the new component split than on the new temperature split. This observation is explained from the model performance, we reported in Fig. 16. The last three temperature ranges show inverse results, e.g., the MAPE error is higher on the new temperature split as in the new component split. However, this observation can be explained by the limited number of samples present on these three temperature ranges in the new component split, as also shown in the Fig. 18.

[0176] Predictive performance per surfactant class

[0177] The temperature effect on the CMC varies through each surfactant class. From a modeling perspective, it is important to analyze the results for each class to identify weak spots in the model. Herein, we calculate the MAPE and RMSE per surfactant class for both split types. We report the results in Fig. 19. The ionic surfactants show slightly lower error for the new component split compared to the new temperature one, even though it contains completely new unseen molecules. For nonionic surfactants, both error metrics are doubled for new component split, while for zwitterionic an error reduction of about 30% in new component split is observed. The high model performance on unseen ionic surfactants can be explained from the steady U-shaped relationship between CMC and temperature that occurs. For nonionic surfactants, where sugar-based surfactants are also considered, the complexity of CMC dependency on temperature leads to higher model errors.

[0178] Predictive performance on sugar-based surfactants and model limitations

[0179] Models' performance was analyzed on some selected sugar-based surfactants present on new component split. As previously done, we only look on the ensembled predictions of our GNN model. We plot our ensembled predictions together with the actual values from the literature in Figure 4. To illustrate our models' error on the CMC we directly plot the actual CMC in mM. We observe the ability of our model to accurately predict two out of the three sugar- based surfactants, namely decyl p-D-glycoside and decyl p-D-maltoside. On the other side, our model fails to accurately predict the real CMC values of octyl p-D-Thioglucoside. The nature of the linker, e.g., the connection between carbohydrate and alkyl chain, impacts the surfactants' properties and consequently the CMC. Replacing the ether -0- with thioether -S- linker, leads to incorrect predictions from our model. We note that our data set contains no other sugar-based surfactant with thioether as the linker. Thus, expanding further the data set could improve the predictions of our model.

[0180] Although the model predictions are close to the real measurements, we observe from Figure 4 that they fail to follow the same pattern as the real measurements, e.g., the U-shaped relationship or monotonic decrease. Further surfactants, present in the new components split, are similarly plotted in Figure S4. According to them, our model can accurately predict the CMC of the test surfactants in multiple temperatures, however it fails to exhibit the same relationship, as was in the case for sugar-based ones.

[0181] Example GNN architecture (as an example of data-driven property model) in particular for temperature-dependency of single surfactant molecules may be described in the following, as e.g. illustrated in Fig. 23.

[0182] Each surfactant may be represented with an isomeric SMILES that may encode information regarding the anomeric configuration und chiral centers, i.e., how are the different bonds structured on the 3D space. This may allow not only to process CMC with canonical smiles strings but allow differentiation between o- / p- anomers. Afterwards, a feature vector each assigned to each atom / node on the graph may be created. In the next step, the surfactant molecules, now represented as graphs with a feature vector assigned to each atom and another feature vector assigned to each bond may be going through a convolutional layer. The following example algorithm may be used to update the node features:

[0183] The term may refer to edge feature vector and may be incorporated in the convolutional layer. This may add expressive power to the convolutional layer without overfitting the data. In the next step, the updated node features may be pooled into one vector (molecular fingerprint) that represents the each surfactant molecule. This vector may be used as an input to a MLP (Multi-layer Perceptron) to predict the logarithmic CMC, an extra neuron may be added in the second layer of the MLP that may account for the temperature. That is, the node may have the temperature value in which the measurement of the CMC was conducted. This node may affect the training as following: During the backpropagation algorithm, information regarding the temperature from the extra node may "travel” until the input features vectors of each node and may influence the update of weight matrix W01mentioned above.

[0184] Example 3

[0185] In this example a data set of temperature-dependent CMC values for a variety of surfactants was obtained. Particularly, the data of example 1 may be extended, to include CMC information at multiple temperatures for each molecule, when such data was available. As in example 2, CMC values measured through tensiometry were prioritized and make up a high percentage of the assembled data for example 3. Duplicates were excluded as in example 2. Each data point includes the CMC, the temperature, and the isomeric SMILES string of the surfactant, allowing e.g. distinguishing between different anomers, e.g., octyl-o / p-D-glycoside, and chiral centers in the sugar head, e.g., glucoside with galactoside. In total, the data set of example 3 consists of 1377 data points, with 492 unique surfactant structures, of which 227 structures were measured at least at two different temperatures. The minimum temperature of the data is 0 °C and the maximum is 90 °C. The different relationships in each surfactant class between temperature and CMC discussed above are also present in this data set. Surfactant examples include three ionic surfactants, namely, S1 , S5, and S6 exhibiting a U-shaped relationship. The minimum CMC is between 30 and 40 °C. An exception to the U-shaped relationship is S2, an ionic surfactant, as an increase in the temperature causes the CMC to increase too. Further, sugar-based nonionic surfactant S3 and zwitterionic surfactant S4 exhibit a U-shaped relationship.

[0186] Data splits

[0187] Two types of data set splitting was used: (I) different temperature for testing the prediction accuracy of the example model at new temperatures; (ii) distinct surfactant for testing the ability of the example model to generalize to unseen surfactant structures, hence structures not included in the model training.

[0188] First all unique surfactant molecules were identified with CMCs in at least two different temperatures and then randomly one data point for each of these molecules was selected to be in the test set. For example, if for a given surfactant, there exist three measurements at three different temperatures in the full data set, one may be randomly assigned to the test set and the other two may remain in the training set. A total of 227 molecules at various temperatures were selected for testing, which accounts for about 16% of the whole data set size. The distinct surfactant test set aims to evaluate the models' predictive performance for completely unseen surfactant molecules at various temperatures. A similar test set size of 218 data points was used. Randomly selected molecules were selected and added all corresponding data points to the test set, such that the temperature-dependent CMC values of these molecules remained completely unseen during training. In this test set, about 70% of the data points were molecules measured at different / multi pie temperatures. For these molecules, all of the available CMC values measured at different temperatures were included in the test set. The remaining 30% of the test set was molecules measured only at one temperature. In total, the test set contains 100 different surfactant structures, with CMC values at various temperatures, thus in total 218 data points.

[0189] The structure of this example GNN model, used ensemble learning, the hyperparameter selection and training settings of the example model (as an example of a data-driven property model) are presented in the following. Further, the baseline QSPR model, i.e., stochastic gradient boosting is described.

[0190] Graph Neural Networks and Ensemble Learning.

[0191] A GNN model (as an example of a data-driven property model) for predicting CMC values of surfactant monomers including the temperature dependency is illustrated in Fig. 23. The GNN model takes as input a surfactant molecule, represented as an undirected molecular graph, where atoms correspond to nodes (vertices) and bonds to edges. Each node and each edge is assigned a feature vector, where chemical information about the corresponding atom / bond is stored. Example node and edge features are presented in the following tables:

[0192] Example atom features used in a molecular graph representation, e.g. implemented as one-hot-encoding:

[0193] Example edge features used in a molecular graph representation, e.g. implemented as one-hot-encoding:

[0194] GNNs may operate directly on the node and edge features of the undirected molecular graph. During graph convolu- tions, the node features, denoted as hidden states, are updated with structural information from their neighborhood. Then, a pooling step is applied. That is, the updated hidden states of all nodes within the graph, which now contain information about the respective node itself and the neighbor nodes, are combined, e.g., by applying the sum operator, into a single unique vector representation, known as the molecular fingerprint. An MLP then may map the molecular fingerprint to the property of interest, herein the temperature-dependent CMC. To include the temperature, concatenating the temperature to the first hidden layer of the MLP, instead of the molecular fingerprint, enhanced model's sensitivity to temperature changes. Note that the error metrics in this example remained similar for both architectures. As an example, the temperature was normalized between a minimum of 0 and a maximum of 10. Increasing the normalization bounds to 0 and 10 instead of 0 and 1 increases the sensitivity of the model to the temperature dependency of the CMC, as the other values in the first hidden layer may have the same magnitude.

[0195] In our previous work, we considered stereochemistry information on the edge feature vector. Chirality and stereochemistry of a surfactant molecule may impact the CMC, so additionally chirality information on the atom feature vector may be used. Using threedimensional (3D) information may be beneficial but may require to determine or calculate 3D coordinates and conformers, which may be computationally expensive compared to using (only) chirality information on the atom feature vector.

[0196] Further, ensemble learning was applied in this example. Ensemble learning may be a technique in machine learning to reduce the noise of randomly chosen training and validation sets. Multiple models may be trained, i.e. , on different splits of training and validation sets, i.e., nontest data, and their predictions may be averaged, thus e.g. leading to more robust and generalized predictions. For example, 40 different models may be trained on both split types mentioned in the previous section, and then average the predictions to report the prediction accuracy of the ensemble of GNNs.

[0197] Implementation and Hyperparameter.

[0198] Hyperparameter values are determined for example by training ther GNN model on 40 different seeded validation sets. The size of the validation set may be kept constant at 200 molecules, which represents about 14% of the whole data set size and thus a general training-validation-test split of 70:14:16. The different temperature test set may be used for hyperparameter tuning, and the ones leading to the minimum root mean squared error (RMSE) on the validation set may be chosen. The CMC (pM) values may be scaled using a (based on 10) logarithmic scale. A grid search may be used to investigate the hyperparameters of the GNN model described e.g. in the following table. The hyperparameter dimensions therein may refer to the size of the molecular fingerprint and the size of the MLP

[0199] Hyperparameter Range Optimized geometry

[0200] Graph convolutional layers (1, 2) 1 Graph convolutional type onv, GINE GINEConv Usage of GRU True, False False Initial learning rate 05, 0.01, 0 0.005 Batch size (16, 32, 64 32 Dimensions (64, 128) 128 Number of MLP layers 3 3 Activation function ReLU ReLU Maximum epochs 300 300 Early stopping patience 30 30 Learning rate decay 0.8 0.8 Patience 3 3 Optimizer Adam Adam

[0201] The first hidden layer of the MLP may for example have a size of 129 neurons for concatenating the normalized temperature to it.

[0202] Each surfactant molecule may be represented with an isomeric SMILES string. RD Kit (version 2022.3.5), an open- source toolkit for cheminformatics, may be used to generate the attributed molecular graph for each surfactant. For the graph convolutions, the GINE-operator62,76 as implemented in PyTorch Geometric (PyG)77 may be applied with sum being the pooling layer of choice. Sum pooling may be used which may be beneficial for molecular-dependent prediction tasks.

[0203] Baseline QSPR Model.

[0204] A baseline QSPR model based on molecular descriptors may be employed. As molecular descriptors, extended-connectivity fingerprints, e.g. Morgan fingerprints, may be used. A radius of 5 and a bit size of 1024 may be chosen. To generate them, for example RDKit (version 2022.3.5) may used and atom chirality information to be encoded may be enabled. Afterward, the normalized temperature, here for example between 0 and 1, may be concatenated to the molecular descriptors. The updated molecular descriptors may be passed into a SGB to predict the CMC, which may be scaled as described above. A SGB may be chosen as it is an ensemble algorithm, and thus, it can be compared with the ensemble of GNNs. A separate validation set may be omitted. The SGB may be implemented in Python with the scikitlearn module and the hyperparameters given as an example in the following table.

[0205] Example model performance is illustrated in Figs. 20a - 22b.

[0206] Figs. 20a and 20b illustrate predictions of an ensemble of GNNs (in particular a single-target GNN model applied to single surfactant at various temperatures (i.e. at least two different temperatures), which may predict CMC values of single surfactant molecules at multiple temperatures) versus experimental data on a nonionic surfactant. In Fig. 20a the predictions are provided on the log CMC scale, while in Fig. 20b on an absolute scale (mM) is used. Model predictions are symbolized with "closed dot” while laboratory experiments with In both scales, the model accurately predicts the order of magnitude of CMC at various temperatures and moreover high sensitivity to temperature variances is observed. For the later, refer to model curvature between 20 and 45 degrees, which follow the measured curve. This surfactant structure was not used during training and therefore the GNN models extrapolate to new surfactant structure with great accuracy. Hence, the model, e.g. an ensemble of GNNs (ensemble = multiple GNNs model with the same architecture but trained on different data subsets to ensure robustness and avoid over- / under- predictions) may have learned an underlying function that accurately maps the surfactant structure and the temperature to the target, e.g., CMC values, even for completely unseen surfactant structures.

[0207] Figs. 21 a and 21 b illustrate predictions of an ensemble of GNNs (in particular a single-target GNN model applied to single surfactant at various temperatures, which may predict CMC values of single surfactant molecules at multiple temperatures) versus the experimental data on a nonionic sugar-based surfactant. In Fig. 21 a the predictions are provided on the log CMC scale, while in Fig. 21 b the predictions are provided on the absolute scale (mM). Model predictions are symbolized with "closed dot” while laboratory experiments with The results interpretation is similar to Fig. 20a, 22b. However, here a sugar-based surfactant with very complex structure is presented. In particular for these kinds of application, training of the GNN model may involve chirality (stereochemistry).

[0208] The GNN model may predict CMCs at multiple temperatures for surfactant mixtures. The model may be config- ured / trained to distinguish between surfactants with different geometries, as was the case in the temperature-dependent one too. An example binary mixture between two cationic surfactants is given below. Again, the predictions on the absolute and the logarithmic scale are provided.

[0209] Example GNN architecture (as an example of data-driven property model) in particular for temperature-dependency of surfactant mixtures may be described in the following, e.g. with reference to the example presented in Fig. 24.

[0210] The architecture is similar to the one described above for single surfactants e.g. Fig. 23, up until the molecular fingerprint. However, in that case more than one molecular fingerprint are generated. For binary mixtures, two molecular fingerprints are generated, for ternary mixtures three molecular fingerprints are generated etc. Afterwards, each molecular fingerprint is multiplied with the molar fraction of each mixture component (from 0 to 1). Two architectures are disclosed in the following.

[0211] A) the two / three / four molecular fingerprints are summed up to one

[0212] B) a new graph is constructed that considers hydrogen bonding information between the components.

[0213] In both cases, the "updated” molecular fingerprint is passed on the MLP to predict the CMC. Note that the temperature neuron is added on the MLP as described above (no changes on the architecture on this part).

[0214] Example 4

[0215] Surfactants may be amphiphilic molecules composed from a hydrophilic (head) and a hydrophobic (tail) part. Mixtures of surfactants may exhibit advantageous properties compared to single surfactants. For example, commercial formulations developed in cosmetics and detergents

[0216] Industries may contain almost exclusively surfactant mixtures. This trend may either be due to the existence of a homologue distribution in an industrial grade surfactant or the combination of surfactants driven by performance, cost and sustainability reasons. A property of a surfactant mixture is the critical micelle concentration (CMC), which may denote the emergence of surfactant micelles in the solution. Surfactant mixtures may generally be described as complex systems due to the difference between bulk and micelle concentration of the components. Mixing two surfactants may lead to a synergistic or antagonistic behavior. A binary mixture exhibits synergism, if at any molar fraction the mixture CMC, denoted as CMCM, is lower than the CMC of either individual component, and antagonism if CMCM is higher than the CMC of either individual component. The mixing behavior of surfactants may be influenced by factors such as the surfactant structure, nature, temperature, presence of electrolytes and the pH. Binary mixtures between two nonionic surfactants may mix ideally possible due to the lack of electrostatic repulsive forces between the head groups. Nonionic / anionic and nonionic / cationic surfactant systems may mix non-ideally and to behave synergistically. However, not all mixtures between ionic and nonionic surfactants may exhibit synergism. The formation of micelles in anionic / cationic mixtures may benefit from the reduction of the repulsive forces between the head groups and thus a synergistic behavior may be observed. Anionic / anionic mixtures may exhibit an antagonistic micellar behavior. Zwitterionic surfactants have a different charge depending on the pH of the solution, and hence mixtures containing may show a complex behavior. Attractive interactions between anionic and zwitterionic surfactants may occur, while almost ideal mixing between cationic and zwitterionic surfactants may occur. Overall, binary surfactant mixtures are complex systems, with interesting performance characteristics.

[0217] CMC data for 108 binary surfactant mixtures from the literature at multiple temperatures may be collected. To treat binary mixtures, two GNN architectures may be considered:

[0218] (I) a weighted linear summation of the molecular fingerprints of each component and (ii) consideration of hydrogen bond information between the mixture components.

[0219] Four different splits may be implemented, so that a wide number of testing scenarios may be covered. Both architectures may be evaluated in all four of them and the model performance per different surfactant class combination may be analyzed. The considered architectures trained only on binary mixtures may also be applied on ternary mixtures.

[0220] Data set

[0221] 108 binary surfactant mixtures at various temperatures were collected from available literature sources. The data set contained 515 mixture points from 68 unique surfactant structures. Since some surfactant structures are present at more than one temperature, the collected data set accounts for 599 data points. In this data set, 1,377 data points were concatenated for single surfactants at various temperatures. In the combined data set duplicate entries between the 68 newly collected surfactants and the 1,377 data points collected before arise, which may be averaged. Overall, the assembled data set consists of 1,924 data points. The minimum experimental temperature of the data is 0 °C and the maximum is 90 °C. The 108 binary surfactant mixtures may be visualized as a mixture network, where each node represents one of the 68 surfactants, and each edge the existence of a binary mixture between two surfactants. Furthermore, the t-distributed neighbor embedding (t-SNE) on generated extended-connectivity fingerprints (ECFP 10) for each surfactant molecule may be applied. Data splits

[0222] Performance of the example model may be evaluated under different test scenarios. Four types of data set splitting may be implemented: (i) composition interpolation (comp-inter), (ii) mixturecompositions extrapolation (mix-comp- extra), (iii) exclude one surfactant mixture (mix-ex one), (iv) mixture extrapolation (mix-ext). An overview of the four splits is provided in the following table.

[0223] Combination Mixture data set Comp-inter Mix-comp-extra Mix-ex-one Mix-ext

[0224] An.-Non. 11 9 1 5 2 An. -Cat. 13 8 1 1 0 An.-An. 8 8 0 4 0

[0225] An.-Zwitt. 6 6 1 5 0 Cat-Non. 24 23 4 9 0 Cat. -Cat. 35 32 8 5 2

[0226] Cat.-Zwitt. 5 5 2 0 0

[0227] Non. -Non. 5 4 2 1 3 Non.-Zwitt. 1 1 0 1 0

[0228] Total 108 (515) 96 (96) 19 (104) 25 (170) 7 (38)

[0229] For example, the comp-inter test set contains mixture points of previously seen mixtures but at different compositions. To select the mixture points, all binary surfactant mixtures with at least two different mixture compositions are identified. Out of the 108 mixtures present in the data set, 96 of them fulfill this criteria. For each of them, a mixture point is randomly selected, removed from the training set and assigned to the test set. By utilizing the comp-inter test set, model's performance in predicting new mixture points of a binary mixture when measured mixture points may be assessed. The mix-comp-extra test set refers to binary surfactant mixtures, where the surfactants were seen during training in other surfactant mixture combinations. In other words, the training set includes the two surfactant structures of a binary mixture, either solely as single molecules or as components of other mixtures. However, their combination (mixture) remains completely unseen. Two subsets, each containing 10 binary mixtures are randomly selected, consisting of 46 and 58 mixture points respectively. Similarly to the comp-inter test set, the mix-comp-extra test set only contains mixture points too. The mix-ex- one test set expands the extrapolation character of the mix-comp-extra split by completely masking out one of the two surfactants in a binary mixture. That is to say, one surfactant is not included in the training set, either as single molecule or as a component of other mixtures. This test scenario reflects the isolation / synthesis of a new surfactant structure, for which no previous measurements are available. To construct the mix-ex one test set, we select 4 representative surfactants, namely: N-decanoyl-N-methylglucamide (Mega-10), n-dodecyl-p-D-maltoside Q3-C12G2), cetylpyridinium chloride (CPC) and sodium dodecyl sulfate (SDS). Removing them from the training set simultaneously, would result in a huge information loss, about 30 percent of the mixture points, and thus diminish model performance. Therefore, each of them (and it is mixture) may be removed separately. Accordingly, four sub-test sets may be constructed. For example, for Mega-10 there are 9 mixtures (45 mixture points) in the assembled data set at 30 °C. The sub-test set may consist both the 45 mixture points and the Mega-10 as pure competent at 30 °C. The data points for Mega-10 at temperatures different than 30 °C58, may be excluded from both the training and the test set, since model's ability on temperature-dependent CMC predictions was already shown. Similarly, the sub-test set for p- C12G2 contains 3 mixtures (14 mixture points) at 25 °C18,21 , for CPC contains 6 mixtures (19 mixture points) at 25 and 30 °C59-62, and for SDS contains 15 mixtures (82 mixture points) at 22, 25, 30, 35 and 40 °C. Mix-ext test set considers a scenario where neither of the surfactants in a mixture are encountered during training, whether as single molecules or as components of other mixtures. In this scenario, the model has to predict CMCs of new / unseen surfactant structures, as well as mixtures of them. First, 5 mixtures composed by 6 surfactants (2 anionic and 4 nonionic) are identified that fulfill the criteria described above. To enhance structural complexity and variety of the mix-ext test set, we further assign 2 mixtures composed by 3 cationic surfactants. However, one of the 3 cationic surfactants (cetalkonium chloride) may be present in a mixture with CPC60, which as described above, may be present in the data set. This mixture may be discarded for the mix-ext split. Thus, the training set contains 100 mixtures and the test set contains 7 mixtures.

[0230] Further a data set containing 6 ternary mixtures was assembled. The ternary dataset contains 16 mixtures points at 25 and 30 °C, composed from 8 surfactant molecules. Note that all 8 surfactant structures exist in the collected data set.

[0231] Example model

[0232] In example GNNs, a surfactant molecule of a binary mixture may be represented as a graph G, = (V, E), where V are the vertices (atoms), E are the edges (bonds) between nodes, and I e {1 , 2} refers to the component of the binary mixture. After graph convolutions and a pooling step, the learned molecular fingerprint of a surfactant may be obtained, which may be denoted as hFp, To treat binary mixtures, a weighted linear summation of two learned molecular fingerprints may be used. The mixture composition (mole fraction), xbmay be used as weight. The mixture fingerprint hFPmmay be calculated through the following equation, which may allow seamless passing of single surfactants, i.e. , X2 = 0, of binary mixtures as well as of ternary ones. This example architecture (as an example of data-driven property model) may be denoted as WS-GNN (weighted sum GNN). hppm= xi ■ hpp1+ X2 ■ hpp2

[0233] The mixture fingerprint may be mapped to the temperature-dependent CMC through a multi-layer perceptron (MLP). More information regarding the MLP architecture and temperature dependency are for example provided above.

[0234] To better capture molecular interactions in mixtures, more advanced geometries that consider hydrogen bonding information may be used. For example, each component of the binary mixture may pass through a graph convolutional layer, so that each molecular fingerprint (hFPl, hpp2) may be obtained. Afterwards, the composition x of surfactant I may be multiplied with the corresponding molecular fingerprint h^,.. From the composition-informed molecular fingerprints, a global graph may be constructed where each node represents a mixture component and each edge the interaction between the two components. For node features, composition-informed fingerprints may be utilized. The edge feature vector may contain information regarding hydrogen bonding between the two components. However, the number of hydrogen acceptors and donors may be calculated through Lipinski's rule of five. The global graph may then be passed into another graph convolutional layer. For instance, the GINE-operator as implemented in PyTorch Geometric (PyG)74 may be used as a graph convolutional layer. To extract the mixture fingerprint hFPm, a summation pooling layer may be added. The mixture fingerprint may be passed through an MLP to predict-tempera- ture dependent CMCs. This example architecture (as an example of data-driven property model) may be denoted as GG-GNN (global graph GNN).

[0235] Hyperparameter tuning based on the WS-GNN architecture may be performed. For instance, the GNN model may be trained on 20 different seeded validation sets, similar to above. To reduce computational cost, the number of seeded validation sets may be decreased from 40 to 20. The size of the validation set may be selected at 385 molecules, which may represent 20% of the whole data set size. The root mean square error (RMSE) may be used on the comp- inter split to find the hyperparameters. The CMC (pM) values may be scaled using a (based on 10) logarithmic scale. Investigated hyperparameters may be given above. The model may be implemented in PyTorch Geometric (PyG)74. Ensemble learning may be employed to enhance the predictive capabilities of the example models. The models trained on different seeded validation sets may produce noisy results. By averaging out the predictions of all them, robust and generalized predictions may be obtained. For example, the predictions of the 20 different trained models for each split type mentioned above may be averaged, and (only) the prediction accuracy of the ensemble of GNNs may be reported. Further, ensemble learning in the predictions of the ensemble of the two developed GNN frameworks described above may be applied. That is to say, the predictions from the two proposed GNNs architectures (WS-GNN and GG-GNN) may be averaged for a result of an example of the data-driven property model.

[0236] Figs. 22a and 22b illustrate predictions of an ensemble of GNNs (in particular a single-target GNN model that is applied to surfactant mixtures at various temperatures, which may predict CMC values of surfactant mixtures (at least two surfactant molecules) at multiple temperatures) versus the experimental data on a mixture between two cationic surfactants. In Fig. 22a the predictions are provided on the log CMC scale, while in Fig.22b they are provided on the absolute scale (mM). The model predictions are highly accurate throughout all mixture compositions although neither the pure surfactant nor any mixtures of them were contained on the training data set. The predictions are highly accurate on both scales. Moreover, the model captures the synergistic effect of the mixture and can predict that in compositions in-between the pure components, the mixture CMC will be decreased. The model may have learned an underlying function that maps the structure of (at least) two surfactants and the experimental temperature to a CMC value. Furthermore, given a pair of cationic surfactants (even with total unknown structure) the developed GNN model may produce more than acceptable predictions to selecting a surfactant e.g. for production.

[0237] So the disclosure above and below may e.g. allow predicting at least two surfactant properties at one temperature, CMCs of one surfactant at multiple temperatures, and CMCs of surfactant mixtures at multiple temperatures. The example aspects of this disclosure may allow differentiation between o- / p- anomers. In Fig. 23 a parity plot of an example GNN model for mixtures is provided demonstrating the high model accuracy. The predicted versus the experimental log CMC values are presented. Here, the model predicts CMC values of binary surfactant mixtures at compositions different than the training set. Most of the points lie on the diagonal and hence high model accuracy may be achieved. In other words, given a binary surfactant mixture measured at least on mixture composition (for example 0.4-0.6 ratio), the GNN model according to this disclosure may make highly accurate predictions for all other unmeasured mixture compositions and hence reduce experimental effort and resources.

[0238] The present disclosure has been described in conjunction with preferred embodiments and examples as well. However, other variations can be understood and effected by those persons skilled in the art and practicing the claimed invention, from the studies of the drawings, this disclosure and the claims.

[0239] Any steps presented herein can be performed in any order. The methods disclosed herein are not limited to a specific order of these steps. It is also not required that the different steps are performed at a certain place or in a certain computing node of a distributed system, i.e. each of the steps may be performed at different computing nodes using different equipment / data processing.

[0240] The following example embodiments shall also be disclosed:

[0241] Clause 1 :

[0242] A method for selecting a surfactant based on at least one surfactant property comprising the steps of: providing one or more molecular structure specification(s) associated with one or more candidate molecular structure(s) of one or more surfactant(s) or associated with one or more surfactant segment(s) for determining one or more candidate molecular structure(s) of one or more surfactant(s); mapping the one or more molecular structure specification(s) to corresponding numeric molecular representations) associated with the atoms and bonds of the molecular structure; providing at least one data-driven property model trained on numeric molecular representation(s) and related property data associated with at least one surfactant property; generating at least one surfactant property by providing the numeric molecular representation(s) to the at least one data-driven property model; providing the at least one generated surfactant property for selecting a surfactant.

[0243] Clause 2:

[0244] The method of clause 1 , wherein the surfactant property relates to at least one surfactant structure associated with a bio-based surfactant, wherein the surfactant property relates to at least one application of the bio-based surfactant, wherein the surfactant property relates to at least one application performance of the bio-based surfactant based on the molecular structure of the bio-based surfactant.

[0245] Clause 3: The method of any of the preceding clauses or example embodiments, wherein the one or more molecular structure specification(s) relate to at least one surfactant type and / or at least one surfactant segment type associated with hydrophilic and / or hydrophobic characteristics of the surfactant.

[0246] Clause 4:

[0247] The method of any of the preceding clauses or example embodiments, wherein the mapping of the one or more molecular structure specification(s) to corresponding one or more numeric molecular representation(s) encodes at least the atom types and the bond types between the atoms, wherein the numeric molecular representation (s) are graph numerical representation(s) with nodes encoding the at-oms and / or edges encoding the bonds.

[0248] Clause 5:

[0249] The method of any of the preceding clauses or example embodiments, wherein one or more data-driven property mod-el(s) are provided dependent on one or more surfactant type(s), property type(s) and / or measurement condition type(s), wherein the surfactant type, property type and / or measurement condition type are provided to select one data-driven property model based on surfactant type, property type and / or measurement condition type.

[0250] Clause 6:

[0251] The method of any of the preceding clauses or example embodiments, wherein at least one data-driven property model trained to map numeric molecular representation (s) to multiple property types or surfactant properties and / or multiple data-driven property models trained to map numeric molecular representation(s) to one or more property type(s) or surfactant properties are provided to generate one or more surfactant properties.

[0252] Clause 7:

[0253] The method of any of the preceding clauses or example embodiments, wherein the at least one generated surfactant property is provided for selecting at least one surfactant and for providing the at least one generated surfactant property in relation to at least one other surfactant and associated surfactant properties.

[0254] Clause 8:

[0255] The method of any of the preceding clauses or example embodiments, wherein property data associated with at least one surfactant property includes one or more measurement condition(s) that influence the one or more surfactant properties.

[0256] Clause 9:

[0257] The method of any of the preceding clauses or example embodiments, wherein one or more measurement condition (s) are provided in relation to one or more surfactant properties, wherein the at least one data-driven property model is trained on numeric molecular representation(s) and related one or more surfacetant properties dependent on one or more measurement condition(s).

[0258] Clause 10: The method of any of the preceding clauses or example embodiments, wherein one or more measurement conditions) are provided in relation to one or more surfactant properties, wherein the one or more surfactant properties are provided to the data-driven property model to generate at least one specific molecular representation from the numeric molecular representation and the one or more measurement condition(s) are provided in relation to the mapping of the generated at least one specific molecular representation to the one or more surfactant properties.

[0259] Clause 11 :

[0260] The method of any of the preceding clauses or example embodiments, wherein the data-driven property model is con-figured to generate at least one specific molecular representation, wherein the data-driven property model is configured to generate at least one surfactant property, wherein the data-driven model is trained to map numeric molecular representation(s) to at least one surfactant property optionally dependent on one or more measurement conditions).

[0261] Clause 12:

[0262] The method of any of the preceding clauses or example embodiments, wherein one or more text input instruction(s) related to at least one target property and / or at least one surfactant class are provided, wherein one or more molecular structure specification(s) and / or numeric structure representation (s) are generated from text instructions.

[0263] Clause 13:

[0264] The method of any of the preceding clauses or example embodiments, wherein at least one target property and at least one surfactant class are provided, wherein multiple molecular structure specifications and / or numeric structure representations are generated based on the at least one target property and at least one surfactant class, wherein one or more surfactant properties are generated per molecu-lar structure specification and / or numeric structure representation, wherein at least one surfac-tant is selected based on the generated one or more surfactant properties.

[0265] Clause 14:

[0266] A surfactant with the surfactant structure produced based on the generated surfactant prop-er-ty and associated surfactant structure generated and selected according to the methods of any of clauses 1-13.

[0267] Clause 15:

[0268] Use of the surfactant property and associated surfactant structure generated and selected according to the methods of any of clauses 1-13 to produce a surfactant with the selected surfac-tant structure.

[0269] As used herein ..determining" also includes ..initiating or causing to determine", "generating" also includes ..initiating and / or causing to generate" and "providing” also includes "initiating or causing to determine, generate, select, send and / or receive”. "Initiating or causing to perform an action” includes any processing signal that triggers a computing node or device to perform the respective action. In the claims as well as in the description the word "comprising” or "including” or similar wording does not exclude other elements or steps and shall not be construed limiting to the elements or steps lined out. The indefinite article "a” or "an” does not exclude a plurality. A single element or other unit may fulfill the functions of several entities or items recited in the claims. The mere fact that certain measures are recited in the mutual different dependent claims does not indicate that a combination of these measures cannot be used in an advantageous implementation or further elements may be included. The expressions "A and / or B” and "at least one of: A or B” are considered interchangeable and meant to comprise any one of the following three scenarios: (i) A, (ii) B, (iii) A and B. More generally, the expression "at least one of the following: ” and "at least one of ” and similar wording, where the list of two or more elements are joined by "and” or "or”, mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements.

[0270] Providing in the scope of this disclosure may include any interface configured to provide data. This may include an application programming interface, a human-machine interface such as a display and / or a software module interface. Providing may include communication of data or submission of data to the interface, in particular display to a user or use of the data by the receiving entity.

[0271] Any disclosure and embodiments described herein relate to methods, systems, apparatuses, devices, chemicals, materials, services, uses, computer program elements lined out above and vice versa. Advantageously, the benefits provided by any of the embodiments and examples equally apply to all other embodiments and examples and vice versa.

[0272] All terms and definitions used herein are understood broadly and have their general meaning.

Claims

Claims1 . A method for selecting a surfactant based on at least one surfactant property comprising the steps of:- providing one or more molecular structure specification(s) associated with one or more candidate molecular structure(s) of one or more surfactant(s) or associated with one or more surfactant segment(s) for determining one or more candidate molecular structure(s) of one or more surfactant(s);- mapping the one or more molecular structure specification(s) to corresponding numeric molecular representations) associated with the atoms and bonds of the molecular structure;- providing at least one data-driven property model trained on numeric molecular representation(s) and related property data associated with at least one surfactant property;- generating at least one surfactant property by providing the numeric molecular representation(s) to the at least one data-driven property model;- providing the at least one generated surfactant property for selecting a surfactant.

2. The method of claim 1 , wherein the at least one data-driven property model is trained on numeric molecular representations) and related property data associated with at least two surfactant properties; and wherein at least two surfactant properties are generated by providing the numeric molecular representation(s) to the at least one data- driven property model; and wherein the at least two generated surfactant properties are provided for selecting the surfactant.

3. The method of claim 1 or 2, wherein one of the at least two surfactant properties is a critical micelle concentration, in particular a critical micelle concentration at at least two different temperatures.

4. The method of any of the preceding claims, wherein the one or more surfactant(s) are at least two surfactants.

5. The method of any of the preceding claims, wherein at least one surfactant property of the at least two surfactant properties relates to at least one surfactant structure associated with a bio-based surfactant, wherein the at least one surfactant property relates to at least one application of the bio-based surfactant, wherein the at least one surfactant property relates to at least one application performance of the bio-based surfactant based on the molecular structure of the bio-based surfactant.

6. The method of any of the preceding claims, wherein the one or more molecular structure specification(s) relate to at least one surfactant type and / or at least one surfactant segment type associated with hydrophilic and / or hydrophobic characteristics of the surfactant.

7. The method of any of the preceding claims, wherein the mapping of the one or more molecular structure specifica- tion(s) to corresponding one or more numeric molecular representation(s) encodes at least the atom types and thebond types between the atoms, wherein the numeric molecular representation (s) are graph numerical representations) with nodes encoding the atoms and / or edges encoding the bonds.

8. The method of any of the preceding claims, wherein one or more data-driven property model(s) are provided dependent on one or more surfactant type(s), property type(s) and / or measurement condition type(s), wherein the surfactant type, property type and / or measurement condition type are provided to select one data-driven property model based on surfactant type, property type and / or measurement condition type.

9. The method of any of the preceding claims, wherein at least one data-driven property model trained to map numeric molecular representation(s) to multiple property types or surfactant properties and / or multiple data-driven property models trained to map numeric molecular representation (s) to one or more property type(s) or surfactant properties are provided to generate one or more surfactant properties.

10. The method of any of the preceding claims, wherein the at least one generated surfactant property is provided for selecting at least one surfactant and for providing the at least one generated surfactant property in relation to at least one other surfactant and associated surfactant properties, optionallywhen dependent on claim 2 the at least two generated surfactant properties are provided for selecting at least one surfactant and for providing the at least two generated surfactant properties in relation to at least one other surfactant and associated surfactant properties.11 . The method of any of the preceding claims, wherein property data associated with at least one surfactant property includes one or more measurement condition(s) that influence the one or more surfactant properties, optionally when dependent on claim 2 , the property data associated with at least two surfactant properties includes one or more measurement condition(s) that influence the at least two surfactant properties.

12. The method of any of the preceding claims, wherein one or more measurement condition(s) are provided in relation to one or more surfactant properties, wherein the at least one data-driven property model is trained on numeric molecular representation(s) and related one or more surfacetant properties dependent on one or more measurement condition(s) , optionally when dependent on claim 2 ,the one or more measurement condition(s) are provided in relation to at least two surfactant properties, wherein the at least one data-driven property model is trained on numeric molecular representation(s) and related at least two surfactant properties dependent on one or more measurement condition(s).

13. The method of any of the preceding claims, wherein one or more measurement condition(s) are provided in relation to one or more surfactant properties, wherein the one or more surfactant properties are provided to the data-driven property model to generate at least one specific molecular representation from the numeric molecular representation and the one or more measurement condition(s) are provided in relation to the mapping of the generated at least one specific molecular representation to the one or more surfactant properties, , optionally when dependent on claim 2,the one or more measurement condition(s) are provided in relation to at least two surfactant properties, wherein the at least two surfactant properties are provided to the data-driven property model to generate at least one specific molecular representation from the numeric molecular representation and the one or more measurement condition(s) are provided in relation to the mapping of the generated at least one specific molecular representation to the at least two surfactant properties.

14. The method of any of the preceding claims, wherein the data-driven property model is configured to generate at least one specific molecular representation, wherein the data-driven property model is configured to generate at least one surfactant property, wherein the data-driven model is trained to map numeric molecular representation(s) to at least one surfactant property optionally dependent on one or more measurement condition(s), , optionally when dependent on claim 2, the data-driven property model is configured to generate at least one specific molecular representation, wherein the data-driven property model is configured to generate at least two surfactant properties, wherein the data- driven model is trained to map numeric molecular representation(s) to at least two surfactant properties optionally dependent on one or more measurement condition(s).

15. The method of any of the preceding claims, wherein one or more text input instruction(s) related to at least one target property and / or at least one surfactant class are provided, wherein one or more molecular structure specification(s) and / or numeric structure representation(s) are generated from text instructions.

16. The method of any of the preceding claims, wherein at least one target property and at least one surfactant class are provided, wherein multiple molecular structure specifications and / or numeric structure representations are generated based on the at least one target property and at least one surfactant class, wherein at least two surfactant properties are generated per molecular structure specification and / or numeric structure representation, wherein at least one surfactant is selected based on the generated at least two surfactant properties, optionally the at least two surfactant properties are generated per molecular structure specification and / or numeric structure representation, wherein at least one surfactant is selected based on the generated at least two surfactant properties.

17. A surfactant with the surfactant structure produced based on the generated surfactant proper-ty and associated surfactant structure generated and selected according to the methods of any of claims 1-16.

18. Use of the at least two surfactant properties and associated surfactant structure generated and selected according to the methods of any of claims 1 -16 to produce a surfactant with the selected surfactant structure.

19. An apparatus comprising respective means for carrying out or performing the steps of any one of claims 1 to 16 or comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to carry out the steps of the method according to any one of claims 1 to 16.