Determining a model and model instructions for generating chemical product data

Data-driven models are used to generate chemical product data, addressing the need for accurate production and processing by selecting suitable models and instructions, thereby improving monitoring and control in chemical product management.

WO2025219379A1PCT designated stage Publication Date: 2025-10-23BASF SE
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/060357
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-15
Filing Date
2025-04-15
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Producing and processing chemical products requires accurate data management due to the significant impact of small structural changes on their properties, necessitating tailored production and processing methods.

Method used

A method utilizing data-driven models, such as large language models, to generate chemical product data by selecting an appropriate model and model instructions based on historical outputs and requests, enabling efficient and accurate production and processing control.

Benefits of technology

This approach allows for improved production and processing of chemical products by generating precise chemical product data, enhancing monitoring and control capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025060357_23102025_PF_FP_ABST
    Figure EP2025060357_23102025_PF_FP_ABST
Patent Text Reader

Abstract

A method (100) for generating chemical product data associated with a chemical product is presented, the method including a) receiving (101) a request for generating chemical product data associated with the chemical product, the request including an indication of the chemical product, b) generating (102) the requested chemical product data by providing model instructions based on the received request as input to a trained data-driven model, and c) providing (103) the generated chemical product data. The data-driven model and / or at least part of the model instructions used for generating the requested chemical product data are determined by a respective training and / or selection process (104a, 104b, 104c). In this way, chemical product data can be generated that allow for an improved production and / or processing of chemical products.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Determining a model and model instructions for generating chemical product data

[0002] FIELD OF THE INVENTION

[0003] The present disclosure relates to a method for generating chemical product data associated with a chemical product, to a method for training a data-driven model to be used in the method, to data processing systems for carrying out such methods, and to a use of any of the foregoing or of the generated chemical product data, particularly for monitoring and / or controlling production and / or processing of the chemical product.

[0004] BACKGROUND OF THE INVENTION

[0005] Producing and / or processing of chemical products requires accurate production and / or processing of data. A small change, e.g., in a chemical structure can completely change the properties of the chemical product. Followingly, production and / or processing of a chemical product needs to be tailored to the properties of the chemical product.

[0006] SUMMARY OF THE INVENTION

[0007] It is an object addressed by the present disclosure to improve production and / or processing of chemical products.

[0008] In an aspect, a method for generating chemical product data associated with a chemical product is presented. The method includes: receiving a request for generating chemical product data associated with the chemical product, the request including an indication of the chemical product, generating the requested chemical product data by providing model instructions based on the received request as input to a data-driven model trained and / or parametrized to generate chemical product data as output in response to receiving model instructions for doing so as input, and providing the generated chemical product data as output of the method, in particular for monitoring and / or controlling production and / or processing of the chemical product, wherein: a) the data-driven model used for generating the requested chemical product data is determined from a plurality of candidate data-driven models, wherein the plurality of candidate data-driven models are trained and / or parametrized to generate chemical product data as output in response to receiving model instructions for doing so as input, wherein the data-driven model used for generating the requested chemical product data is determined based on an assessment of outputs generated by the plurality of candidate data-driven models in response to receiving, as input, model instructions based on one or more historical requests for generating chemical product data, and / or b) at least part of the model instructions used for generating the requested chemical product data are determined based on an assessment of outputs generated by a data- driven model, which is trained and / or parametrized to generate chemical product data as output in response to receiving model instructions for doing so as input, in response to receiving respective candidate model instructions as input.

[0009] It has been realized that data-driven models such as large language models can assist in generating chemical product data. In this way, chemical product data can be generated in a relatively simple manner. Furthermore, it has been realized that the increasing availability of such data-driven models, particularly the increasing number of available architectures and parametric variations, can be leveraged to generate more useful chemical product data. For a given received request for generating chemical product data, this can be achieved by determining in the above defined manner a data-driven model and / or at least part of the model instructions to be used for generating the requested chemical product data from a plurality of respective candidates. Thus, by the method presented herein, useful chemical product data can be generated in a relatively simple manner. Due to the significance of chemical product data in the production and / or processing of chemical products, this allows for an improved production and / or processing of chemical products. This improvement in the production and / or processing may be mediated via an improved monitoring and / or controlling of the production and / or processing by virtue of the chemical product data.

[0010] The model instructions used for generating the requested chemical product data may include or refer to the received request. The part of the model instructions being determined from a plurality of candidates may correspond to a part not including a request for generating chemical product data. For instance, the part of the model instructions being determined from a plurality of candidates may include a reference to a request for generating chemical product data. Such a part of model instructions may subsequently be used with any of a plurality of different requests, and may therefore also be regarded as a universal model instructions part.

[0011] Chemical product data associated with a chemical product can be indicative of the chemical product and / or a property of the chemical product. Accordingly, for instance, the received request for generating chemical product data associated with a chemical product can include an indication of a chemical structure associated with the chemical product and / or of a property of the chemical product. More particularly, the received request can include a representation of the chemical product indicative of a chemical structure associated with the chemical product and / or a property of the chemical product. However, the request may also indicate the chemical product differently, particularly without reference to its chemical structure and / or any of its properties. The indication of the chemical product included in the received request may be an indication including measurement data associated with the chemical product. The request can be received via a user interface. Similarly, the chemical product data generated as output of the method can be provided via the user interface.

[0012] Additionally or alternatively, the received request may be based on chemical product data comprising data obtained from sensor data acquired by producing and / or processing the chemical product. Hence, the request may be generated from within a production and / or processing site. Accordingly, but irrespective of an origin of the received request, the method presented herein may particularly be a method for generating chemical product data associated with a production and / or a processing of a chemical product. The chemical product data provided as output of the method may hence, in particular, be chemical product data for producing and / or processing the chemical product. The output of the method may thus, in particular, be used for producing and / or processing the chemical product. For instance, the chemical product data provided as output of the method may be chemical product data adapted to be used in a monitoring and / or controlling of a production and / or processing of the chemical product. To illustrate this further, the generated chemical product data may be provided in a data format adapted to be received as data input by a monitoring device, a control device, a production device and / or a processing device, wherein the respective device may be used for producing and / or processing the chemical product, and / or for monitoring and / or controlling the production and / or processing of the chemical product. The generated chemical product data may be provided to an interface of a respective device. Similarly, as indicated above, the generated chemical product data may be provided to a user interface, wherein a user may operate a respective monitoring device, control device, production device and / or a processing device based on the provided generated chemical product data. Where a monitoring device is operated by a user, an interface to a monitoring device may be regarded as a user interface.

[0013] While the plurality of candidate data-driven models may be trained and / or parameterized to generate chemical product data as output in response to receiving model instructions for doing so as input, this is not the only option. For instance, the plurality of candidate data- driven models may be configured to generate, and / or suitable for generating, chemical product data as output in response to receiving model instructions for doing so as input. In other words, a training and / or parameterization of the data-driven models may refer to a configuration and / or suitability, and vice versa.

[0014] The term “plurality” is understood as referring to two or more instances of the respective entity. Hence, for instance, the plurality of candidate data-driven models corresponds to two or more candidate data-driven models. Similarly, the candidate model instructions can be two or more model instructions.

[0015] Determining, from the plurality of candidate data-driven models, a data-driven model to be used for generating the requested chemical product data can also be referred to as a selection. The data-driven model used for generating the requested chemical product data may be determined, i.e. selected, based particularly on an assessment in the form of a comparison of the outputs generated by the plurality of candidate data-driven models.

[0016] The model instructions used for determining the selected data-driven model from the candidate data-driven models may be the same for all candidate data-driven models. Hence, the same model instructions may be provided as input to each of the candidate data-driven models, wherein the respective outputs of the candidate data-driven models, which will typically differ from each other, may be compared. The outputs of the candidate data-driven models may also be referred to as candidate data-driven model outputs. The candidate data-driven model outputs may be compared to each other and / or to a reference, which may also be regarded as a target output.

[0017] In particular, the model instructions used for determining, from the candidate data-driven models, the data-driven model used for generating the requested chemical product data, can be historical model instructions for which chemical product data were already generated in the past. The chemical product data which were generated for the historical model instructions may be referred to as known or verified chemical product data. Moreover, they may be used as target output for determining the data-driven model used for generating the requested chemical product data. For instance, the historical model instructions may be provided as input to each of the candidate data-driven models, wherein the chemical product data provided by the respective candidate data-driven models as output may be compared to a target output, wherein the target output may correspond to the known or verified chemical product data associated with the historical model instructions.

[0018] The chemical product data provided as output by each of the candidate data-driven models may be compared individually to the target output, wherein the individual comparison results may be compared to each other. Based on the comparison of the individual comparison results, which could also be understood as a comparison of the candidate data-driven models, one of the candidates may be chosen to be used for generating the requested chemical product data. For instance, the comparison of the individual comparison results may be conducted such that it results in a ranking, wherein the candidate data-driven model that provided the top-ranked chemical product data may be chosen to be used. A ranking of this type may also be provided by use of a data-driven model.

[0019] In one embodiment, the outputs provided by the candidate data-driven models are provided via a user interface. An assessment of the outputs provided by the candidate data-driven models may then the carried out based on one or more respective outputs provided via the user interface.

[0020] In another embodiment, the assessment of the candidate data-driven model outputs is carried out using a data-driven model trained and / or parametrized to generate, as output, a ranking of chemical product data upon receiving, as input, model instructions for doing so as input, wherein the model instructions are based on the chemical product data to be ranked and a request for generating the chemical product data, and wherein the ranking is indicative of a correspondence between the request and the respective chemical product data. This has been found to be an efficient way of selecting a suitable one of the candidate data-driven models for generating the requested chemical product data.

[0021] In particular, it may be preferred that the candidate data-driven models are each provided with an input being or corresponding to the initially received request, i.e., the request for which chemical product are to be generated by the method. This request may be supplemented by (further) model instructions to form the input for the candidate data-driven models. These (further) model instructions may be historical, particularly verified, model instructions, for instance, but could also be newly defined. Thus, the selection of a candidate model to be used does not require a historical request. The outputs from the candidate models and the subsequent ranking of the outputs by one or more of the candidate models can be sufficient to decide which candidate model should be used, even though the outputs have been generated based on the request at hand, in response to which chemical product data still need to be generated, i.e., using the selected one of the plurality of candidates.

[0022] If the chemical product data to be ranked have been generated by the candidate data- driven models based on a historical request, this historical request may be used as a basis for the model instructions, i.e., the model instructions by which the data-driven model used for the ranking is instructed. The correspondence between the request and the respective chemical product data indicated by the generated ranking may indicate a similarity between the respective chemical product data and target chemical product data associated with the historical request. The similarity may be measured by a similarity score, for instance. Two sets of chemical product data, which may also be regarded as two data points in a space of possible chemical product data, may be considered as corresponding to each other if a similarity score determined for them is within a predefined range, particularly above a predetermined threshold.

[0023] Hence, for instance, in order to determine which of the candidate data-driven models to use for generating the requested chemical product data, a historical request and verified chemical product data associated with the historical request may be considered. Based on the historical request and / or associated historical model instructions, the candidate data- driven models may be instructed to generate respective chemical product data, wherein these chemical product data, which have been generated by respective candidate data- driven models, may be determined to correspond to the historical request if they are sufficiently similar to the verified chemical product data as measured using a similarity score.

[0024] In particular, each of the candidate data-driven models may be used to generate a respective ranking. Generally, the candidate data-driven model may be suitable for assessing, i.e., determining an assessment of, outputs generated them. The candidate data-driven models may thus be trained and / or parametrized for two purposes. More generally, they may be general-purpose models such as large language models.

[0025] If several rankings are generated, such as one ranking by each of the candidate data-driven models, for instance, overall top-ranked chemical product data may be determined, wherein the candidate data-driven model that provided the top-ranked chemical product data as output may be selected to be used for generating the requested chemical product data. For instance, an average ranking may be considered for determining the top-ranked set of chemical product data.

[0026] In one embodiment, the plurality of candidate data-driven models include data-driven models of different architectures and / or data-driven models of the same architecture with different parameters. Generally, determining the data-driven model used for generating the requested chemical product data, i.e., the selection of the candidate data-driven model that is to be used for this purpose, may refer to determining a) a model architecture and / or b) model parameters.

[0027] The capabilities of data-driven models can vary not only from architecture to architecture, but also with the parameters of a model having a fixed architecture. Considering variations in both therefore allows for determining a particularly suitable model for generating the requested chemical product data.

[0028] Parameters may refer to numerical values defining the model and / or the data processing by the model. Parameters are understood as including weights.

[0029] An architecture of a model is understood as collectively referring to all characteristics of the model that cannot be described in terms of parameters. In other words, two models are understood as having different architectures if their differences cannot be described just by a change of parameters. Transformers, recurrent neural networks (RNNs) and long-short- term-memory (LSTM) architectures are, for instance, understood as being different architectures.

[0030] Additionally or alternatively to the above understanding, an architecture of a model may be defined as being indicative of a data processing of the input data according to the parameters. Similarly, an architecture of a model may be understood as defining the data processing of the input data according to the parameters.

[0031] In one embodiment, the plurality of candidate data-driven models include data-driven models of different architectures and, per architecture, data-driven models with different parameters, wherein the candidate data-driven models are also trained and / or parametrized to generate, as output, a ranking of chemical product data upon receiving, as input, model instructions for doing so as input, wherein the model instructions are based on the chemical product data to be ranked and a request for generating the chemical product data, wherein the ranking is indicative of a correspondence between the request and the respective chemical product data, and wherein the assessment of the outputs generated by data- driven models of the same architecture is carried out using one of these models for generating a ranking of the respective outputs.

[0032] Hence, the outputs generated by different candidate data-driven models of a same architecture are ranked by one of these models, wherein a ranking is generated for each model architecture. An overall ranking of the outputs may then be considered as indicated above for selecting one of the plurality of candidate data-driven models to be used for generating the requested chemical product data. In this way, data-driven models delivering high-quality chemical product data can be found.

[0033] In addition or as an alternative to determining the data-driven model used for generating the requested chemical product data, at least part of the model instructions used for generating the requested chemical product data may be determined. Hence, while it is possible that the model instructions are partially or entirely just received and this received part or entirety of the model instructions is used for instructing a data-driven model that is determined from a plurality of candidates, it is also possible that the data-driven model is just received without having to determine it from a plurality of candidates, wherein the model instructions used for instructing this received model are partially or entirely determined from a plurality of respective candidates. Both options have been found to allow for simplifying the process of generating useful chemical product data associated with a chemical product. This is particularly the case if both options are combined, i.e., when both a selection of the data-driven model and a selection of the part or entirety of the model instructions used for instructing the selected data-driven model is made.

[0034] In one embodiment, the candidate model instructions are determined based on historical requests and / or verified chemical product data, wherein the assessment of the outputs includes a comparison of the outputs to verified chemical product data.

[0035] Hence, the determined part of the model instructions used for generating the requested chemical product data may be determined based on a plurality of historical requests for generating chemical product data and a plurality of verified chemical product data. The verified chemical product data may be associated with the historical requests. For instance, they may correspond to verified responses to respective historical requests. In particular, the candidate model instructions may be determined based on first historical requests and / or first verified chemical product data, and the outputs (i.e., the model outputs generated in response to the candidate model instruction) may be compared to second verified chemical product data. The second verified chemical product data may be different from the first chemical product data, and may be associated with second historical requests different from the first historical requests. This can allow for a more reliable assessment of the outputs and hence the candidate model instructions.

[0036] In one embodiment, the candidate model instructions are determined by using a data- driven model trained and / or parametrized to generate, in response to receiving chemical product data as input, model instructions for instructing a data-driven model to generate chemical product data as output, wherein the model is provided with verified chemical product data as input to generate the candidate model instructions. In this way, promising candidate model instructions can be generated in an efficient manner.

[0037] Since, in the context of the present disclosure, data-driven models are mainly used for generating chemical product data upon receiving model instructions, using a data-driven model for generating model instructions upon receiving chemical product data may also be regarded as a reverse use of a data-driven model. Since instructing a data-driven model to provide an output may also be viewed as prompting the model, the reverse use may also be referred to as reverse prompting.

[0038] The data-driven model used to determine the candidate model instructions may particularly be the data-driven model to be used for generating the requested chemical product data. Hence, for instance, first a data-driven model to be used for generating the requested chemical product data may be determined from a plurality of candidate models, and subsequently the determined data-driven model may be used for determining candidate model instructions. In this case, the data-driven model therefore is trained and / or parametrized for a further purpose. As already indicated above regarding the candidate data-driven models, the data-driven model may be a general-purpose model such as a large language model. When using the data-driven model to be used for generating the requested chemical product data for determining the candidate model instructions (irrespectively of whether it is determined from a plurality of candidate models or not) the candidate model instructions can be tailored to their subsequent use from the start. In this way, yet more suitable chemical product data can be generated.

[0039] In one embodiment, the model instructions further specify a respective request for generating chemical product data for the respective data-driven model. If instructed using such model instructions, more useful model outputs can be expected. Providing just an arbitrary request for generating chemical product data, which may correspond to a free-text input of a user, to a data-driven model such as a large language model can generally be expected to lead to non-optimal outputs of the model. Conversely, using model instructions further specifying a respective request, the request itself can be less specific. Requests can hence be provided more easily by users. For instance, the model instructions may refer to text indicating background information associated with a request.

[0040] The model instructions may also indicate one or more elements of the chemical product data to be generated. For instance, the model instructions may be indicative of a selection of one or more elements for generating an element of the chemical product data. In particular, the model instructions may be indicative of a temperature of a large language model used and / or its creativity parameter. The real instructions may also indicate a length of the output by specifying, for instance, a token number.

[0041] Model instructions for instructing a data-driven model to generate chemical product data may be suitable for instructing the whole data-driven model or only a part thereof to generate the chemical product data. The data-driven models considered herein may hence comprise more than one part, wherein a part of the model may be sufficient for generating chemical product data as output in response to receiving model instructions as input.

[0042] In one embodiment, the received request comprises a sequence of a plurality of elements, and the data-driven model used for generating the chemical product data is configured, trained and / or parametrized to receive, and / or suitable for receiving, a sequence of a plurality of elements and generate a sequence of a plurality of further elements according to the received sequence of elements. Preferably, the generated chemical product data continue the sequence associated with the request.

[0043] Chemical product data can be defined to continue a sequence associated with a request if they correspond to a sequence indicative of information demanded by the request. For instance, if the sequence associated with the request is a question, an answer to the question could be considered as a sequence continuing the sequence associated with the request, i.e., the question. In particular, a specific parameter value may be requested, in which case a model output may continue the request and / or a sequence associated therewith by providing the requested parameter value.

[0044] In this way, the generated chemical product data can correspond to the request. Chemical product data tailored to the request can be accurately generated. Therefore, precise chemical product data can be generated. Ultimately, this can improve producing and / or processing of the chemical product based on the generated chemical product data. An element may comprise at least a part of a word, a numerical value, a unit of the numerical value, at least a part of a table or a combination thereof. Additionally or alternatively, the sequence of two or more elements may refer to at least a part of a sentence, at least a part of a number sequence, at least a part of a table or a combination thereof. A part of a table may refer to a row number, a column number and / or a table entry.

[0045] The term “property” may refer to a physical property, to a chemical property, a biological property or a combination thereof. A chemical property may be a property that can be established only by changing one or more chemical structures associated with the at least one chemical product. Examples for chemical properties may be acidity, oxidation state or reactivity. A physical property may be one of the following: mechanical properties, electrical properties, optical properties, thermal properties or the like. For example, “physical property” may comprise one or more of the following density, scratch resistance, electrical conductivity, color, absorption, heat capacity or the like. Biological properties may refer to toxicity, biological activity, biodegradability, bioaccumulation or the like.

[0046] The term “chemical product” may refer to a product obtained by means of a chemical production process. “Chemical production process” may refer to a process including one or more chemical reaction(s). The chemical product may be characterized by at least one functional group.

[0047] In one embodiment, the data-driven model used for generating the chemical product data is a pretrained data-driven model. Furthermore, the data-driven model may be a fine-tuned data-driven model. Thus, for instance, the plurality of candidate data-driven models may include, possibly but not necessarily exclusively, data-driven models that are pre-trained and optionally also fine-tuned, wherein, from the candidate data-driven models, a pretrained and optionally also fine-tuned model may be selected to be used for generating the chemical product data.

[0048] The pretrained data-driven model(s) may be parametrized and / or trained based on data with a plurality of contexts and / or unstructured data, in particular text data and optionally numerical data such as tabular data or image data. The pretrained data-driven model(s) may be configured to perform a plurality of task and / or to process data of a plurality of contexts. The pretrained data-driven model(s) may be configured to perform the task according to the provided task instruction. Hence, the pretrained data-driven model may be configured to be provided with a plurality of different task instructions and / or provide a plurality of different types of output data upon receiving different task instructions. The fine-tuned data-driven model(s) may be obtained by training pretrained data-driven model(s) configured to perform a plurality of tasks according to a plurality of task instructions. The fine-tuned data-driven model(s) may trained additionally on a training data set comprising a plurality of task instructions of one type and corresponding output data. The fine-tuned data-driven model may be trained additionally to provide output data of a predefined type according to the training data set. The fine-tuned data-driven model may be configured to be provided with a plurality of different task instructions and / or provide a plurality of different types of output data upon receiving different types of task instructions. Further, the fine-tuned data-driven model may be configured for providing one type of output data upon receiving one type of task instruction with a higher accuracy than providing other types of output data upon receiving other types of task instructions.

[0049] In one embodiment, the method further includes a step of retrieving an indication of a chemical structure and / or a property of the chemical product based on the indication of the chemical product included in the received request by providing the indication of the chemical product to a data source. The data source may associate indications of chemical products with indications of their respective chemical structures and / or properties.

[0050] In one embodiment, the providing of the indication of the chemical product to the data source comprises providing a query obtained from the indication of the chemical product included in the received request, wherein the chemical structure and / or the property of the chemical product is retrieved according to the query.

[0051] In one embodiment, the providing of the indication of the chemical product to the data source comprises providing a numerical representation of the indication of the chemical product, wherein the chemical structure and / or the property of the chemical product is retrieved according to a similarity score associated with a) the numerical representation of the indication of the chemical product and b) a numerical representation of indications of chemical structures and / or properties of chemical products retrievable from the data source.

[0052] The present disclosure also relates, in an aspect, to a method for further training a pretrained data-driven model, wherein the method includes: providing the pre-trained data-driven model, providing training data including pairs of training input data and training output data, wherein the training input data are associated with model instructions based on requests for generating chemical product data associated with a respective chemical product, and the training output data are associated with verified chemical product data, and training the pre-trained data-driven model further using the training data such that the further trained data-driven model is trained and / or parameterized to generate chemical product data as output in response to receiving model instructions for doing so as input.

[0053] The further training of a pre-trained data-driven model may also be understood as a fine- tuning of the pre-trained data-driven model. The above method may hence also be viewed as a method for fine-tuning a pre-trained data-driven model.

[0054] The verified chemical product data with which the training output data are associated may refer to historical chemical product data, particularly historical chemical product data that are considered appropriate outputs for the respective model instructions with which the training input data are associated.

[0055] The use of the training data for the further training, or fine-tuning, of the pre-trained data- driven model may refer to providing at least part of the training input data as input to the data-driven model and evaluating the respective outputs of the data-driven model with respect to the training output data associated with the used training input data. The evaluation may be carried out using an evaluation function such as a loss function. A result of the evaluation may be used for adapting parameters of the data-driven model. This process may be repeated until a respective result of the evaluation is considered acceptable, such as until a value of the evaluation function is below a predetermined threshold.

[0056] In a further aspect of the present disclosure, a data processing system comprising a processor configured to carry out the steps of a method as defined above is presented. Thus, in other words, the method defined above is preferably computer-implemented.

[0057] A further aspect of the present disclosure relates to a use of a method as defined above and / or the above-indicated data processing system for generating chemical product data. Also disclosed herewith is a use of the generated chemical product data, particularly for monitoring and / or controlling a production and / or processing of the chemical product. Accordingly, another aspect of the present disclosure concerns a method for monitoring and / or controlling production and / or processing of a chemical product based on chemical product data associated with the chemical product, wherein the chemical product data are generated as outlined above. Thus, the chemical product data based on which the production and / or processing of the chemical product is monitored and / or controlled according to this further aspect is generated by: receiving a request for generating chemical product data associated with the chemical product, the request including an indication of the chemical product, generating the requested chemical product data by providing model instructions based on the received request as input to a data-driven model trained and / or parametrized to generate chemical product data as output in response to receiving model instructions for doing so as input, and providing the generated chemical product data as output of the method, particularly for monitoring and / or controlling production and / or processing of the chemical product, wherein: a) the data-driven model used for generating the requested chemical product data is determined from a plurality of candidate data-driven models, wherein the plurality of candidate data-driven models are trained and / or parametrized to generate chemical product data as output in response to receiving model instructions for doing so as input, wherein the data-driven model used for generating the requested chemical product data is determined based on an assessment of outputs generated by the plurality of candidate data-driven models in response to receiving, as input, model instructions based on one or more historical requests for generating chemical product data, and / or b) at least part of the model instructions used for generating the requested chemical product data are determined based on an assessment of outputs generated by a data- driven model, which is trained and / or parametrized to generate chemical product data as output in response to receiving model instructions for doing so as input, in response to receiving respective candidate model instructions as input. Optionally, the generated chemical product data are provided, in particular for monitoring and / or controlling the production and / or processing of the chemical product. Also a data processing system comprising a processor configured to carry out the steps of the method for monitoring and / or controlling the production and / or processing of the chemical product is disclosed. Thus, also the disclosed method for monitoring and / or controlling is preferably a computer-implemented.

[0058] It shall be understood that the method of claim 1 , the system of claim 14 and the use according to claim 15 have similar and / or identical preferred embodiments as defined in the dependent claims.

[0059] It shall be understood that a preferred embodiment of the invention can also be any combination of the dependent claims with the respective independent claim.

[0060] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter.

[0061] In the following, the present disclosure is further described with reference to the enclosed figures. The same reference numbers in the drawings and this disclosure are intended to refer to the same or like elements, components, and / or parts.

[0062] BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In the following, the present disclosure is further described with reference to the enclosed figures. The same reference numbers in the drawings and this disclosure are intended to refer to the same or like elements, components, and / or parts.

[0064] FIG. 1 illustrates an embodiment of a method for generating chemical product data.

[0065] FIG. 2 illustrates a selection of a data-driven model to be used for generating chemical product data according to an embodiment.

[0066] FIG.3A illustrates a selection of model instructions to be used for generating chemical product data according to an embodiment.

[0067] FIG. 3B illustrates a combined selection of a data-driven model and model instructions to be used for generating chemical product data according to an embodiment. FIG. 4 illustrates a data processing system for generating chemical product data, particularly by carrying out a model selection as illustrated in FIG. 2, according to an embodiment.

[0068] FIG. 5 illustrates a data processing system for generating chemical product data, particularly by carrying out a model instructions selection as illustrated in FIG. 3A, according to an embodiment.

[0069] FIG. 6 illustrates an embodiment of a training of an embedding layer.

[0070] FIG. 7 illustrates an embodiment of a transformer encoder architecture.

[0071] FIG. 8 illustrates an embodiment of a transformer decoder architecture.

[0072] FIG. 9 illustrates an embodiment of a transformer encoder-decoder architecture.

[0073] FIG. 10 illustrates an embodiment of training and / or deploying the transformer encoder, the transformer decoder and / or the transformer encoder-decoder.

[0074] FIG. 11 illustrates an embodiment of input embedding.

[0075] FIG. 12 illustrates a further embodiment of input embedding.

[0076] DETAILED DESCRIPTION OF EMBODIMENTS

[0077] The following embodiments are mere examples for implementing the method and system disclosed herein and shall not be considered limiting.

[0078] FIG. 1 shows schematically and exemplarily a method 100 for generating chemical product data associated with a chemical product.

[0079] In a first step 101 of the method 100, a request for generating chemical product data associated with the chemical product is received. The request includes an indication of the chemical product. This indication may be an indication of a chemical structure associated with the chemical product and / or of a property of the chemical product. However, the chemical structure and / or properties of the chemical product may also be retrieved separately based on another type of indication of the chemical product included in the received request.

[0080] In steps 102a, 102b of the method 100, which may jointly be understood as a step 102, the requested chemical product data are generated by providing, in step 102a, model instructions based on the received request as input to a data-driven model trained and / or parameterized to generate chemical product data as output in response to receiving model instructions for doing so as input. The generated chemical product data, which are received in step 102b from the data-driven model, are, in a subsequent step 103, provided as output of the method 100. The generated chemical product data may particularly be provided in step 103 for monitoring and / or controlling a production and / or processing of the chemical product.

[0081] As shown in FIG. 1 , the data-driven model used for generating the chemical product data and / or the model instructions used as model input for generating the chemical product data may be selected from respective candidates, in the course of the method 100.

[0082] Considering first the option of selecting the data-driven model to be used, i.e., the left branch in FIG. 1 , it is shown that, in a step 104a-1 , a historical request and / or associated model instructions is / are provided as input to a plurality of candidate data-driven models. The candidate data-driven models are each trained and / or parameterized generate chemical product data as output in response to receiving model instructions for doing so as input. For instance, model instructions including a historical request for generating chemical product data may be provided to each of the candidate data-driven models, wherein the historical request includes an indication of the chemical product for which chemical product data are to be generated according to the request.

[0083] In a subsequent step 104a-2, the outputs generated by the candidate data-driven models in response to step 104a-1 are received. Hence, from each candidate model, the chemical product data generated by this candidate model based on the historical request are received.

[0084] In a further step 104a-3, an assessment of the outputs from the plurality of candidate data- driven models is carried out. The assessment may refer to a comparison of the plurality of outputs a) to verified chemical product data associated with the historical request and / or b) with each other. This will be described in more detail further below. As indicated as step 104a-4, the data-driven model to be used for generating the chemical product data requested in terms of the request received in step 101 is selected from the plurality of candidate data models based on the assessment carried out in step 104a-3. The selected model is then used in step 102a.

[0085] Considering the option of selecting the model instructions to be used in step 102a, i.e., the right branch in FIG. 1 , it is shown that, in a step 104b-3, for instance, a plurality of candidate model instructions based on respective historical requests may be provided as a respective input to a data-driven model. Also this model may be trained and / or parameterized to generate chemical product data as output in response to receiving model instructions for doing so is input. In fact, the model used in step 104b-3 may be, for instance, the model selected in step 104a-4. The candidate model instructions may each comprise a respective historical request and a candidate instructions part further specifying the historical request.

[0086] In a subsequent step 104b-4, the plurality of sets of chemical product data generated as outputs by the data-driven model in step 104b-3 may be received, and thereafter an assessment of these outputs may be carried out in a step 104b-5. For instance, the assessment may involve a comparison of the outputs to the verified chemical product data associated with the historical requests on which the candidate model instructions were based. In step 104b-6, model instructions to be used for generating the initially requested chemical product data in step 102a are selected from the plurality of candidate model instructions based on the assessment, in this case the comparison, carried out in step 104b-5. In particular, the selection may refer to the part of the candidate model instructions further specifying the respective request.

[0087] As also indicated in FIG. 1 , the candidate model instructions used in step 104b-3 may, for instance, be generated by use of a reverse trained data-driven model, i.e., a model trained and / or parameterized generate, in response to receiving chemical product data and optionally also a request for generating the same as input, model instructions for instructing a data-driven model to generate chemical product data as output. In a step 104b- 1 , verified chemical product data and optionally also historical requests associated therewith may be provided as input to the reverse trained data-driven model. The verified chemical product data and the optionally used historical requests on which the candidate model instructions are determined may be different from the verified chemical product data and the associated historical requests on which the assessment of the candidate model instructions or instruction parts is based. This will be further clarified with reference to FIG. 3A. Based on the received verified chemical product data, the reverse-trained model generates candidate model instructions as output, which are received in step 104b-2 and thereafter used as model inputs in step 104b-3. The models used in steps 104b- 1 and 104b-3 may be the same or different from another.

[0088] The verified chemical product data and the optionally used historical requests on which the candidate model instructions are determined may be different from the verified chemical product data and the associated historical requests on which the assessment of the candidate model instructions or instruction parts are based. This will be further clarified with reference to FIG. 3A.

[0089] FIG. 2 illustrates schematically and exemplarily how a model that is to be used for generating the requested chemical product data can be selected from a plurality of candidate models. In other words, FIG. 2 illustrates a particular realization of the left branch of FIG. 1 .

[0090] According to FIG. 2, the plurality of candidate data-driven models include data-driven models of three different architectures and, for each architecture, data-driven models with three different sets of parameters. Moreover, all candidate data-driven models are large language models in this case.

[0091] Each of the in total nine candidate models considered according to FIG. 2 is provided with the same input, wherein the input corresponds to a request for generating chemical product data and / or associated model instructions, particularly to model instructions based on the request received in step 101. Hence, for instance, the request received in step 101 , which may be provided by a user, may be used in the same way as input to all candidate data- driven models. If it is supplemented by (further) model instructions for forming an input, also these further model instructions may be the same for each of the candidate data-driven models. Due to the differences in architecture and / or parameters, each candidate model generates, in response to the input, a different output, such that in total also nine different outputs are generated.

[0092] Even if just pre-trained, but particularly if they have been fine-tuned in a corresponding supplementary training, each of the candidate data-driven models is in this case trained and / or parameterized to generate, as output, a ranking of chemical product data upon receiving, as input, model instructions for doing so as input. Hence, the candidate models can be used themselves to rank the chemical product data responding to the nine different outputs. In particular, a model of a first of the architectures considered may be used for generating a first ranking of the nine different outputs, a model of a second of the architectures considered may be used for generating a second ranking of the nine different outputs, and a model of a third of the architectures considered may be used for generating a third ranking of the nine different outputs. This is the case indicated in FIG. 2, wherein only the respective four top-ranked outputs are shown for each of the three generated rankings.

[0093] The model instructions used for causing the respective candidate model to generate the respective ranking are based on the nine different sets of output chemical product data to be ranked and the request that was used as input for generating the chemical product data. Accordingly, the rankings are indicative of a correspondence between the request and the respective chemical product data.

[0094] Based on the nine generated rankings, a final ranking of the nine outputs may be generated, wherein the candidate model that generated the top-ranked output in this final ranking may be selected to be used for generating the initially requested chemical product data. The final ranking may, for instance, correspond to an average of the previous rankings.

[0095] FIG. 3A illustrates schematically and exemplarily how model instructions or a part thereof to be used for generating the requested chemical product data can be selected from a plurality of candidates. In other words, FIG. 3A illustrates a particular realization of the right branch of FIG. 1 .

[0096] In FIG. 3A, access to a plurality of historical requests for generating chemical product data and associated verified chemical product data is assumed. For illustrative purposes, access to eight pairs of a) historical requests (i.e., model inputs) and b) verified chemical product data (i.e., model outputs) is assumed, wherein four of the pairs are considered as training pairs and the remaining four of the pairs are considered as test pairs. A further, ninth request is assumed to correspond to the request received in step 101 of the method 100, wherein this ninth request is to be processed to generate chemical product data indicated as “Output 9” in FIG. 3A. From the training pairs of historical requests and associated verified chemical product data, the latter, i.e. the verified chemical product data, or training outputs, are used in step 104b-1 . In the illustrated case, the reverse-trained model to which they are provided as input may be the LLM that is also to be used, in step 102a, for generating the initially requested chemical product data, indicated in FIG. 3A as the ninth output. Moreover, to this same LLM, the candidate instructions generated by the LLM as output in response to receiving the verified chemical product data as input may be provided as model input in step 104b-3. The outputs generated in response thereto by the LLM may be assessed in order to select the model instructions to be used in step 101. In a variant, the LLM used for generating the candidate model instructions from the verified chemical product data in the training pairs may be parameterized using the model parameters that were also used for generating the verified chemical product data from the respective historical requests. For this purpose, the historical data being accessed may not only include pairs of historical requests and associated verified chemical product data, but also historical model parameters associated with these pairs. The historical model parameters associated with a historical request and associated verified chemical product data may be the model parameters of the respective data driven-model that was used for generating, based on the historical requests as input, the verified chemical product data as output.

[0097] In particular, the test pairs of a) historical requests and b) verified chemical product data can be used for the assessment of the candidate model instructions. For instance, it may be preferred that the model instructions further specify a respective request for generating chemical product data for the LLM. The candidate model instructions may then be provided together with, or based on, the historical requests of the test pairs as input to the LLM. As parameters of the LLM, those that were also used for generating the verified chemical product data from the associated historical requests in the test pairs may be used. The resulting outputs generated by the LLM, i.e., the output chemical product data, can then be compared, in step 104b-5, to the respective verified chemical product data of the test pairs. The candidate model instructions that lead to a highest similarity between the generated chemical products outputs and the verified chemical product data may be selected to be used in step 103.

[0098] FIG. 3B illustrates an embodiment corresponding to a variant of the embodiment illustrated by FIG. 3A. In this variant, the historical requests and associated verified chemical product data are not only used to determine the model instructions that are to be used for generating the requested chemical product data as output of the method 100, but also to select the data-driven model that is to be used for this purpose, particularly its model parameters. Hence, this variant could be understood as a variant according to which a combined selection of a data-driven model and model instructions is carried out.

[0099] In an embodiment according to FIG. 3B, initially a plurality of candidate data-driven models and candidate model instructions may be provided. As may also be the case in embodiments according to FIG. 2, the candidate data-driven models may include models of different architectures and / or with different model parameters. Moreover, as may also be the case in embodiments according to FIG. 3A, the candidate model instructions may include a plurality of different sequences that, when provided together with a sequence corresponding to a request to generate chemical product data as input to a data driven model, cause the data-driven model to generate an output sequence.

[0100] Furthermore, as already explained with reference to FIG. 3A, the historical data may be split into training data and test data. While, however, FIG. 3A showed the option of including also model parameters as historical data, this option, which was already just that in embodiments according to FIG. 3A, may particularly not be realized in an embodiment according to FIG. 3B and is therefore also not shown in FIG. 3B. In particular, in a step 104c-1 , just the plurality of historical requests and associated verified chemical product data may be used as training data for determining a data-driven model and model instructions that are to be used for generating the requested chemical product data to be provided as output of the method 100.

[0101] The training carried out in step 104c-1 can generally follow known principles of machine learning according to which, for instance, outputs generated by a model under training upon being provided with training inputs are compared to training outputs associated with the training inputs in order to adapt the model based on the comparison, wherein this procedure is carried out repetitively, leading to successive adaptations of the model and corresponding successively changed outputs of the model in response to the training inputs, wherein the successive adaptations of the model are carried out such that the outputs generated by the model in response of the training inputs become successively more similar to the training inputs until a predefined similarity threshold is reached, at which point the model may be considered trained. However, this principle may be extended in step 104c-1 to the model instructions. Namely, not only the model may be adapted from a training step (or epoch) to the next, but also the model instructions. For instance, instead of using fixed model instructions together with the respective training inputs as combined training inputs during the training, the model instructions may be adapted from a training step (or epoch) to the next. In this way, not only a trained model may result from the training, but also trained model instructions.

[0102] The trained data-driven model and model instructions arising from the combined training carried out in step 104c-1 may be verified using the historical test data. Again, also for this verification only historical requests and associated verified chemical product data may be used as historical data, i.e., no historical model parameters. For instance, the plurality of historical requests may be provided, together with the trained model instructions, as input to the trained data-driven model in a step 104c-2. In response thereto, the trained data- driven model may generate, in a step 104c-3, a plurality of corresponding outputs. These outputs could also be referred to as trained outputs, since they are generated using a data- driven model and model instructions resulting from a training as carried out in step 104c-1 . On the other hand, the trained data-driven model, the trained model instructions and the trained outputs could also be referred to as candidate data-driven model, candidate model instructions and candidate outputs, respectively, since they may still be subject to verification with respect to historical test data as illustrated.

[0103] The verification may be carried out in a step 104c-4 by comparing the verified chemical product data from the historical test data to the outputs provided step 104c-3. If the comparison results in a degree of similarity, which may be encoded in terms of a similarity score, for instance, that is higher than a predefined threshold, the trained data-driven model and the trained model instructions may be considered verified and may therefore be selected to be used for generating the requested chemical product data to be output by the method 100. This case is shown in FIG. 3B, wherein, like in FIG. 3A, the request initially received in step 101 of the method 100 is referred to as a new, ninth request. Since the model instructions and the data-driven model used for processing this new, ninth request have been determined in a manner different from the one illustrated by FIG. 3A and will therefore likely be different from the model instructions and the model used for this purpose according to FIG. 3A, also the generated output will differ. Nevertheless, the output, exemplarily referred to as ninth output, is indicated identically in FIG. 3A and FIG. 3B.

[0104] If the verification step 104c-4 results in the trained data-driven model and the trained model instructions not being verified, the training of step 104c-1 may be resumed or restarted.

[0105] FIG. 4 and FIG. 5 illustrate schematically and exemplarily a data processing system for carrying out the method 100, wherein, apart from elements of the system, also data being processed and / or exchanged by the elements are shown. In particular, FIG. 4 illustrates processing steps taken to select a data-driven model from two candidates, and FIG. 5 illustrates processing steps taken to select model instructions from two candidates, wherein it shall be understood that these process steps may also be combined partially or completely to select both a data-driven model and model instructions.

[0106] As illustrated, the data processing system may comprise a processor 112 and a user interface 108. Furthermore, while not shown, the system may comprise a database for storing a plurality of candidate data-driven models 104, 106, 204, and / or an interface for accessing the same, such as from a remote source via a communication network. Similarly, while also not shown, the system may comprise a database for storing, and / or an interface for accessing, historical requests and associated chemical product data 207, as well as candidate model instructions 209.

[0107] The determination of the data-driven model to be used and the model instructions to be used for generating the requested chemical product data can be carried out separately or in combination. For instance, such determination may be initiated by receiving, via the user interface 108, from the database or via a network interface, a historical request for generating chemical product data and first model instructions for instructing a first data-driven model 104, 204 and / or a second data-driven model 106 to generate chemical product data. Furthermore, second model instructions for instructing the first data-driven model 104, 204 and / or the second data-driven model 106 to generate chemical product data may be received via the user interface 108, from the database or via a network interface. In FIG. 5, the first and the second model instructions are collectively indicated by reference numeral 209.

[0108] Subsequently, as indicated in FIG. 4, the historical request and the first model instructions may be provided to the first data-driven model 104 for generating first chemical product data and to the second data-driven model 106 for generating second chemical product data. Additionally or alternatively, as indicated in FIG. 5, the historical request and the first model instructions may be provided to the first data-driven model 204 for generating first chemical product data, and the second model instructions may be provided to the first data- driven model 204 for generating second chemical product data as well. In a combined selection process for selecting both a data-driven model and model instructions, the second model instructions may be provided to a second data-driven model 206.

[0109] Furthermore, target chemical product data associated with the historical request may be received, wherein it may be determined if the target chemical product data correspond to the first chemical product data. The target chemical product data may be verified in the sense that they are considered a desirable model output for the historical request. Alternatively, the first chemical product data may be selected. In both cases, the first model instructions may be provided to the first data-driven model 104, 204 for generating the requested chemical product data, i.e., the first model instructions and the first data-driven model 104, 204 may be selected for use.

[0110] In particular, the first model instructions may be provided to the first data-driven model 104, 204 for generating the requested chemical product data in response to determining that the first chemical product data correspond to the target chemical product data or in response to receiving the selection of the first chemical product data. If the target chemical product data correspond to the second chemical product data, the second data-driven model 106 or the second model instructions may be determined for use, depending on how the second chemical product data have been generated (i.e. , if they have been generated by providing the first model instructions to the second data-driven model 106, the second data-driven model 106 may be determined to be used, and if they have been generated by providing the second model instructions to the first data-driven model 204, the second model instructions may be determined to be used).

[0111] In the above, the target chemical product data may be received via the user interface 108 and / or form the first data-driven model 104, 204 and / or the second data-driven model 206. Moreover, the determination of whether the target chemical product data correspond to the first chemical product data (or the second chemical product data) may be carried out by providing the target chemical product data and the first chemical product data to the first data-driven model 104, 204 and / or the second data-driven model 206. Furthermore, the first model instructions and / or the second model instructions may be provided to the first data-driven model 104, 204 and / or the second data-driven model 206. Alternatively, the target chemical product data and the first chemical product data (and / or the second chemical product data) may be provided to a third data-driven model, which may particularly be a classification model.

[0112] In the following, particular features of possible data-driven models as considered herein will be described with reference to FIG. 6 to FIG. 12.

[0113] FIG. 6 illustrates an embodiment of obtaining an embedding layer usable in a data- driven model. The embedding layer may be obtained by training for example a continuous bag of words model (CBOW) or a skip-gram model. The embedding layer may be suitable for generating embedded input data based on input data. Generating embedded input data may refer to embedding input data.

[0114] As is the case throughout the subsequent description of FIG. 6 to FIG. 12, the input data may be unstructured or structured data. For instance, the input data may be or comprise general text, numerical and / or image data. However, more particularly, the input data throughout the subsequent description of FIG. 6 to FIG. 12 may also refer to a request for generating chemical product data and / or model instructions associated with such a request. Such particular input data may comprise text, numerical and / or image data structured in a particular form. An exemplary form would be, for instance, such that the input data split into a part corresponding to the request and a part corresponding to model instructions further specifying how the request is to be handled by the model. In such a case, the model instructions may or may not be specific to the request. A further particular kind of input data may refer to chemical product data. For instance, in a “reverse” use of a model as indicated further above by steps 104b-1 and 104b-2, chemical product data can be used as input data.

[0115] Embedding input data may result in a representation associated with the input data. Thus, the embedded input 114 may be the representation associated with the input data. The input data may comprise one or more elements. The one or more elements may be represented by the input vector 106. In particular, the embedded input 114 and / or the input vector 106 may be machine-readable and / or processable by a processor. For this purpose, the embedded input 114 and / or the input vector 106 may be a tensor, in particular a first-rank tensor. Specifically, the input vector 106 may be a one-hot vector or a summation of a plurality of one-hot vectors. A one-hot vector may be a vector with one entry unequal to zero. Examples for one-hot vectors may be 108, 110 and 112. The entries unequal to zero in the one-hot vector and / or in the input vector 106 may indicate the element. For example, a lookup table may define the relation between the position of the entries unequal to zero and the element indicated by the one-hot vector. The lookup table may specify a plurality of different elements. The number of different elements may be equal to the number of entries in the one- hot vector. The number of different elements may be referred to as vocabulary size. In an example, the elements may be represented by tokens and a sequence of elements may refer to at least a part of a sentence. The at least a part of the sentence may be represented by a plurality of tokens. A token may represent at least a part of the element and / or word. For example, where one element would be associated with only one word, words such as “embeddings", “embedding” or “embed” would constitute different elements. A first token may represent the stem “embed” and the endings, typically appearing in a plurality of words, may be represented by a second token, a third token and a fourth token. The second token, the third token and the fourth token may be used for representing other words such as “look”, “looking” or the like, preferably together with a fifth token representing the stem “look”. Ultimately, this tokeniza- tion of elements associated with a plurality of stems and a plurality of endings results in less tokens to be used for representing a plurality of elements and thus, uses less computational resources. A lookup table specifying a subset of the vocabulary size e.g. of the English language may comprise 10,000 words or more. The embedded input 114 may be a lower-dimensional representation than the input vector 106. For example, typical embedded inputs 114 may comprise some hundreds of different entries. Followingly, the embedded inputs 114 constitute a densified representation of one or more elements using less computational resources. More than that, the embedded input 114 may represent a relation between two or more elements. For example, the words “Italy” and “Germany” may be similar or may be more closely related since they both define European countries, whereas the word “embodiment” may be very different from the two respective words. The smaller the dot product between two embedded inputs 114 may be, the more similar the two elements associated with the embedded inputs 114 may be. Hence, the embedded inputs 114 may represent one or more elements accurately and lead to accurate results based on processing the embedded inputs 114.

[0116] For transforming the input vector 106 into the embedded input 114, the embedding layer may comprise a number of neurons equal to the number of entries in the embedded input 114. Based on the embedded inputs 114, the output layer may generate the output vector 116. The output vector may be a vector and / or may indicate one or more elements. The output vector 116 may indicate one or more elements different from the input vector 106 and / or the one-hot vectors associated with the input vector 106. For this purpose, the output layer may comprise a number of neurons equal to the number of entries of the input vector 106 and / or the output vector 116. The output layer may apply a softmax function to the embedded inputs 114. By doing so, the output vector may comprise the probabilities associated with the elements associated with the entries of the output vector 116 unequal to zero. Hence, from the output vector 116 one or more elements may be obtained with a corresponding probability. Where the input vector 106 may specify one or more sequence(s) of elements, the output vector 116 may specify one or more elements corresponding to the sequence(s) of elements specified by the input vector 106. In the example of FIG. 6, the element associated with vector 118 may correspond to the input vector with a probability of 71 %. Additional or alternative elements may correspond to the input vector as indicated by the output vector with lower probability. By defining a threshold to which the probability may be compared, the selection of the corresponding elements may be tailored to the needs of the user. The elements generated by the model comprising the embedding layer 102 and the output layer 104 may refer to the most probable elements indicated by the output vector 116. Hence, the model depicted in FIG. 6 may generate the element associated with the vector 118 with a confidence score of 71 %. The model of FIG. 6 may be a continuous bag of words (CBOW) model. The CBOW model may be trained based on a training data set comprising a plurality of input vectors and corresponding output vectors. As the training data set may not be labeled, the training of the CBOW model may be referred to as self-supervised. Before training of the CBOW model, the CBOW model may be initialized with random values assigned to the weights of the neurons. During the training of the CBOW model, the input vectors may be passed through the initialized embedding layer and the output layer and a loss may be determined by comparing the output vector obtained by passing the input vector 106 through the model to the output vector corresponding to the input vector 106 as specified by the training data set. Based on the determined loss, back- propagation may be applied to determine the gradients associated with the neurons of the embedding layer 102 and the output layer 104 to lower the loss. According to the determined gradients, the weights of the neurons may be updated by using a gradient descent algorithm. If a predetermined loss may be achieved by the CBOW model, the training may be terminated and a trained CBOW model may be obtained. From the trained CBOW model, the embedding layer 102 may be suitable for embedding input data comprising one or more elements. This embedding layer 102 may be used in other machine-learning architectures requiring an embedding layer 102 such as a transformer encoder, transformer decoder or transformer encoder decoder architecture as described within the context of FIG. 7, FIG. 8 and FIG. 9, all of which are possible architecture for the data-driven models considered herein. For training these architectures, a trained embedding layer 102 may be required. Hence, a model such as a CBOW model may be trained prior to training the transformer encoder, transformer decoder or transformer encoder decoder architecture.

[0117] FIG. 7 illustrates an embodiment of a transformer encoder architecture. The transformer encoder comprises an encoder input 278, one or more encoder blocks 274, 214 and an encoder output. The transformer encoder architecture may be derived from the transformer encoder-decoder architecture as known in the art and shown in FIG. 9. In particular, the transformer encoder may be referred to as X-former. The transformer encoder architecture may correspond to the encoder architecture associated with the transformer encoder-decoder architecture with an additional encoder output instead of connecting the encoder block directly to the decoder of the transformer encoder-decoder architecture. A plurality of transformer encoder architectures are available in the art, such as the bi-directional encoder representations from transformers (BERT). The input data may be received at the encoder input 278. The encoder input 278 may apply an input embedding 202. Applying the input embedding 202 may refer to passing the input data through an embedding layer, e.g. as described within the context of FIG. 6. Further, the encoder input 278 may apply positional encoding 203. Applying positional encoding 203 may refer to adding a positional factor to the embedded input obtained via input embedding. Preferably, the input data may specify a sequence of elements. The positional factor Ppos may be indicative of the position of the elements within the sequence. For example, the positional factor Ppos may be obtained based on the following equation: \ / pos

[0118] Ppos 2« + 1 = cos - —

[0119] \ J \ 1000( / where pos may refer to the position of the element within the sequence, / may refer to the dimension associated with the input embedding and d may refer to the dimension of the model, e.g. transformer decoder, transformer encoder or transformer encoderdecoder. This may be referred to as absolute positional embeddings. Alternatively, the positional encoding may be based on rotary positional embeddings (RoPE). Positional encoding is beneficial since it enables the processing of sequential data without requiring further dimensions indicating the position of each element. Followingly, the positional encoding 203 reduces the computational resources needed for embedding the input data. By passing the input data through the encoder input, the input data may be transformed into a second-rank tensor representing the sequence of elements. This second-rank tensor may be referred to as embedded input data. The embedded input data may be processed by the encoder block. The embedded input data may be provided to the layer normalization 208 by a residual connection. Multi-head self-attention 206 may be applied to the embedded input data. Multi-head self-attention 206 may comprise the two components multi-head and self-attention. Self-attention may be understood as being a filter applied to the embedded input data. By applying the filter to the embedded input data, the elements associated with the embedded input data contributing to the to be generated output data may be identified for generating the output data. Hence, the filter may represent the degree of contributing to the to be generated output data by the elements associated with the embedded input data. Applying the filter may be referred to as weighting the elements associated with the em- bedded input data. This is advantageous specifically regarding long sequences of elements. The filter may be learned and improved during the training by learning to identify the contribution of elements associated with the embedded input data. For example, in the partial sentence “I went to the bakery to buy a” the last word may be generated by the data-driven model such as the transformer encoder. The self-attention may focus the transformer encoder to attend to the word “bakery” and “buy” mostly to generate the word “bread”. Self-attention may refer to attention generated based on the input data. Hence, the filter may be determined based on the input data, preferably the embedded input data. The embedded input data may serve as query Q, key K and value V with respect to the self-attention operation. The self-attention may refer to attention based on the received input data. Hence, the filter may be calculated based on the following formula by inserting the respective tensors based on the embedded input data: where corresponds to the dimension of the key.

[0120] For improving the efficiency of the transformer encoder further, the multiple heads are used to apply the filter resulting in the multi-head self-attention 206. Multi-head selfattention 206 may comprise applying the filter to two or more parts of the embedded input data. Hence, the tensor may be split into two or more parts and the filter may be applied to the two or more parts separately by two or more heads according to the following equation: head i = Attenti(m(QWiQ, KWiK, VWiv) with parameter matrices

[0121] WiQe RdxdQ, WiKe Rdx, WiVe Rdxdv- where / may refer to the number of heads, dy, dx and dQmay refer to the dimensions of the value, key and query.

[0122] The result of the two or more heads may be concatenated according to the following equation: where j -^hdvxd and h may refer to the number of heads.

[0123] The embedded input data may be transformed via the multi-head self-attention 206 into a context tensor. The context tensor may represent the sequence of elements and the relation between two or more elements of the input data. The context tensor may be a second rank tensor and / or may comprise one or more first rank tensor(s). After the multi-head self-attention 206, layer normalization 208 may be applied based on the context tensor and / or the embedded input data from the residual connection. Applying layer normalization 208 may refer to normalizing the context tensor. Normalizing the context tensor may lower the values of the entries of the context tensor. This reduces the computational cost associated with processing the context tensor. Further, it improves the training by contributing the loss to converge and preventing instabilities.

[0124] Layer normalization 208 may be followed by passing the context tensor to a feedforward layer 210, again followed by layer normalization 212 based on the residual connection to the context tensor and / or the output of the feed-forward layer 210. The feed-forward layer 210 may be a feed-forward neural network. The feed-forward neural network may comprise of a plurality of fully connected neurons. Passing the context tensor through the feed-forward neural network may result in transforming the context tensor linearly. Additionally or alternatively, the neural network may comprise one or more activation functions such as a rectified linear unit (ReLU). Hence, the neural network may be configured for performing one or more non-linear operations to the context tensor and / or transforming the context tensor non-linearly. After the context tensor has been transformed and / or normalized by the feed-forward layer 210 and the layer normalization 212, the context tensor may be provided to one or more further encoder blocks 214. Having passed the context tensor through the feed-forward layer 210 may adapt the context tensor for the processing by a further attention layer of the one or more further encoder blocks 214 for applying a self-attention filter, preferably multi-head self-attention 206. The context vector after being transformed by the layer normalization 212 and the feed-forward layer 210 may be referred to as hidden state.

[0125] The encoder output 276 comprises of a linear layer 216 and a softmax layer 218. The linear layer 216 may transform the context vector into a logits vector. The linear layer may be fully-connected. The logits vector obtained by passing the context tensor through the linear layer 216 may be passed through the softmax layer 218. Passing the logits vector through the softmax layer 218 may refer to applying the softmax function to the logits vector. Applying the softmax function to the logits vector may result in a probability distribution of one or more elements corresponding to the sequence of elements in the input data. From the probability distribution based on predefined selection criteria, one or more elements may be chosen. The one or more chosen elements may be referred to as the one or more elements generated by the transformer encoder. The one or more generated elements may be provided to the encoder input for generating further one or more elements corresponding to the sequence of the input data and the one or more elements generated by the transformer encoder as described within the context of FIG. 10.

[0126] FIG. 8 illustrates an embodiment of a transformer decoder architecture.

[0127] The transformer decoder comprises a decoder input 284, one or more decoder blocks 280, 232 and a decoder output 292. The transformer decoder architecture may be derived from the transformer encoder-decoder architecture as known in the art and shown in FIG. 9. The transformer decoder may be referred to as X-former. The transformer decoder architecture may correspond to the decoder architecture associated with the transformer encoder-decoder architecture independent of receiving one or more hidden states from the encoder of the transformer encoder-decoder. A plurality of transformer decoder architectures are available in the art, such as the generative pretrained transformers (GPT).

[0128] The decoder input 284 may apply input embedding 220 and positional encoding 222 analogous to the input embedding 202 and the positional encoding 203 as described within the context of FIG. 7.

[0129] The decoder block 280 may comprise the layer normalizations 226, the masked multihead self-attention 224, the feed-forward layers 228 and / or the layer normalization 230. The embedded input data resulting from passing the input data through the decoder input 284 may be provided to the layer normalization 226 via a residual connection. Further, masked multi-head self-attention 224 may be applied to the embedded input data. Masked multi-head self-attention 224 corresponds to the multi-head selfattention 206 as described within the context of FIG. 7 with additionally masking a part of the embedded input data associated with elements later in the sequence than the element to be generated. Additionally or alternatively, the part of the input data associated with elements later in the sequence than the element to be generated may not be received and / or transformed into the embedded input data. Thus, the transformer decoder may be suitable for generating a subsequent element to a sequence, whereas the transformer encoder may be suitable for generating a missing element in within one sequence and / or between two or more sequences. Therefore, the transformer encoder may be configured for classification tasks. The transformer decoder may be configured for text generation.

[0130] Similar to the transformer encoder as described within the context of FIG. 7, a context tensor may be generated by applying the masked multi-head self-attention 224 and the layer normalization 226. The context tensor may be provided to the layer normalization 230 via a residual connection. Further, the feed-forward layer 228 and the layer normalization 230 may be analogous to the feed-forward layer 210 and the layer normalization 212 as described within the context of FIG. 7. The context tensor may be provided to one or more further decoder blocks 232.

[0131] The decoder output 292 may comprise a linear layer 234 and a softmax layer 236. The linear layer 234 and the softmax layer 236 may be analogous to the linear layer 216 and the softmax layer 218 as described within the context of FIG. 7.

[0132] FIG. 9 illustrates an embodiment of a transformer encoder-decoder architecture. The transformer encoder-decoder may comprise the encoder input 288, the one or more encoder blocks 286, 264, the decoder input 294, the decoder block 290 and the decoder output 292. The encoder input 288 may correspond to the encoder input 278 of FIG. 7. The one or more encoder block(s) 286, 264 may correspond to the one or more encoder blocks 274, 214 of FIG. 7. The decoder input 294 may correspond to the decoder input 284 of FIG. 8.

[0133] The decoder block 290 may comprise a masked multi-head self-attention 270, a layer normalization 272, a feed-forward layer 238 and a layer normalization 240 analogous to the masked multi-head self-attention 224, the layer normalization 226, the feedforward layer 228 and the layer normalization 230 as described within the context of FIG. 8. The decoder block 290 may further comprise a multi-head self-attention 250 and a layer normalization 248. Analogous to the description of FIG. 8, the context tensor may be obtained from the masked multi-head self-attention 270 and the layer normalization 272. Multi-head self-attention 250 analogous to the multi-head self-attention 206 of FIG. 7 may be applied to the context vector obtained from the layer normalization 272 and the hidden states of the one or more encoder blocks 286, 264. Layer normalization 248 may be applied to the context vector obtained from the multi- head self-attention 250 and the context vector obtained from the layer normalization 272 provided via a residual connection. The context vector resulting from the layer normalization 248 may be processed via the feed-forward layer 238 and the layer normalization 240 analogous to the description of FIG. 8. The context vector resulting from the layer normalization 240 may be provided to further decoder blocks 242 analogous to the decoder block 290. The context vector obtained from the one or more decoder blocks 290, 242 may be provided to the decoder output 292. The decoder output 292 may correspond to the decoder output 282 of FIG. 8.

[0134] With the above-described architecture, the transformer encoder-decoder may receive and process input data at the encoder input 288 and the one or more encoder blocks 286, 264 and the decoder block 290 and the decoder output 292. Based on the input data, the transformer encoder-decoder may generate output data part by part or sequentially. The sequentially generated output data may be provided to and / or may be processed by the decoder input 294, the one or more decoder blocks 290, 242 and the decoder output 292. Preferably, a sequence may be provided to the encoder input 288 and after having generated at least a part of the output data, the decoder input 294 may be provided with at least the part of the elements of the output data already generated. By doing so, the next elements of the output data may be generated with a higher accuracy by taking the input data and the generated output data into account since more data may be received by the transformer encoder-decoder over time.

[0135] Because of the transformer encoder-decoder architecture, the transformer encoderdecoder may be configured for transforming a sequence into another representation of the sequence. An example for transforming one sequence into another representation may be translation of one sentence into another language. A plurality of transformer encoder-decoders are available in the art, such as BART, T5 or the like.

[0136] In an embodiment, the layer normalization 208, 212 may be applied prior to the masked multi-head self-attention 224, multi-head self-attention 206 and / or the feedforward layer 210 in the transformer decoder, the transformer encoder and / or the transformer encoder-decoder. By doing so, the computational resources for applying the multi-head self-attention 206 and / or the feed-forward layer 210 to the embedded input data and / or the context tensor may be decreased as the entries of the respective tensors may be lower after normalization.

[0137] In an embodiment, the decoder output 292 may comprise a classification neural network, further feedforward layers, convolutional layers, fully connected layers or the like. For example, the transformer encoder-decoder may be configured for choosing between a plurality of options. For this purpose, the transformer encoder-decoder may be provided with three different input data sets and may classify the context vectors obtained from the one or more decoder blocks 290 via one or more linear layers. Fol- lowingly, the architecture may be extended depending on the use case to be solved.

[0138] FIG. 10 illustrates an embodiment of training and / or deploying the transformer encoder, the transformer decoder and / or the transformer encoder-decoder.

[0139] The encoder / decoder / encoder-decoder architecture 302 may correspond to the transformer decoder, the transformer encoder and / or the transformer encoder-decoder as described within the context of FIG. 7- FIG. 9.

[0140] The output data generated by the encoder / decoder / encoder-decoder architecture 302 may comprise one or more elements, in particular a sequence of elements. The previously generated elements of the output data may be provided as input for generating the next element in the sequence of the output data.

[0141] If, for instance, the input data did correspond to a request for generating chemical product data and / or associated model instructions, then the output data may correspond to chemical product data. If, as in the reverse case, the input data did correspond to chemical product data, then output data may correspond to a request for generating chemical product data and / or model instructions.

[0142] In the example of FIG. 10, the input data may comprise N elements, in particular input tokens. For instance, any request for generating chemical product data and / or associated model instructions, which may initially be input in a format comprising text and / or numerical data, may be tokenized, thereby converting it into a sequence of tokens. An input token may be a token dedicated to be inputted into a data-driven model such as the transformer decoder, the transformer encoder or the transformer encoder-decoder. The output data to be generated may comprise M elements. The encoder / decoder / encoder-decoder architecture 302 may generate one element of the output data based on receiving the input data and optionally previously generated elements of the output data at a timestep. Hence, for generating M elements M time steps are required. A time step comprises providing input 310, 312, 314 to the en- coder / decoder / encoder-decoder architecture 302 and receiving output data 304, 308, 306 from the encoder / decoder / encoder-decoder architecture 302. In a first timestep, the input 310 may comprise of N input tokens. The N input tokens may be associated e.g. with N words, stems or endings. Preferably, the N input tokens may specify a question or other form of request, and / or associated model instructions. One or more input tokens may specify the beginning of the sequence of tokens and / or the end of the sequence of tokens. The input 310 may be processed by the encoder / decoder / en- coder-decoder architecture 302. Based on the input 310 at least a part of the output data 304 may be generated. The at least a part of the output data may comprise a first output token. In the next timestep, the generated first output token may be provided together with the input 312. Specifically, where the input 312 may be received by a transformer encoder-decoder the input tokens may be received at the encoder input 288 and the first output token may be received at the decoder input 294. Where the input 312 may be received by the transformer encoder, the input 312 may be received by the encoder input 278 and analogously regarding the transformer decoder and the decoder input 284. Based on the input 312, the output data 308 comprising the first output token and a second output token may be generated. Generating the output data 308 based on the input 312 may refer to generating the second token based on the first token and the N input tokens, wherein the first token may have been generated based on the N input tokens. This process may be repeated until the last token in the sequence of the output data 306 may be generated. Preferably, the last token may be an end token. The end token may terminate the generation of a further output token.

[0143] Similarly, to the data processing during deployment of the encoder / decoder / encoder- decoder architecture 302, the encoder / decoder / encoder-decoder architecture 302 may be trained. The training data set may comprise a plurality of sequences comprising a plurality of elements. The sequences may be associated with the input data and / or the output data. Additionally or alternatively, the sequences may be independent of the input data and / or the output data. For example, where the input data and the output data may refer to chemical compositions represented via text, the training data set may comprise sequential text data independent of chemical compositions. In this example, the training data set may comprise sequences of words originating from a conversation. In an embodiment, the training data set may comprise at least partially input data sets and / or output data sets.

[0144] The training may be initialized by initializing the encoder / decoder / encoder-decoder architecture 302. In an embodiment, the parameters associated with the encoder / de- coder / encoder-decoder architecture 302 may be initialized randomly. Additionally or alternatively, the input embedding of the encoder / decoder / encoder-decoder architecture 302 may be obtained by training a CBOW model or a skip gram model as described within the context of FIG. 6. The trained embedding layer may be used during training. The parameters associated with the embedding layer may be kept constant and / or may be updated after a predefined number of training epochs. By doing so, the number of parameters to be updated is lower, enabling a faster and less computational resources-consuming training. Further, the accuracy associated with the embedding layer may be constant and / or may be increased by avoiding error compensation in relation to the just initialized encoder / decoder / encoder-decoder architecture 302.

[0145] During the training of the encoder / decoder / encoder-decoder architecture 302, at least a part of the sequences of the training data set may be provided to the encoder / de- coder / encoder-decoder architecture 302 one by another and one or more elements may be generated based on the sequences of the training data set one by another. The elements generated based on the sequences may follow the elements of the parts of sequences the encoder / decoder / encoder-decoder architecture 302 may have been provided with. The generated one or more elements may be compared to the one or more elements following the at least a part of the sequences provided to the en- coder / decoder / encoder-decoder architecture 302 as specified by the training data set. Hence, during the training the encoder / decoder / encoder-decoder architecture 302 may generate a guess on the next element and the guess on the next element in a sequence may be compared to the ground truth specifying the actual next element according to the training data set. Based on the guess on the next element and the ground truth a loss may be determined. The loss may define the similarity between the guess on the next element and the ground truth. The loss may be determined by forming a vector dot product between the token associated with the one or more elements and the token associated with the ground truth. A loss unequal to zero may result in updating the parameters associated with encoder / decoder / encoder-decoder architecture 302. Preferably the parameters associated with the encoder / decoder / en- coder-decoder architecture 302 may be independent of the embedding layer. For example, the parameters associated with the encoder / decoder / encoder-decoder architecture 302 may be weights of the neurons of the encoder / decoder / encoder-decoder architecture 302.

[0146] Based on the determined loss, backpropagation may be applied to determine the gradients associated with the parameters of the parameters associated with encoder / de- coder / encoder-decoder architecture 302 to lower the loss. According to the deter- mined gradients, the parameters associated with the encoder / decoder / encoder-de- coder architecture 302, preferably the weights of the neurons associated with the en- coder / decoder / encoder-decoder architecture 302, may be updated by using a gradient descent algorithm.

[0147] The training data set may be unlabeled. The sequences of elements within the training data set may inherently comprise the ground truth for determining the loss with respect to the one or more elements generated during the training of the encoder / decoder / en- coder-decoder architecture 302. Hence, the encoder / decoder / encoder-decoder architecture 302 may be trained self-supervised. This is advantageous since time and resources for creating a labeled training data set may be saved. Furthermore, this enables the usage of large training data sets associated with a size of several terabytes. Consequently, the data-driven model may be accurate in generating elements of a sequence. In addition, the large training data set enables few shot predictions or even zero shot predictions. Hence, the data-driven models trained as described above are versatile contributing to saving resources needed for training and / or hosting a plurality of purpose-driven models such as convolutional neural networks. The training described above may be referred to as pretraining. The data-driven model may be configured for performing few shot or even zero shot predictions with respect to a plurality of use cases after pretraining. The performance of the data-driven model may be increased further by additional training referred to as fine-tuning. The training data used for fine-tuning may comprise pairs of training input data and training output data. For fine-tuning the model to generate chemical product data as output upon being provided with a request for doing so and / or associated model instructions as input, the training input data may comprise historical requests and / or associated model instructions, and the training output data may comprise verified chemical product data, i.e., product data that have been considered to be useful, suitable and / or desirable outputs for the respective historical requests and / or associated model instructions. For fine- tuning the model to generate, upon being provided with chemical product data, model instructions for generating the chemical product data as output, the training input data may comprise verified chemical product data and optionally also historical requests associated with the verified chemical products, and the training output data may comprise model instructions that were used for instructing a model and thereby led to the verified chemical product data being generated by the model.

[0148] FIG. 11 illustrates an embodiment of input embedding. Where the sequence of elements associated with the input data, preferably comprised in the input data, may be of one type, the input embedding 202, 220, 252, 266 as described within the context of FIG. 7 - 9 may be used. For example, a type of input data may be text where the elements may be associated with at least a part of a word, a punctuation character, a start token specifying the beginning of one or more sequences associated with the input data and / or the end token. In another example, the input data may be at least partially numerical. Hence, the input data may comprise a plurality of numbers. A request for generating chemical product data, for instance, may comprise both text and numerical data. Numerical input data may be for example tabular data. Tabular data may specify one or more rows and / or one or more columns. Hence, the tabular data may comprise one or more cells, wherein the cells may be associated with one or more numerical values.

[0149] Numerical input data may require a different embedding than text input data. Input embeddings for numerical input data may comprise a token embedding, a positional embedding, a column embedding, a row embedding or a combination thereof.

[0150] Applying a token embedding to one or more elements, in particular tokens associated with the input data may result in a machine-processable representation associated with the one or more elements, in particular tokens. Applying the token embedding to one or more elements may refer to passing the one or more elements through the embedding layer, e.g. as described within the context of FIG. 6. Hence, token embeddings may specify the one or more elements, in particular tokens in a machine-processable representation. For example, the token embedding may transform a numerical value into a vector. This is advantageous since this representation can be enriched by further information such as the position of the token within the sequence and / or within a table associated with the sequence of tokens. The positional embedding may be analogous to the positional embedding as described within the context of FIG. 6, FIG. 7-9. Where the input data may be tabular data, column embedding may be applied. Applying a column embedding to one or more elements, in particular tokens associated with the input data may result in a machine-processable representation specifying the location of the one or more elements within a table 402, preferably within the columns of the table 402. Applying the column embedding may refer to adding a column factor to the input data embedded via token embeddings, in particular the embedded input data. The column factor may be the same for elements associated with the same column and / or may differ between two or more elements associated with different columns. Analogous, row embeddings may be applied where the input data may be tabular data. Applying a row embedding to one or more elements, in particular tokens associated with the input data may result in a machine-processable representation specifying the location of the one or more elements within a table 402, preferably within the rows of the table 402. Applying the row embedding may refer to adding a column factor to the input data embedded via token embeddings, in particular the embedded input data. The row factor may be the same for elements associated with the same row and / or may differ between two or more elements associated with different rows.

[0151] In an embodiment, input data may be at least partially numerical and at least partially text. As indicated above, this may be the case for a request for generating chemical product data. Hence, the input data may comprise two or more types of data. A type of data may refer to a modality. Followingly, different embeddings may be applied to the input data. To parts of the input data comprising text the input embedding referred to in FIG. 6, FIG. 7-9 may be applied. To parts of the input data being numerical token embeddings, positional embeddings, column embeddings and row embeddings may be applied. Further, segment embeddings may be applied to the input data independent of the type of input data. The segment embedding may specify the type of input data one or more elements may be associated to. For example, if the input data comprises of text and numbers, the input data may comprise of two types of input data. Applying the segment embedding to the input data may refer to adding a segment factor to the input data, preferably the embedded input data and / or the input data after having applied the token embedding. The segment factor may specify the type of data associated with the one or more elements. The segment factor may be the same for one or more elements associated with the same type of input data and / or may differ between two or more elements associated with different types of input data.

[0152] Applying the token embedding, the positional embedding, the segment embedding, the column embedding, the row embedding or a combination thereof may result in embedded input data and / or may be the output of any one of the encoder input 278, 284, 288 or decoder input 284, 294. The data obtained by applying the token embedding, the positional embedding, the segment embedding, the column embedding, the row embedding or a combination thereof may be processed by the encoder block 274, 286, decoder block 280, 290, encoder output 276, decoder output 292, 282.

[0153] FIG. 12 illustrates a further embodiment of input embedding.

[0154] Input data to the data-driven model, in particular to the encoder input and / or the decoder input as described in the context of FIG. 7-9, may comprise image data. Also a request for generating chemical product data may comprise image data, such as in the form of an image associated with the chemical product to which the request relates. The data-driven model may be parametrized to receive image data. For processing image data as input data, the data-driven model may comprise one or more encoder blocks and / or one or more decoder blocks and / or one or more encoder outputs and / or one or more decoder outputs as described within the context of FIG. 7-9. FIG. 12 may show an embodiment of an encoder input and / or a decoder input. When processing image data, the encoder input and / or the decoder input of the data-driven model may be as described within the context of FIG. 12. The encoder input and / or decoder input may comprise one or more linear projection layers 514 for a linear projection of one or more images, preferably one or more partial images, more preferably a sequence of two or more partial images. The one or more linear projection layers 514 may be suitable for changing the dimension of the one or more received images, preferably one or more partial images, preferably passing the one or more images, preferably partial images, through the one or more linear projection layers 514 may result in applying image embedding, preferably partial image embedding to the one or more images and / or partial images.

[0155] Furthermore, when a sequence of two or more images and / or partial images may be received, positional embedding may be applied to the sequence, preferably by passing the sequence of one or more images and / or partial images through the one or more linear projection layers 514. Applying positional embedding may refer to adding a positional factor. The positional factor may be different depending on the position of the image and / or the partial image within the sequence. In particular, the positional factor added to a first element of the sequence may be different to the positional factor added to a second element of the sequence. The first element of the sequence may be a first image and / or first partial image. The second element of the sequence may be a second image and / or a second partial image.

[0156] The representation of the one or more images, preferably one or more partial images, may be obtained based on the following equation: where x class is the image class embedding 528 ,XNis the n-th image, in particular p partial image in the sequence,zo is the representation of the one or more images, preferably one or more partial images, (H,W) are the resolution of the image, in particular the image the partial images are generated on, C is the number of channels associated with the one or more image, in particular the one or more partial images and D is the dimension of the representation of the one or more images, preferably one or more partial images. Applying the partial image embedding may refer to forming the product ofx'iwith E above-described equation. Applying the positional embed- ding may refer to adding the factor pos according to the above-described equation.

[0157] By doing so, text-based data, numerical data, tabular data, image data or the like may be processed by one data-driven model.

[0158] The present disclosure has been described in conjunction with preferred embodiments and examples as well. However, other variations can be understood and effected by those persons skilled in the art and practicing the claimed invention, from the studies of the drawings, this disclosure and the claims. Notably, in particular, any steps presented can be performed in any order, i.e. the present invention is not limited to a specific order of these steps. Moreover, it is also not required that the different steps are performed at a certain place or at one node of a distributed system, i.e. each of the steps may be performed at different nodes using different equipment / data processing.

[0159] As used herein ..determining" also includes ..initiating or causing to determine", “generating" also includes ..initiating and / or causing to generate" and “providing” also includes “initiating or causing to determine, generate, select, send and / or receive”. “Initiating or causing to perform an action” includes any processing signal that triggers a computing node or device to perform the respective action.

[0160] In the claims as well as in the description the word “comprising” does not exclude other elements or steps. The indefinite article “a” or “an” and the definite article “the” does not exclude a plurality. In particular, indefinite article “a” or “an” may be replaced with one or more and the definite article “the” may be replaced with the one or more. A single element or other unit may fulfill the functions of several entities or items recited in the claims. The mere fact that certain measures are recited in the mutual different dependent claims does not indicate that a combination of these measures cannot be used in an advantageous implementation. Procedures like the receiving of a request for generating chemical product data, the providing of model instructions and / or a request for generating chemical product data as input to a data-driven model, the providing of chemical product data as input to a data-driven model, any assessment of outputs generated by the data-driven model, etc., performed by one or several units or devices, can be performed by any other number of units or devices. These procedures can be implemented as program code means of a computer program and / or as dedicated hardware. A computer program product may be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium, supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems.

[0161] Any disclosure and embodiments described herein relate to the methods, the systems, devices, any computer program element lined out above and vice versa. Advantageously, the benefits provided by any of the embodiments and examples equally apply to all other embodiments and examples and vice versa.

[0162] Any reference signs in the claims should not be construed as limiting the scope.

[0163] A method for generating chemical product data associated with a chemical product is presented, the method including a) receiving a request for generating chemical product data associated with the chemical product, the request including an indication of the chemical product, b) generating the requested chemical product data by providing model instructions based on the received request as input to a trained data-driven model, and c) providing the generated chemical product data. The data-driven model and / or at least part of the model instructions used for generating the requested chemical product data are determined by a respective training and / or selection process. In this way, chemical product data can be generated that allow for an improved production and / or processing of chemical products.

Claims

Claims:1 . A method (100) for generating chemical product data associated with a chemical product, the method including: receiving (101) a request for generating chemical product data associated with the chemical product, the request including an indication of the chemical product, generating (102) the requested chemical product data by providing model instructions based on the received request as input to a data-driven model trained and / or parametrized to generate chemical product data as output in response to receiving model instructions for doing so as input, and providing (103) the generated chemical product data, in particular for monitoring and / or controlling production and / or processing of the chemical product, wherein: a) the data-driven model used for generating the requested chemical product data is determined (104a; 104c) from a plurality of candidate data-driven models, wherein the plurality of candidate data-driven models are trained and / or parametrized to generate chemical product data as output in response to receiving model instructions for doing so as input, wherein the data-driven model used for generating the requested chemical product data is determined based on an assessment of outputs generated by the plurality of candidate data-driven models in response to receiving, as input, model instructions based on one or more historical requests for generating chemical product data, and / or b) at least part of the model instructions used for generating the requested chemical product data are determined (104b; 104c) based on an assessment of outputs generated by a data-driven model, which is trained and / or parametrized to generate chemical product data as output in response to receiving model instructions for doing so as input, in response to receiving respective candidate model instructions as input.

2. The method (100) as defined in claim 1 , wherein the assessment of the candidate data-driven model outputs is carried out using a data-driven model trained and / or parametrized to generate, as output, a ranking of chemical product data upon receiving, as input, model instructions for doing so as input, wherein the model instructions are basedon the chemical product data to be ranked and a request for generating the chemical product data, and wherein the ranking is indicative of a correspondence between the request and the respective chemical product data.

3. The method (100) as defined in claim 1 or 2, wherein the plurality of candidate data-driven models include data-driven models of different architectures and / or data-driven models of the same architecture with different parameters.

4. The method (100) as defined in any of the preceding claims, wherein the plurality of candidate data-driven models include data-driven models of different architectures and, per architecture, data-driven models with different parameters, wherein the candidate data-driven models are also trained and / or parametrized to generate, as output, a ranking of chemical product data upon receiving, as input, model instructions for doing so as input, wherein the model instructions are based on the chemical product data to be ranked and a request for generating the chemical product data, wherein the ranking is indicative of a correspondence between the request and the respective chemical product data, and wherein the assessment of the outputs generated by data-driven models of the same architecture is carried out using one of these models for generating a ranking of the respective outputs.

5. The method (100) as defined in any preceding claims, wherein the candidate model instructions are determined based on historical requests and / or verified chemical product data, and wherein the assessment of the outputs includes a comparison (104b-5) of the outputs to verified chemical product data.

6. The method (100) as defined in claim 5, wherein the candidate model instructions are determined by using a data-driven model trained and / or parametrized to generate, in response to receiving chemical product data as input, model instructions for instructing a data-driven model to generate chemical product data as output, wherein the model is provided (104b-1) with verified chemical product data as input to generate the candidate model instructions.

7. The method (100) as defined in any of the preceding claims, wherein the model instructions further specify a respective request for generating chemical product data for the respective data-driven model.

8. The method (100) as defined in any of the preceding claims, wherein the received request comprises a sequence of a plurality of elements, and the data-driven model used for generating the chemical product data is trained and / or parametrized to receive a sequence of a plurality of elements and generate a sequence of a plurality of further elements according to the received sequence of elements.

9. The method (100) as defined in any of the preceding claims, wherein the data- driven model used for generating the requested chemical product data is a pretrained data- driven model.

10. The method (100) as defined in any of the preceding claims, wherein the data- driven model used for generating the requested chemical product data is a fine-tuned data- driven model, wherein the fine-tuned data-driven model is trained based on historical model instructions and corresponding historical chemical product data.11 . The method (100) as defined in any of the preceding claims, further including a step of retrieving an indication of a chemical structure and / or a property of the chemical product based on the indication of the chemical product included in the received request by providing the indication of the chemical product to a data source.

12. The method (100) as defined in claim 11 , wherein the providing of the indication of the chemical product to the data source comprises providing a query obtained from the indication of the chemical product included in the received request, wherein the chemical structure and / or the property of the chemical product is retrieved according to the query.

13. The method (100) as defined in claim 11 or 12, wherein the providing of the indication of the chemical product to the data source comprises providing a numerical representation of the indication of the chemical product, wherein the chemical structure and / or the property of the chemical product is retrieved according to a similarity score associated with a) the numerical representation of the indication of the chemical product and b) a numerical representation of indications of chemical structures and / or properties of chemical products retrievable from the data source.

14. A data processing system comprising a processor configured to carry out the steps of any of the methods (100) as defined in claims 1 to 13.- M -15. A use of the chemical product data generated according to any of the methods (100) as defined in the claims 1 to 13, particularly for monitoring and / or controlling production and / or processing of the chemical product, and / or of the data processing system as defined in claim 14 for generating chemical product data.

Citation Information

Patent Citations

  • Formulation graph for machine learning of chemical products

    US11862300B1

  • Augmentation of multimodal time series data for training machine-learning models

    US20230045548A1

  • Methods and apparatuses for characterizing chemical substances, measuring physicochemical properties and generating control data for synthesizing chemical substances

    WO2023198927A1