Determining a property of a chemical product

By employing a data-driven model to convert product requests into retrieval requests for measurement data, the method addresses the underutilization of measurement data in the chemical industry, enhancing data reuse and resource efficiency.

WO2025133141A1PCT designated stage expired Publication Date: 2025-06-26BASF SE
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/087937
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-12-20
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

In the chemical industry, measurement data acquired during research and development are often underutilized as they are typically analyzed only a few times for selected properties shortly after acquisition, leading to a waste of valuable resource information.

Method used

A method is developed that uses a data-driven model to efficiently retrieve measurement data relevant to determining specific properties of chemical products by converting product requests into retrieval requests, allowing for the reuse of accumulated measurement data.

Benefits of technology

This approach enables the efficient retrieval and reuse of measurement data, saving laboratory resources by avoiding repeated data acquisition and providing more insights into chemical product properties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024087937_26062025_PF_FP_ABST
    Figure EP2024087937_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method (100) for determining a property of a chemical product, including a) receiving (101) a product request for determining a target property value associated with a target property of a target product, and b) acquiring (102) a retrieval request for retrieving, from a data management system (20), measurement data associated with the target product and suitable for determining based thereon the target property value. The retrieval request is acquired by providing the product request to a data-driven model (10) trained to relate product requests to retrieval requests. Furthermore, the method includes c) retrieving (103) measurement data associated with the target product from the data management system using the acquired retrieval request, and d) determining (104a, 104b) the target property value based on the retrieved measurement data. The invention allows to make more efficient use of measurement data associated with chemical products.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Determining a property of a chemical product

[0002] FIELD OF THE INVENTION

[0003] The invention relates to a method and a data processing system for determining one or more properties of one or more chemical products, and to a use of the method and / or the data processing system.

[0004] BACKGROUND OF THE INVENTION

[0005] Measurement data acquired during research and development in the chemical industry are often stored for a long time, but are nevertheless, if not just once, only analyzed a few times with respect to selected properties of a respective chemical product, typically relatively shortly after their acquisition. Hence, large amounts of measurement data can accumulate over time whose potential for gaining further insights into chemical product properties therefrom may not yet have been fully exploited. Part of the complexity in making more efficient use of this valuable resource derives from the fact that the measurement data are typically acquired with various different modalities and are being stored in various different formats.

[0006] SUMMARY OF THE INVENTION

[0007] It is regarded as an object of the present invention to make more efficient use of measurement data associated with chemical products.

[0008] In a first aspect, the invention relates to a method for determining one or more properties of one or more chemical products, the method including a) receiving a product request for determining one or more target property values associated with a target property of a target product, b) acquiring a retrieval request for retrieving, from a data management system, measurement data associated with the target product and suitable for determining based thereon the one or more target property values, the retrieval request being acquired by providing the received product request to a data-driven model trained to relate product requests to retrieval requests, c) retrieving measurement data associated with the target product from the data management system using the acquired retrieval request, and d) determining the one or more target property values based on the retrieved measurement data.

[0009] Hence, i) a product request for determining one or more target property values associated with a target property of a target product and ii) a trained data-driven model are used to acquire a retrieval request for retrieving, from a data management system, measurement data associated with the target product and suitable for determining based thereon the one or more target property values. This has been found to allow for an efficient retrieval of measurement data relevant to the product request. In particular, the use of the trained data-driven allows to deal with product requests associated with target products being indicated in multimodal form. This is particularly relevant for chemical products, which are usually associated with a plurality of different names and / or representations.

[0010] Since, moreover, measurement data associated with the target product are retrieved from the data management system using the acquired retrieval request, and since the one or more target property values are determined based on the retrieved measurement data, measurement data, whose potential for gaining further insights therefrom otherwise may have been lost, are put to a further use. Thus, more efficient use can be made of measurement data associated with chemical products. In other words, a sustainable re-use of measurement data can be facilitated, which allows to save laboratory resources, since a repeated acquisition of measurement data can be avoided.

[0011] The method is preferably a computer-implemented method. In particular, the method may be implemented in terms of instructions of a computer program executable on a general- purpose computer and / or on one or more nodes of a computing network. The computer or the one or more nodes carrying out the method may communicate with another computer and / or other network nodes for carrying out the method. For instance, the step of acquiring the retrieval request can include sending a request, from a local computer or computing network, to a remote server hosting the data-driven model to access the model, using the model upon being granted access to the model to generate the retrieval request, and receiving the generated retrieval request.

[0012] The product request may be received via a user interface. Hence, the product request may be initiated by a user providing an input to a user interface. The product request may then also be referred to as a user request. However, the product request can in principle also be initiated otherwise, such as by a different part of a computing environment in which the method is running. The product request may then be generated by, and received from, that different part of the computing environment.

[0013] The product request may be understood as being associated with the target product, which is preferably a chemical product, and one or more target properties of the target product. The product request can be a request to determine one or more target property values associated with one or more target properties of the target product. The target product is preferably a chemical target product. The word “product”, as used herein, can refer to any substance or material. The modifier “target” in the expressions “target product” and “target property” / ”tar- get properties” is preferentially to be understood as indicating that the respective product and / or property is of interest to the issuer of the request. The association between the one or more target property values and the one or more target properties may be such that the one or more target property values shall be values of a respective target property.

[0014] Target properties of a target product are understood broadly herein so as to possibly include any property associated with the target product. A property associated with the target product may also be a property associated with a chemical reaction for which the target product is used or from which it arises. In particular, a property associated with the target product may also be a property associated with a production (e.g., a synthesis) of the target product. One possible target property would be, for instance, a conversion ratio associated with the target product, such as a conversion ratio achievable in producing the target product under various production conditions.

[0015] The receiving of the product request may follow a step of providing the product request, wherein such a providing of the product request before its receipt may correspond to an issuing or a transmitting of the product request, such as by a user. The receiving of the product request may also precede a providing of the product request, wherein such a providing of the product request after its receipt may correspond to a providing of the product request for further processing.

[0016] The product request preferably includes an indication of the target product and the target property. Hence, for instance, the product request may indicate a chemical product of interest and additionally a property of interest. The request may then be understood as a request to determine, for the indicated product, a value of the indicated property. In other words, the request may be a request to determine a value for the indicated property of the indicated product. Hence, the indication of the target property is not to be understood so as to necessarily refer already to a value of the target property. Instead, the indication of the target prop- erty preferably identifies which property, among a plurality of possible properties, is of interest and hence the ’’target” property. Like the indication of the target product, the indication of the target property may therefore refer to a name or other identifier. Thus, for instance, the product request may include an identifier of the target property and an identifier of the target product. The target product may also be indicated in terms of a digital representation thereof.

[0017] While the indication of the target property by the product request does not necessarily indicate a value of the target property, the method presented herein is preferably suitable for determining respective values of one or more properties of one or more chemical products. Hence, the method presented herein may also be considered a method for evaluating one or more properties of a chemical product instead of a method for determining one or more properties of a chemical product.

[0018] While the product request may be a request to explicitly determine the one or more target property values in order to, for instance, provide the determined one or more target property values as output, this is not necessarily so. In fact, the request may be formulated not explicitly as a request, but rather as a question, from which a request may be derived. For instance, the product request may be a closed-ended question, which may be processed as a request to determine an answer to the closed-ended question. A closed-ended question regarding a target property of a target product may be provided, for instance, in terms of a range of values for the target property, which may also be understood as an assumption about possible values which the target property may take for the target product. The product request may be processed in order to determine whether this assumption is true or not, wherein this result may be provided as output of the method.

[0019] In a particular embodiment, the product request may indicate a class of chemical products for which a value of a target property shall be determined. The class of chemical products may be indicated in terms of one or more properties shared by all products in the class. For instance, a range of values may be indicated for each of one or more properties, wherein for each product in the class the one or properties are required to have a value within the respective range. An indication of a range of values of a property may be made explicitly, such as in terms of numerical values, or implicitly, such as in terms of categories such as “heat resistant”, “nickel-based”, “non-toxic” orthe like. Requests ofthis kind, i.e., requests fortarget property classes, can be advantageous in situations such as the development of a compound product where it is not yet known or fixed which specific chemical product is to be used in the compound product and hence which specific chemical product information is needed. Instead, the compound product may just put some boundary conditions on the possible chemical products to be used. These boundary conditions can then be translated into a product class indication in the product request. It will be understood that if, in the following, reference is made to “a” or even “the” target product, the reference may also be replaced by a reference to a class of target products.

[0020] The retrieval request may preferably be understood as a request to retrieve, from the data management system, measurement data associated with the target product, wherein the measurement data to be retrieved shall be suitable for determining, based thereon, the one or more values of the one or more target properties. Hence, preferentially measurement data are to be retrieved based on which the one or more target property values requested according to the product request can actually be determined. This is achievable by the training of the data-driven model.

[0021] Particularly if the product request is received from a user via an input to a user interface, the retrieval request may be understood as a structured version of the product request, wherein the product request may be considered unstructured. Being structured or unstructured may in this regard refer to whether the respective request can serve as a request to the data management system to retrieve measurement data or not. Additionally or alternatively, the product request being unstructured may refer to the product request being provided in text form, whereas the retrieval request being structured may refer to the retrieval request being provided in terms of a numerical representation.

[0022] The data management system may be associated with a database for storing a plurality of sets measurement data, each set stemming from a respective acquisition, such as a respective experiment, chemical reaction or test, for instance. The data management system may include and / or provide access to the database. In particular, the data management system may retrieve a set of measurement data from the data base upon receipt of a retrieval request. The measurement data may be stored in conjunction with metadata, wherein the retrieval request may indicate the measurement data to be retrieved in terms of the metadata. Metadata associated with a set of measurement data can include, for instance, a data acquisition modality with which the measurement data may have been acquired, a format in which the measurement data are stored and / or chemical and / or physical conditions associated with the measurement. A chemical condition of this kind for a chemical reaction may be, for instance, a set of other substances (i.e., apart from the target product) involved in the reaction, amounts thereof, their roles (e.g., whether they function as educts or products), a temperature at which the reaction took place and / or a pressure under which the reaction took place. In a particular example, the data management system could be SQL-based. In this case, the retrieval requests may be queries.

[0023] The acquiring of the retrieval request may also refer to a determining or a generating of the retrieval request. In the following, the term “acquiring” is used in this regard, but may likewise be replaced by “determining” or “generating”.

[0024] The retrieval request is acquired by providing the received product request to a data-driven model trained to relate product requests to retrieval requests. In particular, the product request may be provided as input to a data-driven model trained to provide, as output, retrieval requests in response to being provided, as input, with product requests, wherein the retrieval request provided as output by the data-driven model after having been provided with the received product request as input is the acquired retrieval request.

[0025] That the data-driven model is trained to relate product request to retrieval requests, such as in an input-to-output manner as exemplarily indicated above, may be understood such that the data-driven model relates, after being trained as indicated above, any product request like the one received to a suitable retrieval request of the kind above. In the previous sentence, “like” and “suitable” are to be understood such that a product request for determining one or more target property values associated with a target product of a target property is, upon being provided to the trained data-driven model, related by the trained data-driven model to a retrieval request for retrieving, from the data management system, measurement data associated with the target product and suitable for determining based thereon the one or more target property values.

[0026] Since the data-driven model may be understood as being constructed from a set of model parameters, the data-driven model being “trained” may also be referred to as the data-driven model being “parameterized”. The model is referred to as “data-driven” since it is trained and / or parametrized using data, i.e., training data. The training data may or may not include product requests and retrieval requests. The data-driven model may be trained and / or parametrized to relate unstructured data to structured data. While, for instance, a product request may correspond to unstructured data, a retrieval request may correspond to structured data. Since the data-driven model is used for acquiring the retrieval request, it may also be referred to as a retrieval request providing model.

[0027] Using the retrieval request, measurement data associated with the target product are retrieved from the data management system. For instance, the retrieval request is provided to the data management system for retrieving the measurement data. The data management system may be configured to retrieve measurement data in response to being provided with a data retrieval request.

[0028] Furthermore, the one or more target property values are determined based on the retrieved measurement data, such as by analyzing the retrieved measurement data with respect to the one or more target properties. Optionally, the determined one or more target property values may furthermore be provided, i.e., for instance, as an output to the user via the user interface.

[0029] Preferentially, the one or more target property values are determined based on the retrieved measurement data by a) acquiring analysis instructions for analyzing the retrieved measurement data with respect to the target property using an analysis engine, the analysis instructions being acquired by providing the retrieved measurement data to a data-driven model trained to relate measurement data to analysis instructions, and b) instructing the analysis engine to analyze the retrieved measurement data with respect to the target property according to the acquired analysis instructions.

[0030] Hence, preferentially both for retrieving the measurement data (i.e., for providing the retrieval request based on which the measurement data can be retrieved) and for determining the one or target property values based on the measurement data (i.e., for providing the analysis instructions according to which the analysis engine is instructed to analyze the measurement data) a trained data-driven model is used. One or more data-driven models may be used. That is to say, a single data-driven model may be trained to both a) relate product requests to retrieval requests and b) measurement data to analysis instructions, wherein this single data-driven model may be used both for retrieving the measurement data and for determining the one or target property values based on the measurement data, or two separate models may be used. In the latter case, while the model used for retrieving the measurement data may, as indicated above, be referred to as retrieval request providing model, the model used for determining the one or target property values based on the measurement data may be referred to as analysis instructions providing model. However, the terms “retrieval request providing model” and “analysis instructions providing model” may also be used to refer to the respective functionalities of a single data-driven model used for both functionalities.

[0031] By using a data-driven model to acquire the analysis instructions for the retrieved measurement data, a reliable analysis of the measurement data can be achieved in an efficient manner irrespective of the format of the measurement data. Even if, as is usually the case, measurement data from various different modalities are stored in various formats, i.e., even in the presence of “multi-modal” stored measurement data, by using an accordingly trained data- driven model it is possible to efficiently find suitable analysis instructions for analyzing retrieved measurement data with respect to a target property in a reliable manner.

[0032] The analysis instructions may preferably be understood as comprising a possible input for an analysis engine that causes the analysis engine to analyze the measurement data with respect to the one or more target properties. The analysis engine may be configured to take, as a further input, the measurement data to be analyzed. Furthermore, the analysis instructions may indicate an analysis engine to be used for analyzing the measurement data.

[0033] The one or more analysis engines can be software tools for analyzing measurement data. Hence, for instance, an analysis engine can be a computer program or a part thereof. Moreover, an analysis engine can have one or a plurality of functionalities, i.e., it can be a singlepurpose analysis engine for carrying out a specific kind of analysis on a specific type of measurement data, or it can be a multi-purpose analysis engine for carrying out a selectable one of a plurality of possible types of analyses on a plurality of types of measurement data. As indicated above, an analysis engine can be configured to receive analysis instructions and optionally measurement data as input. Moreover, an analysis engine can be configured to provide, as an output in response to be provided with an input such as analysis instructions and optionally also measurement data, an analysis result. The analysis results can particularly have the form of one or more target property values, i.e., one or more target property values which the analysis engine was instructed to determine from the measurement data according to the analysis instructions.

[0034] The acquiring of the analysis instructions may also refer to a determining or a generating of the analysis instructions. In the following, the term "acquiring" is used in this regard, but may likewise be replaced by "determining" or "generating".

[0035] The acquisition of the analysis instructions by providing the retrieved measurement data to the data-driven model trained to relate measurement data to analysis instructions may particularly be embodied such that the retrieved measurement data are provided as input to the data-driven model. The data-driven model may be a pretrained model. The pretrained model may be trained on unstructured data such as, for instance, string data. Furthermore, the data-driven model may be trained to provide, as output, analysis instructions in response to being provided, as input, with measurement data, wherein the analysis instructions provided as output by the data-driven model after having been provided with the retrieved measurement data as input are the acquired analysis instructions. The remarks made further above regarding the training of the data-driven model trained to relate product requests to retrieval requests apply analogously to the training of the data- driven model trained to relate measurement data to analysis instructions. If a single data- driven model is used both for acquiring the retrieval requests and for acquiring the analysis instructions, the remarks made further above regarding the data-driven model trained to relate product requests to retrieval requests apply analogously to the training of the data-driven model trained to relate measurement data to analysis instructions.

[0036] Thus, that a data-driven model is trained to relate measurement data to analysis instructions, such as in an input-to-output manner as exemplarily indicated above, may be understood such that the data-driven model relates, after being trained as indicated above, any measurement data like the ones retrieved to suitable analysis instructions of the kind above. In the previous sentence, "like" and "suitable" are to be understood such that measurement data from which it is possible to determine one or more target property values associated with a target property of a target product are, upon being provided to the trained data-driven model, related by the trained data-driven model to analysis instructions based on which an analysis engine can be caused to analyze the measurement data with respect to the target property of the target and thereby to determine the one or more target property values.

[0037] Also insofar as the ability to relate measurement data to analysis instructions is concerned, a data-driven model may be understood as being constructed from a set of model parameters, such that the data-driven model being "trained" may also be referred to as the data- driven model being "parameterized", and the model may be referred to as "data-driven" since it is trained and / or parameterized using data, i.e., training data. The training data may or may not include measurement data and analysis instructions. As indicated above, as far as a data-driven model is used for acquiring analysis instructions, it may also be referred to as analysis instructions providing model.

[0038] It will be understood that if a single data-driven model is used for acquiring the retrieval request and for acquiring the analysis instructions, both the above remarks regarding the training of a data-driven model to provide a retrieval request in response to being provided with a product request and the above remarks regarding the training of a data-driven model to provide analysis instructions in response to being provided with measurement data apply to the training of the single data-driven model, which may also be regarded as a combined data-driven model. This single, combined data-driven model may hence be trained a) to relate product requests to retrieval requests and b) to relate measurement data to analysis instructions. The training of the single, combined model may be carried out using training data which may or may not comprise product requests, retrieval requests, measurement data and / or analysis instructions.

[0039] In an embodiment, any of the one or more data-driven models may comprise one or more machine-learning architectures and model parameters. The one or more machine-learning architectures may be at least one of one or more convolutional layers, one or more fully connected layers, one or more pooling layers, one or more transformer encoders, one or more classification layers, one or more transformer decoders, one or more feed forward layers, one or more linear layers, one or more transformer encoder-decoders or a combination thereof.

[0040] The data-driven model may be trained and / or parametrized based on a training data set. The training data set may comprise historical data related to chemical products and properties thereof. Further, the training data set may comprise at least one data set comprising product requests and corresponding retrieval requests and / or measurement data associated with a chemical product and corresponding analysis instructions. The training data may comprise one or more sequences of elements. Preferably, at least one data set within the training data set may comprise a sequence of elements specifying product requests and corresponding retrieval requests and / or measurement data associated with a chemical product and corresponding analysis instructions. The data-driven model may be trained and / or parametrized based on the training data set to provide the retrieval requests in response to being provided by the product requests and / or to provide the analysis instructions in response to being provided by the measurement data. The data-driven model may be a generative model. The generative model may be trained and / or parametrized based on the training data set to generate retrieval requests and / or analysis instructions, preferably based on being provided by product requests and / or measurement data, respectively. Preferably, the training data set may comprise numerical and / or sentence-based historical data related to chemical products and properties thereof. Numerical historical data may refer to historical data comprising one or more numbers. Sentence-based historical data may refer to historical data comprising at least a part of a sentence. The data-driven model may be trained and / or parametrized to generate the retrieval requests and / or analysis instructions sequentially, in particular based on being provided with product requests and / or measurement data, respectively.

[0041] The data-driven model may be trained and / or parametrized to generate retrieval requests and / or analysis instructions sequentially, in particular based on being provided with product requests and / or measurement data may referto the data-driven model may be trained and / or parametrized to generate retrieval requests and / or analysis instructions comprising a sequence of elements by generating an element of the sequence of elements, preferably the first element of the sequence of elements, based on product requests and / or measurement data, respectively, and generating the further elements, preferably the elements following the first element in the sequence, based on product requests and / or measurement data, respectively, and the previously generated elements in the sequence. By taking previously generated elements into account, the data-driven model may generate the retrieval requests and / or analysis instructions more accurately and better linked to the other elements within the sequence. Ultimately, this enables more reliable analysis results, i.e., more accurate target property values being determined.

[0042] In an embodiment, the data-driven model (i.e., any of the one or more data-driven models) may be a pre-trained data-driven model. The pre-trained data-driven model may be trained based on general data to provide output data based on being provided by input data. In an embodiment, the data-driven model may be a fine-tuned data-driven model. The fine-tuned data-driven model may be trained based on general data to provide output data based on being provided by input data, preferably in a first training process. The fine-tuned data-driven model may be further trained, preferably in a second training based on historical data related to chemical products and properties thereof to provide retrieval requests and / or, respectively, analysis instructions based on being provided by product requests and / or, respectively, measurement data. General data may comprise historical input data and / or historical output data. Input data and / or output data may refer to data processable by the data-driven model, preferably by the data-driven model as parametrized and / or initialized. The input data and / or the output data may comprise and / or may represent one or more sequences of elements, wherein a sequence of elements comprises two or more elements. In particular, the sequence of elements may be indicative of a sequence of the two or more elements of the sequence.

[0043] Training the data-driven model may refer to, and / or the training process may be a process of, building the data-driven model, in particular determining and / or updating parameters of the data-driven model. During the training process, the data-driven model may adjust to achieve best fit with the training data set, e.g. relating the at least one input value with best fit to the at least one desired output value. For example, if the neural network is a feedforward neural network such as a CNN, a backpropagation-algorithm may be applied for training the neural network. In case of a RNN, a gradient descent algorithm may be employed fortraining purposes. Gradient descent algorithm may use gradient for updating parameters. Gradient may indicate the degree of change for a parameter of the data-driven model. The gradient may be obtained by backpropagation. Thus, gradient descent algorithm may be based on backpropagation. A training process may be terminated when a deviation of the output generated by the data-driven model in comparison to a target output specified by the training data set falls within a predetermined range. The determining and / or updating of parameters of the data-driven model may be terminated when the training process may be terminated. The output generated by the data-driven model may be a retrieval request and / or, respectively, a set of analysis instructions. The target output specified by the training data set may be a retrieval request and / or, respectively, a set of analysis instructions. During the training process and / or the training of the data-driven model the training data set may comprise one or more sequences of elements and the one or more sequences of elements may be provided to the data-driven model sequentially and / or the data-driven model may generate the output sequentially. Hence, the data-driven model may generate a first element of the one or more sequences of elements based on a product request and / or, respectively, a set of measurement data and may generate further elements such as a second element of the one or more sequences based on one or more previously generated elements and the product request and / or, respectively, the set of measurement data. A second training process following a first training process may be advantageous to tailor the data-driven model to the use cases of retrieving measurement data from a given data management system and / or determining one or more target property values associated with a character product with a given set of analysis engines.

[0044] In an embodiment, the transformer encoder may comprise an encoder input, one or more encoder blocks and / or an encoder output. The encoder input may generate an embedded product request and / or, respectively, embedded measurement data based on receiving the product requests and / or, respectively, measurement data in a non-embedded form. The one or more encoder blocks may generate a context tensor based on receiving the embedded product request and / or measurement data, respectively. The context tensor may be indicative for a context associated with the product request and / or the measurement data, respectively. The encoder output may generate an encoded version of the respective input based on receiving the context tensor, wherein the encoded version of the respective input may refer, for instance, to an element in a sequence of elements to be provided by the transformer encoder. The transformer decoder may comprise a decoder input, one or more decoder blocks and / or a decoder output. The decoder input may generate an embedded product request and / or embedded measurement data, respectively (i.e., a further embedded version thereof) based on being receiving the output of the transformer encoder. The one or more decoder blocks may generate a context tensor based on receiving the embedded product request and / or measurement data as generated by the decoder input. Also this context tensor may be indicative for a context associated with the product request and / or the measurement data, respectively (i.e., a further kind of context indication, which may be different from the one generated and / or indicated by the encoder blocks). The decoder may generate a decoded version of the output of the output of the respective input based on receiving the context tensor, wherein the decoded version of the respective input may refer, for instance, to an element in a sequence of elements to be provided by the transformer decoder. The transformer encoder-decoder may comprise a transformer encoder comprising an encoder input and one or more encoder blocks and a transformer decoder comprising a decoder input, one or more decoder blocks and a decoder output.

[0045] In an embodiment, the method may further comprise embedding the retrieved product request and / or, respectively, the retrieved measurement data by the data-driven model, preferably in response to being provided by the retrieved product request and / or, respectively, the retrieved measurement data. Embedding the retrieved product request and / or, respectively, the retrieved measurement data may refer to applying input embedding and / or positional encoding to the retrieved product request and / or, respectively, the retrieved measurement data. Embedding the retrieved product request and / or, respectively, the retrieved measurement data may comprise transforming the retrieved product request and / or, respectively, the retrieved measurement data into a machine-processable form of the retrieved product request and / or, respectively, the retrieved measurement data, such as a tensor. The data-driven model may process the machine-processable form of the retrieved product request and / or, respectively, the retrieved measurement data into a retrieval request and / or, respectively, analysis instructions. By doing so, the retrieved product request and / or, respectively, the retrieved measurement data may be received in user language and accurate, robust and efficient processing of the user input may be processed by the data-driven model. This enables a barrier-free usage of the data-driven model to generate the retrieval request and / or, respectively, analysis instructions. In turn, providing access to all users provides the opportunity to all users to deploy the methods and systems as presented herein. This will increase the number of issues treated by the methods and systems resulting in more tailored analysis results associated with the target product while lowering the amount of errors in the one or more target property values associated with the target product. Hence, resources for determining one or more properties of chemical products are used more efficiently and with less errors.

[0046] Additionally or alternatively, embedding the received product request and / or, respectively, the retrieved measurement data by the data-driven model may comprise passing the received product request and / or, respectively, the retrieved measurement data by the data- driven model through an embedding layer. The embedding layer may be suitable for transforming the received product request and / or, respectively, the retrieved measurement data into a machine-processable format. The machine-processable format may referto a numberbased, in particular tensor-based representation of the received product request and / or, respectively, the retrieved measurement data. Embedding the received product request and / or, respectively, the retrieved measurement data may result in an embedded received product request and / or, respectively, embedded retrieved measurement data. The embedded received product request and / or, respectively, the embedded retrieved measurement data may comprise a tensor representing the received product request and / or, respectively, the retrieved measurement data. The tensor-based representation of the received product request and / or, respectively, the retrieved measurement data may be referred to as an embedded received product request and / or, respectively, embedded retrieved measurement data. The embedded received product request and / or, respectively, the embedded retrieved measurement data may be referred to as machine-processable format of the received product request and / or, respectively, the retrieved measurement data.

[0047] In an embodiment, the method may further comprise processing the embedded received product request and / or, respectively, the embedded retrieved measurement data to a context tensor by the data-driven model. Processing the embedded received product request and / or, respectively, the embedded retrieved measurement data to the context tensor may refer to transforming the embedded received product request and / or, respectively, the embedded retrieved measurement data to the context tensor. Preferably the embedded received product request and / or, respectively, the embedded retrieved measurement data may be transformed into the context tensor by forming one or more tensor products associated with the embedded received product request and / or, respectively, the embedded retrieved measurement data, and / or one or more representations associated with the embedded received product request and / or, respectively, the embedded retrieved measurement data. The one or more representations associated with the embedded received product request and / or, respectively, the embedded retrieved measurement data may be obtained by applying one or more mathematical operations to the embedded received product request and / or, respectively, the embedded retrieved measurement data. One or more mathematical operations may include for example, summing, subtracting, dividing, integrating, forming a derivative, multiplying, normalizing or a combination thereof.

[0048] In an embodiment, the embedded received product request and / or, respectively, the embedded retrieved measurement data may be sequence-specific. The sequence-specific embedded received product request and / or, respectively, the embedded retrieved measurement data may refer to a representation of the embedded received product request and / or, respectively, the embedded retrieved measurement data taking the sequence of data points associated with the embedded received product request and / or, respectively, the embedded retrieved measurement data into account. The sequence-specific embedded received product request and / or, respectively, embedded retrieved measurement data may be obtained by applying input embedding and / or positional encoding to the embedded received product request and / or, respectively, the embedded retrieved measurement data. For example, applying positional encoding may refer to adding and / or multiplying the one or more parts of the embedded received product request and / or, respectively, the embedded retrieved measurement data by a positional factor indicative of the position of the one or more parts of the embedded received product request and / or, respectively, the embedded retrieved measurement data within the embedded received product request and / or, respectively, the embedded retrieved measurement data.

[0049] In an embodiment, the provided product request may be sentence-based. Followingly, the data-driven model used for acquiring retrieval requests may be a natural language processing model. The sentence-based product request may be embedded resulting in the embedded product request. The sentence-based product request may be embedded by means of input embedding, preferably word embedding. Word embedding may referto transforming the sentence-based product request into a machine-processable format.

[0050] In an embodiment, the retrieved measurement data may be number-based and may hence. Nevertheless, the measurement data may be embedded, resulting in the embedded measurement data, such as by means of input embedding, preferably number embedding. Number embedding may refer to transforming the number-based measurement data into another machine processable format.

[0051] In an embodiment, the retrieval request and / or the analysis instructions, respectively, may comprise two or more parts, specifically a first part of the retrieval request and / or the analysis instructions, respectively, and a second part of the retrieval request and / or the analysis instructions, respectively. Where the retrieval request and / or the analysis instructions, respectively, may be sentence-based, the two or more parts may refer to words, numbers and / or parts of a word. A first part of the retrieval request and / or the analysis instructions, respectively, may refer to the first part in the sequence of retrieval request and / or the analysis instructions, respectively, and a second part of the retrieval request and / or the analysis instructions, respectively, may refer to the second part in the sequence of the retrieval request and / or the analysis instructions, respectively. The first part may be determined at a first time step and the second part may be determined at a second time step.

[0052] First part of the retrieval request and / or the analysis instructions, respectively, may refer to one element associated with the retrieval request and / or the analysis instructions, respectively, preferably the first element associated with the retrieval request and / or the analysis instructions, respectively. Second part of the retrieval request and / or the analysis instructions, respectively, may referto one element associated with the retrieval request and / or the analysis instructions, respectively, preferably the second element associated with the retrieval request and / or the analysis instructions, respectively.

[0053] In the first time step, the first part of the retrieval request and / or the analysis instructions, respectively, may be determined based on the embedded product request and / or the embedded measurement data, respectively. In the second time step, the second part of the retrieval request and / or the analysis instructions may be determined based on the embedded product request and / or the embedded measurement data, and further based on an embedded first part of the retrieval request and / or the analysis instructions. Embedded first part of the retrieval request and / or the analysis instructions may refer to a first part of the retrieval request and / or the analysis instructions, respectively, embedded analogous to the product request and / or the measurement data, respectively. Followingly, the second part of the retrieval request and / or the analysis instructions, respectively, may be generated based on the product request and / or the measurement data, respectively, and at least partially based on at least a part of the generated (or determined, or provided) retrieval request and / or analysis instructions, respectively, preferably by further providing at least a part of the generated (or determined, or provided) retrieval request and / or analysis instructions, respectively, to the data-driven model, in particular providing at least a first part of the generated (or determined, or provided) retrieval request and / or analysis instructions, respectively, to the data-driven model.

[0054] It may be preferred that the product request is indicative of a chemical reaction to which the target property relates. In this way, measurement data associated with the target product and suitable for determining based thereon the one or more target property values can be retrieved particularly efficiently. For instance, only measurement data acquired from a chemical reaction as indicated by the product request might be considered. The retrieval request providing model may, for instance, be trained to provide retrieval requests based on which only a subset of all data sets managed by the data management system may need to be searched for measurement data, wherein the subset may be defined as corresponding to all sets of measurement data acquired during chemical reactions as indicated by the product request. In particular, the product request may indicate the chemical reaction in terms a solvent, a reagent, an educt and / or a byproduct, and / or an amount of any of the foregoing. Moreover, as part of or additionally to the indication of the chemical reaction, the product request may indicate conditions associated with the chemical reaction, such as a temperature and / or a pressure. Instead of a chemical reaction, also a physical or chemical test carried out on the target product may be indicated by the product request. Additionally or alternatively, it may be preferred that the product request indicates the target product in terms of a digital representation of a chemical structure of the chemical product. Also this can allow for an efficient retrieval of the measurement data in terms of a corresponding retrieval request. For instance, the retrieval request providing model may be trained to provide retrieval requests based on which only a subset of all data sets managed by the data management system may need to be searched for measurement data, wherein the subset may be defined as corresponding to all sets of measurement data stored in association with the chemical product structure indicated by the product request. In particular, the product request may indicate the chemical structure in terms of a digital representation such as, for instance, SMILES, SMARTS, a graph structure or one or more mol files.

[0055] Furthermore, the product request may indicate a class of target products in terms of respective one or more values of one or more properties of chemical products in the class. As already indicated further above, this may be useful in situations where not a particular chemical product is of interest, but a target product among a class of chemical products with similar properties in some regard is still to be found based on properties in another regard.

[0056] In an embodiment, the data management system may include a plurality of databases. It may then be preferred that the model trained to relate product requests to retrieval requests is trained such that it relates product requests to one or more retrieval requests for one or more selected databases of the data management system. Also this can allow for a more efficient retrieval of relevant measurement data, since less data has to be searched and / or data can be searched in parallel. For instance, databases for which retrieval requests are to be acquired and hence to be provided by the model may be selected by a user via a user interface. However, the one or more selected databases may also be selected based on the product request. The product request may, for this purpose, indicate, or more generally be associated with, a selection of one or more databases. Separate databases may be installed for different kinds of properties of chemical products. For instance, a viscosity database and a conversion ratio database may be installed, wherein a retrieval request requesting the data management system to search for relevant measurement data in the viscosity database may be provided in case the product request indicated that the target property is a viscosity, and / or a retrieval request requesting the data management system to search for relevant measurement data in the conversion ratio database may be provided in case the product request indicated that the target property is a conversion ratio. A database may be selected based further on a specification of the database, wherein the specification of a database may be indicative of a type of measurement data stored in the database and / or one or more properties associated with the measurement data stored in the database. Particularly if the data management system includes a plurality of databases it may also be preferred that the retrieval request for retrieving the measurement data associated with the target product is acquired by providing the received product request to the data-driven model trained to relate product requests to retrieval requests in an adapted form, wherein the adapted form of the received product request is indicative of the product request and data retrieval specifications for the plurality of databases. An adaptation of the product request before providing it, in adapted form, to the data-driven model can lead to retrieval requests being acquired with a significantly increased quality, i.e., to retrieval requests allowing for a significantly more precise retrieval of relevant measurement data. The adapted form of the product request may be built by enhancing the product request with a prompt template. The prompt template may, for instance, include an indication of at least one of the target property and the target product, and furthermore data retrieval specifications for the plurality of databases.

[0057] Similarly, particularly if the analysis engine is one of a plurality of analysis type-specific analysis engines for carrying out respective specific types of analyses, it may be preferred that the analysis instructions are indicative of a respective analysis engine to be used for analyzing the retrieved measurement data. Also this has been found to allow for an increased efficiency in determining the one or more target property values.

[0058] If not the analysis instructions indicate the analysis engine to be used, a selection of the analysis engine to be used may be made based on, for instance a user input provided via a user interface. Moreover, an analysis engine may be selected based on the product request, wherein, for this purpose, the product request may indicate, or generally be associated with, a type of analysis suitable for determining the one or more target property values. The analysis engine may then be further selected based on a specification of the analysis engine, which may indicate one or more types of analyses that can be carried out with the analysis engine.

[0059] In particular, the analysis instructions providing model may be trained to relate the retrieved measurement data to analysis engine-specific analysis instructions. Hence, the analysis instructions may indicate an analysis engine to be used for analyzing the retrieved measurement data and may furthermore be specific for the indicated analysis engine. Analysis engines tend to be rather specific regarding how they are configured to receive analysis instructions. Therefore, by training the analysis instructions providing model to provide the analysis instructions already in an analysis engine-specific form, the efficiency in determining the one or more target property values can be further increased. The analysis instructions may be acquired by providing the retrieved measurement data to the data-driven model trained to relate measurement data to analysis instructions in an adapted form, wherein the adapted form of the measurement data is indicative of the product request and analysis specifications for the plurality of analysis engines. An adaptation of the measurement data before providing it, in adapted form, to the data-driven model can lead to analysis instructions being acquired with a significantly increased quality, i.e., to analysis instructions allowing for a more valuable analysis of the measurement data. The adapted form of the measurement data may be built by enhancing the measurement data with a prompt template. The prompt template may, for instance, include an indication of the target property and specifications regarding functionalities of the analysis engines.

[0060] In an embodiment, analysis instructions are acquired for several of the analysis type-specific analysis engines, wherein the one or more target property values are determined by analyzing the retrieved measurement data with the several analysis-specific analysis engines. Hence, the retrieved measurement data can be analyzed in several different ways with respect to a target property, which can lead to a more accurate determination of the one or more target property values. The analysis of the retrieved measurement data with the several analysis type-specific analysis engines may lead to several analysis results, i.e., for instance, several preliminary versions of the one or more target property values. One or more final target property values may be determined based on the several preliminary versions of the one or more target property values by, for instance, selecting one of the preliminary versions based on a confidence score being assigned to each of the preliminary versions, or by averaging the preliminary versions. The final target property values may then be provided as output to a user, such as in terms of being displayed on a display. Instead of determining a further, final version of the one or more target property values, also the preliminary versions may be considered as final and may be provided as output to the user. Hence, the analysis results from all of the several analyses carried out with the several analysis type-specific analysis engines may be displayed, for instance.

[0061] Optionally, the analysis instructions are indicative of a part of the respective measurement data to be analyzed. This can again increase the efficiency of the analysis, since the respective analysis engine does not need to analyze other parts then the indicated part of the measurement data. In an example, the analysis instructions may indicate a spectral region known to carry spectral information relevant to the target property.

[0062] It may be preferred that the retrieval of the measurement data from the data management system based on the retrieval request and / or the selection of one or more analysis engines to be used for analyzing the retrieved measurement data is further based on a respective similarity search. For instance, sets of measurement data in a database managed by the data management system may be compared, optionally in terms of metadata associated with the respective measurement data set, to one or more parts of a retrieval request based on a similarity measure, wherein those measurement data sets may be retrieved for which the similarity measure indicates a degree of similarity with the one or more parts of the retrieval request beyond a predefined threshold. Similarly, analysis engine specifications may be compared with one or more parts of analysis instructions based on a similarity measure, wherein those analysis engines may be selected for analysis of the retrieved measurement data for which the similarity measure indicates a degree of similarity with the one or more part of the analysis instructions beyond a predefined threshold. In both cases, the similarity measure may be defined, for instance, in an embedding space into which the elements to be compared may be embedded. The embedding space may refer to a vector space, for instance. Hence, for example, the measurement data sets managed by the data management system and stored in one or more databases may be represented within a same, possibly vector-type, embedding space as one or more parts of a retrieval request, wherein the degree of similarity between the two may be determined between the respective representations. Similarly, the analysis engines may be represented in a same, possibly vector-type, embedding space as one or more parts of the analysis instructions, wherein the degree of similarity between the two may be determined between the respective representations.

[0063] The one or more analysis engines, which may be operating engines, may be configured to provide one or more target properties in response to receiving measurement data. Measurement data may be provided to an analysis engine and hence be received by an analysis engine, for instance, via an application programming interface (API). An analysis engine may correspond to or be configured to run a script for processing the received measurement data, wherein the processing may involve an analysis of the measurement data with respective to one or properties predefined for a respective analysis engine.

[0064] As already indicated above, one or more analysis engines of a plurality of available analysis engines may be selected to be used for analyzing the retrieved measurement data. In particular, the selection may be made using the analysis instructions providing model based on a technical specification of the analysis engines. The technical specification may be provided to the analysis instructions providing model together with the retrieved measurement data, wherein the analysis instructions providing model may be trained to select, in response thereto, the one or more analysis engines to be used, i.e., as part of or in addition to the analysis instructions. The technical specification may, for instance, be a list of available analysis engines in which for each analysis engine a data structure is indicated according to which the analysis engine is configured to receive input data. The analysis instructions are preferably related to the measurement data and are preferably provided as input to an analysis engine. The analysis instructions may be formatted by the analysis instructions providing model. If the analysis instructions refer to an analysis engine capable to carry out several different types of analyses, the analysis instructions may include an identifier indicating which analysis of the several types of analyses is to be carried out by the analysis engine. The analysis engine may be selected based on the retrieved measurement data, the product request and a functional (or, as indicated above, “technical”) specification of the available analysis engines. In fact, however, the selection of the one or more analysis engines could again be made by a trained data-driven model, which may be the same as the retrieval request providing model and / or the analysis instructions providing model, or which may be a further data-driven model.

[0065] As indicated above, any data-driven model referred to herein may be a generative model. Furthermore, any data-driven model referred to herein may be large language model (LLM).

[0066] As also mentioned above, the retrieved measurement data should be suitable for determining based thereon the one or more target property values. If this is not the case, additionally or alternatively to the retrieval request, an insufficiency indication may be provided by any of the one or more data-driven models. Hence, for instance, the retrieval request providing model may be trained to provide, in response to being provided with a product request, an indication of whether the measurement data accessible from the data management are suitable for determining the one or more target property values. Such an indication may be based, for instance, on the product request and a specification of the data management system and / or the databases. Additionally or alternatively, the analysis instructions providing model may be trained to provide, in response to being provided with retrieved measurement data, an indication of whether and / or in how far the measurement data are suitable for determining based thereon the one or more target property values. If the retrieved measurement data are not suitable for determining based thereon the one or more target property values, which may be the case for certain product requests or if, exceptionally, the retrieval request has been inaccurate, then optionally no analysis instructions are provided by the analysis instructions providing model.

[0067] It may be preferred that the method further includes a step of acquiring a response to the product request by providing analysis results of the several analysis engines to a data-driven model trained to relate analysis results from analysis engines to responses to product requests. In this way, an output of the method can be provided in a more useful form. For instance, it may not be of particular assistance for a user if he / she is provided with the raw analysis results from the analysis engines. Instead, he / she may be provided with a response to the product request that summarizes the analysis results in a form tailored to the use case intended by the user.

[0068] The response to the product request may be acquired by providing the analysis results of the several analysis engines to the data-driven model trained to relate analysis results from analysis engines to responses to product requests in an adapted form, wherein the adapted form of the analysis results may be indicative of the product request and output specifications forthe response to the product request. An adaptation of the analysis results before providing them, in adapted form, to the data-driven model can lead to responses being acquired that are particularly useful for a user or for further processing. The adapted form of the analysis results may be built by enhancing the analysis results with a prompt template. The prompt template may, for instance, include an indication of the target property, the target product and specifications regarding a required or desired output in response to product requests.

[0069] Regarding the data-driven model trained to relate analysis results to responses to product requests, the remarks made above regarding the relation of the data-driven model trained to relate measurement data to analysis instructions to the data-driven model trained to relate product requests to retrieve requests apply. Hence, in particular, the data-driven model trained to relate analysis results to responses to product requests can be a further data- driven model or the data-driven model trained to relate product requests to retrieval requests and / or measurement data to analysis instructions. The data-driven model trained to relate analysis results to responses to product requests may be referred to as response providing model, wherein, again, this term may refer to a further data-driven model or a further functionality of the one or more data-driven models described above.

[0070] Also regarding the training, the remarks made above regarding the retrieval request providing model and the analysis instructions providing model apply analogously. Hence, the response providing model may be trained as described above regarding the retrieval request providing model and the analysis instructions providing model, a possible difference being that the training data used for training the response providing model may comprise or correspond to analysis results and responses to product requests instead of product requests and retrieval requests and / or measurement data and analysis instructions.

[0071] In an aspect, the invention also relates to a training method for training a data-driven model to relate a) product requests to retrieval requests, b) measurement data to analysis instructions, and / or c) analysis results to responses to product requests. The training method includes providing training data comprising training input data and training output data, wherein a) the training input data comprise a plurality of training product requests for determining one or more target property values associated with a target property of a target product and the training output data comprise a plurality of training retrieval requests associated with respective training product requests, the training retrieval requests being retrieval requests for retrieving, from a data management system, measurement data associated with the respective target property of the respective target products, and / or b) the training input data comprise a plurality of training measurement data sets associated with target properties of target products and the training output data comprise a plurality of training analysis instructions associated with respective training measurement data sets, the training analysis instructions being analysis instructions for instructing one or more analysis engines to analyze the respective training measurement data sets with respect to the respective target properties, and / or c) the training input data comprise a plurality of training analysis results from respective analysis engines and the training output data comprise a plurality of training responses to product requests associated with the analysis results. The training method further includes a step of training the data-driven model based on the training data.

[0072] The training method may be a supplementary training method, carried out supplementary after a foundational training of the data-driven model. For instance, the data-driven model may previously have been trained with more general training data, wherein the supplementary training can then be understood as a fine-tuning of the foundationally trained data-driven model to the use case of relating a) product requests to retrieval requests, b) measurement data to analysis instructions, and / or c) analysis results to responses to product requests.

[0073] It shall be understood that any features mentioned above regarding the use of the respective one or more data-driven models, such as any features of the product requests, the data management system, the retrieved measurement data and / or the analysis engines, and particularly the adaptations of the product requests, the measurement data and the analysis results before being provided to the respective data-driven model, may also apply to the training, particularly the training data and the use of the training data for training the respective data-driven model.

[0074] In a further aspect, the invention relates to a data processing system comprising a processor configured to carry out the steps of any of the above methods.

[0075] The inventions also relates in an aspect to a use of any of the methods above and / or the above data processing system for determining one or more properties of one or more chemical products. It shall be understood that the method of claim 1 , the system of claim 14 and the use according to claim 15 have similar and / or identical preferred embodiments as defined in the dependent claims.

[0076] It shall be understood that a preferred embodiment of the invention can also be any combination of the dependent claims with the respective independent claim.

[0077] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter.

[0078] In the following, the present disclosure is further described with reference to the enclosed figures. The same reference numbers in the drawings and this disclosure are intended to refer to the same or like elements, components, and / or parts.

[0079] BRIEF DESCRIPTION OF THE DRAWINGS

[0080] FIG. 1 illustrates an embodiment of a method for determining one or more properties of one or more chemical products.

[0081] FIG. 2 illustrates an embodiment of a data processing system for determining one or more properties of one or more chemical products.

[0082] FIG. 3 illustrates an embodiment of training an embedding layer.

[0083] FIG. 4 illustrates an embodiment of a transformer encoder architecture.

[0084] FIG. 5 illustrates an embodiment of a transformer decoder architecture.

[0085] FIG. 6 illustrates an embodiment of a transformer encoder-decoder architecture.

[0086] FIG. 7 illustrates an embodiment of training and / or deploying the transformer encoder, the transformer decoder and / or the transformer encoder-decoder.

[0087] FIG. 8 illustrates an embodiment of input embedding.

[0088] FIG. 9 illustrates an embodiment of input embedding. DETAILED DESCRIPTION OF EMBODIMENTS

[0089] The following embodiments are mere examples for implementing the method and system disclosed herein and shall not be considered limiting.

[0090] FIG. 1 illustrates schematically and exemplarily a method 100 for determining one or more properties of one or more chemical products. The method is initiated by a step 101 of receiving a product request for determining one or more target property values associated with a target property of a target product. In the illustrated embodiment, the product request is provided by a user via a user interface, possibly in free text form. In a following subsequence of steps 102, a retrieval request for retrieving, from a data management system 20 (shown in FIG. 2), measurement data associated with the target property and suitable for determining based thereon the one or more target property values is acquired. The retrieval request is acquired by providing the received product request to a data-driven model 10 (shown in FIG. 2) trained to relate product requests to retrieval requests. In the illustrated case, an adapted form of the product request is provided to the model 10, wherein the adapted form is indicative of the product request and data retrieval specifications for a plurality of databases included in the data management system 20. The model 10 is trained to relate product requests to one or more retrieval requests for one or more selected databases of the data management system 20. The databases may be selected based on an indication of the target property and / or the target product included in the product request, and further based on specifications of the databases regarding types of measurement data stored therein.

[0091] The acquired retrieval request, i.e., the retrieval request as received from the model 10, is subsequently used in a subsequence of steps 103 for retrieving measurement data associated with the target product from the data management system 20 in order to, in a further subsequence of steps 104a, 104b of the method 100, determine the one or more target property values based on the retrieved measurement data.

[0092] In the illustrated case, the retrieval request is forwarded to the data management system 20, wherein the data management system is configured to respond thereto by providing measurement data from one or more of its databases that "match" the retrieval request (steps 103). Which measurement data "match" a retrieval request can be defined depending on a similarity measure between parts of the retrieval request and stored measurement data sets. For instance, a degree of similarity can be determined between a) parts of the retrieval request indicating the target property and / or the target product and b) parts of the measurement data sets or metadata associated therewith. Moreover, in the illustrated case, the one or more target property values are determined based on the retrieved measurement data by acquiring analysis instructions for analyzing the retrieved measurement data with respect to the target property using one or more analysis engines 40 (shown in FIG. 2), the analysis instructions being acquired by providing the retrieved measurement data to a data-driven model 30 (shown in FIG. 2) trained to relate the measurement data to analysis instructions, and by subsequently instructing the one or more analysis engines 40 to analyze the retrieved measurement data with respect to the target property according to the acquired analysis instructions (steps 104 b).

[0093] The retrieved measurement data are provided to the data-driven model 30 in an adapted form, wherein the adapted form is indicative of the product request and analysis specifications for the one or more analysis engines 40. Moreover, the measurement data may be provided to the data-driven model 30 in combination with an indication of the target property of the target product and an indication of an input structure of the one or more analysis engines 40.

[0094] The one or more analysis engines 40 may be analysis type-specific analysis engines for carrying out respective specific types of analyses. In that case, the analysis instructions may be indicative of a respective analysis engine to be used for analyzing the retrieved measurement data. Where several analysis engines are to be used, the analysis instructions may be specific for the respective analysis engines used. Preferably, analysis instructions are acquired for several of the analysis engines 40, wherein the one or more target property values are determined by analyzing the retrieved measurement data with the several analysis engines 40.

[0095] The input structure of the one or more analysis engines 40 can be such that they are configured to receive measurement data to be analyzed in combination with analysis instructions. Hence, as illustrated in FIG. 1 , the analysis results may be obtained by providing the retrieved measurement data in combination with the analysis instructions to the one or more analysis engines 40.

[0096] Since it may not be desired to provide a raw output of the one or more analysis engines 40 back to the user, the method 100 further includes a subsequence of steps 105 for acquiring a response to the product request based on the analysis results. The response is acquired by providing the analysis results of the several analysis engines to a data- driven model 50 (shown in FIG. 2) trained to relate analysis results from analysis engines 40 to responses to product requests. The thus required response is then provided as output to the user, such as in text form, for instance.

[0097] In the illustrated case, the analysis results of the several analysis engines 40 are provided to the trained model 50 in an adapted form for acquiring the response to the product request, wherein the adapted form is indicative of the product request and output specifications for the response to the product request. In particular, the analysis results may be provided to the trained model 50 in combination with a structure format in which the response is to be provided as output to the user, wherein the structure format may be predefined based on the product request. For instance, if the product request is in text form, also the response may be provided in text form, and / or if the product request indicates the target property in a particular system of units, also the response may do so.

[0098] FIG. 2 shows schematically and exemplarily a data processing system 200 for carrying out the method 100. The data processing system 200 comprises the data-driven models 10, 30, 50, the data management system 20 and the one or more analysis engines 40 already indicated above regarding FIG. 1. Moreover, the data processing system 200 comprises a user interface 60 by which, as also already indicated above with respect to FIG. 1 , the product request can be received from a user and the response thereto can be provided to the user.

[0099] The data processing system 200 further comprises an orchestrating agent 70. In fact, the steps of the method 100 may be carried out solely by the orchestrating agent 70. For doing so, the orchestrating agent 70 may be configured to communicate with the further elements 10, 20, 30, 40, 50, 60 of the system 200, such as by providing respective input data and receiving respective output data from the respective elements. For instance, the orchestrating agent 70 may be configured to provide the product request to the retrieval request providing model 10, to provide the retrieved measurement data to the analysis instructions providing model 30, to provide the retrieved measurement data and the analysis instructions as received from the model 30 to the one or more analysis engines, et cetera. Nevertheless, the orchestrating agent 70 may be a data processing device in its own right.

[0100] FIG. 3 illustrates an embodiment of obtaining an embedding layer. The embedding layer may be obtained by training for example a continuous bag of words model (CBOW) or a skip-gram model. The embedding layer may be suitable for generating embedded input data based on input data. The input data may be product requests, measurement data associated with product requests and / or analysis results associated with product requests as described with reference to FIG. 1. Generating embedded input data may refer to embedding input data. Embedding input data may result in a representation associated with the input data. Thus, the embedded input 114 may be the representation associated with the input data. The input data may comprise one or more elements. The one or more elements may be represented by the input vector 106. In particular, the embedded input 114 and / or the input vector 106 may be machine- readable and / or processable by a processor. For this purpose, the embedded input 114 and / or the input vector 106 may be a tensor, in particular a first-rank tensor. Specifically, the input vector 106 may be a one-hot vector or a summation of a plurality of one-hot vectors. A one- hot vector may be a vector with one entry unequal to zero. Examples for one-hot vectors may be 108, 110 and 112. The entries unequal to zero in the one-hot vector and / or in the input vector 106 may indicate the element. For example, a lookup table may define the relation between the position of the entries unequal to zero and the element indicated by the one-hot vector. The lookup table may specify a plurality of different elements. The number of different elements may be equal to the number of entries in the one-hot vector. The number of different elements may be referred to as vocabulary size. In an example, the elements may be represented by tokens and a sequence of elements may refer to at least a part of a sentence. The at least a part of the sentence may be represented by a plurality of tokens. A token may represent at least a part of the element and / or word. For example, where one element would be associated with only one word, words such as “embeddings", “embedding” or “embed” would constitute different elements. A first token may represent the stem “embed” and the endings, typically appearing in a plurality of words, may be represented by a second token, a third token and a fourth token. The second token, the third token and the fourth token may be used for representing other words such as “look”, “looking” or the like, preferably together with a fifth token representing the stem “look”. Ultimately, this tokenization of elements associated with a plurality of stems and a plurality of endings results in less tokens to be used for representing a plurality of elements and thus, uses less computational resources.

[0101] A lookup table specifying a subset of the vocabulary size e.g. of the English language may comprise 10,000 words or more. The embedded input 114 may be a lower-dimensional representation than the input vector 106. For example, typical embedded inputs 114 may comprise some hundreds of different entries. Followingly, the embedded inputs 114 constitute a densified representation of one or more elements using less computa- tional resources. More than that, the embedded input 114 may represent a relation between two or more elements. For example, the words “Italy” and “Germany” may be similar or may be more closely related since they both define European countries, whereas the word “embodiment” may be very different from the two respective words. The smaller the dot product between two embedded inputs 114 may be the more similar the two elements associated with the embedded inputs 114 may be. Hence, the embedded inputs 114 may represent one or more elements accurately and lead to accurate results based on processing the embedded inputs 114.

[0102] For transforming the input vector 106 into the embedded input 114, the embedding layer may comprise a number of neurons equal to the number of entries in the embedded input 114. Based on the embedded inputs 114, the output layer may generate the output vector 116. The output vector may be a vector and / or may indicate one or more elements. The output vector 116 may indicate one or more elements different from the input vector 106 and / or the one-hot vectors associated with the input vector 106. For this purpose, the output layer may comprise a number of neurons equal to the number of entries of the input vector 106 and / or the output vector 116. The output layer may apply a softmax function to the embedded inputs 114. By doing so, the output vector may comprise the probabilities associated with the elements associated with the entries of the output vector 116 unequal to zero. Hence, from the output vector 116 one or more elements may be obtained with a corresponding probability. Where the input vector 106 may specify one or more sequence(s) of elements, the output vector 116 may specify one or more elements corresponding to the sequence(s) of elements specified by the input vector 106. In the example of FIG. 3, the element associated with vector 118 may correspond to the input vector with a probability of 71 %. Additional or alternative elements may correspond to the input vector as indicated by the output vector with lower probability. By defining a threshold to which the probability may be compared, the selection of the corresponding elements may be tailored to the needs of the user. The elements generated by the model comprising the embedding layer 102 and the output layer 104 may refer to the most probable elements indicated by the output vector 116. Hence, the model depicted in FIG. 3 may generate the element associated with the vector 118 with a confidence score of 71 %.

[0103] The model of FIG. 3 may be a continuous bag of words (CBOW) model. The CBOW model may be trained based on a training data set comprising a plurality of input vectors and corresponding output vectors. As the training data set may not be labeled, the training of the CBOW model may be referred to as self-supervised. Before training of the CBOW model, the CBOW model may be initialized with random values assigned to the weights of the neurons. During the training of the CBOW model, the input vectors may be passed through the initialized embedding layer and the output layer and a loss may be determined by comparing the output vector obtained by passing the input vector 106 through the model to the output vector corresponding to the input vector 106 as specified by the training data set. Based on the determined loss, backpropagation may be applied to determine the gradients associated with the neurons of the embedding layer 102 and the output layer 104 to lower the loss. According to the determined gradients, the weights of the neurons may be updated by using a gradient descent algorithm. If a predetermined loss may be achieved by the CBOW model, the training may be terminated and a trained CBOW model may be obtained. From the trained CBOW model, the embedding layer 102 may be suitable for embedding input data comprising one or more elements. This embedding layer 102 may be used in other machine-learning architectures requiring an embedding layer 102 such as a transformer encoder, transformer decoder or transformer encoder decoder architecture as described within the context of FIG. 4, FIG. 5 and FIG. 6. For training these architectures, a trained embedding layer 102 may be required. Hence, a model such as a CBOW model may be trained prior to training the transformer encoder, transformer decoder or transformer encoder decoder architecture.

[0104] FIG. 4 illustrates an embodiment of a transformer encoder architecture. The transformer encoder comprises an encoder input 278, one or more encoder blocks 274, 214 and an encoder output. The transformer encoder architecture may be derived from the transformer encoder-decoder architecture as known in the art and shown in FIG. 6. In particular, the transformer encoder may be referred to as X-former. The transformer encoder architecture may correspond to the encoder architecture associated with the transformer encoder-decoder architecture with an additional encoder output instead of connecting the encoder block directly to the decoder of the transformer encoder-decoder architecture. A plurality of transformer encoder architectures are available in the art such as the bi-directional encoder representations from transformers (BERT).

[0105] The input data may be received at the encoder input 278. The encoder input 278 may apply an input embedding 202. Applying the input embedding 202 may refer to passing the input data through an embedding layer e.g. as described within the context of FIG. 3. Further, the encoder input 278 may apply positional encoding 204. Applying positional encoding 204 may refer to adding a positional factor to the embedded input obtained via input embedding. Preferably, the input data may specify a sequence of elements. The positional factor Ppos may be indicative of the position of the elements within the sequence. For example, the positional factor Ppos may be obtained based on the following equation: where pos may refer to the position of the element within the sequence, / may refer to the dimension associated with the input embedding and d may refer to the dimension of the model, e.g. transformer decoder, transformer encoder or transformer encoder-decoder. This may be referred to as absolute positional embeddings. Alternatively, the positional encoding may be based on rotary positional embeddings (RoPE). Positional encoding is beneficial since it enables the processing of sequential data without requiring further dimensions indicating the position of each element. Followingly, the positional encoding 204 reduces the computational resources needed for embedding the input data. By passing the input data through the encoder input, the input data may be transformed into a second-rank tensor representing the sequence of elements. This second-rank tensor may be referred to as embedded input data. The embedded input data may be processed by the encoder block. The embedded input data may be provided to the layer normalization 208 by a residual connection. Multi-head self-attention 206 may be applied to the embedded input data. Multi-head self-attention 206 may comprise the two components multi-head and self-attention. Self-attention may be understood as being a filter applied to the embedded input data. By applying the filter to the embedded input data, the elements associated with the embedded input data contributing to the to be generated output data may be identified for generating the output data. Hence, the filter may represent the degree of contributing to the to be generated output data by the elements associated with the embedded input data. Applying the filter may be referred to as weighting the elements associated with the embedded input data. This is advantageous specifically regarding long sequences of elements. The filter may be learned and improved during the training by learning to identify the contribution of elements associated with the embedded input data. For example, in the partial sentence “I went to the bakery to buy a” the last word may be generated by the data-driven model such as the transformer encoder. The self-attention may focus the transformer encoder to attend to the word “bakery” and “buy” mostly to generate the word “bread”. Selfattention may refer to attention generated based on the input data. Hence, the filter may be determined based on the input data, preferably the embedded input data. The embedded input data may serve as query Q, key K and value V with respect to the selfattention operation. The self-attention may refer to attention based on the received input data. Hence, the filter may be calculated based on the following formula by inserting the respective tensors based on the embedded input data: where dk corresponds to the dimension of the key.

[0106] For improving the efficiency of the transformer encoder further, the multiple heads are used to apply the filter resulting in the multi-head self-attention 206. Multi-head selfattention 206 may comprise applying the filter to two or more parts of the embedded input data. Hence, the tensor may be split into two or more parts and the filter may be applied to the two or more parts separately by two or more heads according to the following equation:head i =Attention QWiQ, KW , VWiV) with parameter matrices may refer to the number of heads, dy, d% and dqmay refer to the dimensions of the value, key and query.

[0107] The result of the two or more head may be concatenated according to the following equation: MultiHead(Q, K, V) = Concat(head 1, . . . , headh)W° e ^hdv*d and h may refer to the number of heads.

[0108] The embedded input data may be transformed via the multi-head self-attention 206 into a context tensor. The context tensor may represent the sequence of elements and the relation between two or more elements of the input data. The context tensor may be a second rank tensor and / or may comprise one or more first rank tensor(s). After the multihead self-attention 206 layer normalization 208 may be applied based on the context tensor and / or the embedded input data from the residual connection. Applying layer normalization 208 may refer to normalizing the context tensor. Normalizing the context tensor may lower the values of the entries of the context tensor. This reduces the computational cost associated with processing the context tensor. Further, it improves the training by contributing the loss to converge and preventing instabilities.

[0109] Layer normalization 208 may be followed by passing the context tensor to a feed-forward layer 210 again followed by layer normalization 212 based on the residual connection to the context tensor and / or the output of the feed-forward layer 210. The feedforward layer 210 may be a feed-forward neural network. The feed-forward neural network may comprise of a plurality of fully connected neurons. Passing the context tensor through the feed-forward neural network may result in transforming the context tensor linearly. Additionally or alternatively, the neural network may comprise one or more activation functions such as a rectified linear unit (ReLU). Hence, the neural network may be configured for performing one or more non-linear operations to the context tensor and / or transforming the context tensor non-linearly. After the context tensor has been transformed and / or normalized by the feed-forward layer 210 and the layer normalization 212, the context tensor may be provided to one or more further encoder blocks 214. Having passed the context tensor through the feed-forward layer 210 may adapt the context tensor for the processing by a further attention layer of the one or more further encoder blocks 214 for applying a self-attention filter, preferably multi-head self-attention 206. The context vector after being transformed by the layer normalization 212 and the feed-forward layer 210 may be referred to as hidden state.

[0110] The encoder output 276 comprises of a linear layer 216 and a softmax layer 218. The linear layer 216 may transform the context vector into a logits vector. The linear layer may be fully-connected. The logits vector obtained by passing the context tensor through the linear layer 216 may be passed through the softmax layer 218. Passing the logits vector through the softmax layer 218 may refer to applying the softmax function to the logits vector. Applying the softmax function to the logits vector may result in a probability distribution of one or more elements corresponding to the sequence of elements in the input data. From the probability distribution based on predefined selection criteria, one or more elements may be chosen. The one or more chosen elements may be referred to as the one or more elements generated by the transformer encoder. The one or more generated elements may be provided to the encoder input for generating further one or more elements corresponding to the sequence of the input data and the one or more elements generated by the transformer encoder as described within the context of FIG. 7.

[0111] FIG. 5 illustrates an embodiment of a transformer decoder architecture. The transformer decoder comprises a decoder input 284, one or more decoder blocks 280, 232 and a decoder output 292. The transformer decoder architecture may be derived from the transformer encoder-decoder architecture as known in the art and shown in FIG. 6. The transformer decoder may be referred to as X-former. The transformer decoder architecture may correspond to the decoder architecture associated with the transformer encoder-decoder architecture independent of receiving one or more hidden states from the encoder of the transformer encoder-decoder. A plurality of transformer decoder architectures are available in the art such as the generative pretrained transformers (GPT).

[0112] The decoder input 284 may apply input embedding 220 and positional encoding 222 analogous to the input embedding 202 and the positional encoding 204 as described within the context of FIG. 4.

[0113] The decoder block 280 may comprise the layer normalizations 226, the masked multihead self-attention 224, the feed-forward layers 228 and / or the layer normalization 230. The embedded input data resulting from passing the input data through the decoder input 284 may be provided to the layer normalization 226 via a residual connection. Further, masked multi-head self-attention 224 may be applied to the embedded input data. Masked multi-head self-attention 224 corresponds to the multi-head selfattention 206 as described within the context of FIG. 4 with additionally masking a part of the embedded input data associated with elements later in the sequence than the element to be generated. Additionally or alternatively, the part of the input data associated with elements later in the sequence than the element to be generated may not be received and / or transformed into the embedded input data. Thus, the transformer decoder may be suitable for generating a subsequent element to a sequence, whereas the transformer encoder may be suitable for generating a missing element in within one sequence and / or between two or more sequences. Therefore, the transformer encoder may be configured for classification tasks. The transformer decoder may be configured for text generation.

[0114] Similar to the transformer encoder as described within the context of FIG. 4, a context tensor may be generated by applying the masked multi-head self-attention 224 and the layer normalization 226. The context tensor may be provided to the layer normalization 230 via a residual connection. Further, the feed-forward layer 228 and the layer normalization 230 may be analogous to the feed-forward layer 210 and the layer normalization 212 as described within the context of FIG. 4. The context tensor may be provided to one or more further decoder blocks 232. The decoder output 292 may comprise of a linear layer 234 and a softmax layer 236. The linear layer 234 and the softmax layer 236 may be analogous to the linear layer 216 and the softmax layer 218 as described within the context of FIG. 4.

[0115] FIG. 6 illustrates an embodiment of a transformer encoder-decoder architecture. The transformer encoder-decoder may comprise the encoder input 288, the one or more encoder blocks 286, 264, the decoder input 294, the decoder block 290 and the decoder output 292. The encoder input 288 may correspond to the encoder input 278 of FIG. 4. The one or more encoder block 286, 264 may correspond to the one or more encoder blocks 274, 214 of FIG. 4. The decoder input 294 may correspond to the decoder input 284 of FIG. 5.

[0116] The decoder block 290 may comprise a masked multi-head self-attention 270, a layer normalization 272, a feed-forward layer 238 and a layer normalization 240 analogous to the masked multi-head self-attention 224, the layer normalization 226, the feed-forward layer 228 and the layer normalization 230 as described within the context of FIG. 5. The decoder block 290 may further comprise a multi-head self-attention 250 and a layer normalization 248. Analogous to the description of FIG. 5, the context tensor may be obtained from the masked multi-head self-attention 270 and the layer normalization 272. Multi-head self-attention 250 analogous to the multi-head self-attention 206 of FIG. 4 may be applied to the context vector obtained from the layer normalization 272 and the hidden states of the one or more encoder blocks 286, 264. Layer normalization 248 may be applied to the context vector obtained from the multi-head self-attention 250 and the context vector obtained from the layer normalization 272 provided via a residual connection. The context vector resulting from the layer normalization 248 may be processed via the feed-forward layer 238 and the layer normalization 240 analogous to the description of FIG. 5. The context vector resulting from the layer normalization 240 may be provided to further decoder blocks 242 analogous to the decoder block 290. The context vector obtained from the one or more decoder blocks 290, 242 may be provided to the decoder output 292. The decoder output 292 may correspond to the decoder output 282 of FIG. 5.

[0117] With the above-described architecture, the transformer encoder-decoder may receive and process input data at the encoder input 288 and the one or more encoder blocks 286, 264 and the decoder block 290 and the decoder output 292. Based on the input data, the transformer encoder-decoder may generate output data part by part or sequentially. The sequentially generated output data may be provided to and / or may be processed by the decoder input 294, the one or more decoder blocks 290, 242 and the decoder output 292. Preferably, a sequence may be provided to the encoder input 288 and after having generated at least a part of the output data, the decoder input 294 may be provided with at least the part of the elements of the output data already generated. By doing so, the next elements of the output data may be generated with a higher accuracy by taking the input data and the generated output data into account since more data is received by the transformer encoder-decoder may be received over time.

[0118] Because of the transformer encoder-decoder architecture, the transformer encoder-decoder may be configured for transforming a sequence into another representation of the sequence. An example for transforming one sequence into another representation may be translation of one sentence into another language. A plurality of transformer encoderdecoders are available in the art such as BART, T5 or the like.

[0119] In an embodiment, the layer normalization 208, 212 may be applied prior to the masked multi-head self-attention 224, multi-head self-attention 206 and / or the feed-forward layer 210 in the transformer decoder, the transformer encoder and / or the transformer encoder-decoder. By doing so, the computational resources for applying the multi-head self-attention 206 and / or the feed-forward layer 210 to the embedded input data and / or the context tensor may be decreased as the entries of the respective tensors may be lower after normalization.

[0120] In an embodiment, the decoder output 292 may comprise of a classification neural network, further feedforward layers, convolutional layers, fully connected layers or the like. For example, the transformer encoder-decoder may be configured for choosing between a plurality of options. For this purpose, the transformer encoder-decoder may be provided with three different input data sets and may classify the context vectors obtained from the one or more decoder blocks 290 via one or more linear layers. Followingly, the architecture may be extended depending on the use case to be solved.

[0121] FIG. 7 illustrates an embodiment of training and / or deploying the transformer encoder, the transformer decoder and / or the transformer encoder-decoder.

[0122] The encoder / decoder / encoder-decoder architecture 302 may correspond to the transformer decoder, the transformer encoder and / or the transformer encoder-decoder as described within the context of FIG. 4- FIG. 6. The output data generated by the encoder / decoder / encoder-decoder architecture 302 may comprise of one or more elements, in particular a sequence of elements. The previously generated elements of the output data may be provided as input for generating the next element in the sequence of the output data. The output data may be retrieval requests, analysis instructions and / or responses to product requests as described with reference to FIG. 1 .

[0123] In the example of FIG. 7, the input data, which may be product requests and / or measurement data, may comprise of N elements, in particular input tokens. An input token may be a token dedicated to be inputted into a data-driven model such as the transformer decoder, the transformer encoder or the transformer encoder-decoder. The output data to be generated may comprise of M elements. The encoder / decoder / encoder- decoder architecture 302 may generate one element of the output data based on receiving the input data and optionally previously generated elements of the output data at a timestep. Hence, for generating M elements M time steps are required. A time step comprises of providing input 310, 312, 314 to the encoder / decoder / encoder-decoder architecture 302 and receiving output data 304, 308, 306 from the encoder / decoder / en- coder-decoder architecture 302. In a first timestep, the input 310 may comprise of N input tokens. The N input tokens may be associated eg with N words, stems or endings. Preferably, the N input tokens may specify a question. One or more input tokens may specify the beginning of the sequence of tokens and / or the end of the sequence of tokens. The input 310 may be processed by the encoder / decoder / encoder-decoder architecture 302. Based on the input 310 at least a part of the output data 304 may be generated. The at least a part of the output data may comprise a first output token. In the next timestep, the generated first output token may be provided together with the input 312. Specifically, where the input 312 may be received by a transformer encoderdecoder the input tokens may be received at the encoder input 288 and the first output token may be received at the decoder input 294. Where the input 312 may be received by the transformer encoder, the input 312 may be received by the encoder input 278 and analogously regarding the transformer decoder and the decoder input 284. Based on the input 312, the output data 308 comprising the first output token and a second output token may be generated. Generating the output data 308 based on the input 312 may refer to generating the second token based on the first token and the N input tokens, wherein the first token may have been generated based on the N input tokens. This process may be repeated until the last token in the sequence of the output data 306 may be generated. Preferably, the last token may be an end token. The end token may terminate the generation of a further output token. Similarly, to the data processing during deployment of the encoder / decoder / encoder- decoder architecture 302, the encoder / decoder / encoder-decoder architecture 302 may be trained. The training data set may comprise a plurality of sequences comprising a plurality of elements. The sequences may be associated with the input data and / or the output data. Additionally or alternatively, the sequences may be independent of the input data and / or the output data. For example, where the input data and the output data may refer to chemical compositions represented via text, the training data set may comprise sequential text data independent of chemical compositions. In this example, the training data set may comprise sequences of words originating from a conversation. In an embodiment, the training data set may comprise at least partially input data sets and / or output data sets.

[0124] The training may be initialized by initializing the encoder / decoder / encoder-decoder architecture 302. In an embodiment, the parameters associated with the encoder / de- coder / encoder-decoder architecture 302 may be initialized randomly. Additionally or alternatively, the input embedding of the encoder / decoder / encoder-decoder architecture 302 may be obtained by training a CBOW model or a skip gram model as described within the context of FIG. 3. The trained embedding layer may be used during training. The parameters associated with the embedding layer may be kept constant and / or may be updated after a predefined number of training epochs. By doing so, the number of parameters to be updated is lower enabling a faster and less computational resources- consuming training. Further, the accuracy associated with the embedding layer may be constant and / or may be increased by avoiding error compensation in relation to the just initialized encoder / decoder / encoder-decoder architecture 302.

[0125] During the training of the encoder / decoder / encoder-decoder architecture 302, at least a part of the sequences of the training data set may be provided to the encoder / de- coder / encoder-decoder architecture 302 one by another and one or more elements may be generated based on the sequences of the training data set one by another. The elements generated based on the sequences may follow the elements of the parts of sequences the encoder / decoder / encoder-decoder architecture 302 may have been provided with. The generated one or more elements may be compared to the one or more elements following the at least a part of the sequences provided to the encoder / de- coder / encoder-decoder architecture 302 as specified by the training data set. Hence, during the training the encoder / decoder / encoder-decoder architecture 302 may generate a guess on the next element and the guess on the next element in a sequence may be compared to the ground truth specifying the actual next element according to the training data set. Based on the guess on the next element and the ground truth a loss may be determined. The loss may define the similarity between the guess on the next element and the ground truth. The loss may be determined by forming a vector dot product between the token associated with the one or more elements and the token associated with the ground truth. A loss unequal to zero may result in updating the parameters associated with encoder / decoder / encoder-decoder architecture 302. Preferably the parameters associated with the encoder / decoder / encoder-decoder architecture 302 may be independent of the embedding layer. For example, the parameters associated with the encoder / decoder / encoder-decoder architecture 302 may be weights of the neurons of the encoder / decoder / encoder-decoder architecture 302.

[0126] Based on the determined loss, backpropagation may be applied to determine the gradients associated with the parameters of the parameters associated with encoder / de- coder / encoder-decoder architecture 302 to lower the loss. According to the determined gradients, the parameters associated with the encoder / decoder / encoder-decoder architecture 302, preferably the weights of the neurons associated with the encoder / de- coder / encoder-decoder architecture 302, may be updated by using a gradient descent algorithm.

[0127] The training data set may be unlabeled. The sequences of elements within the training data set may inherently comprise the ground truth for determining the loss with respect to the one or more elements generated during the training of the encoder / decoder / en- coder-decoder architecture 302. Hence, the encoder / decoder / encoder-decoder architecture 302 may be trained self-supervised. This is advantageous since time and resources for creating a labeled training data set may be saved. Furthermore, this enables the usage of large training data sets associated with a size of several tera bytes. Consequently, the data-driven model may be accurate in generating elements of a sequence. In addition, the large training data set enables few shot predictions or even zero shot predictions. Hence, the data-driven models trained as described above are versatile contributing to saving resources needed for training and / or hosting a plurality of purpose-driven models such as convolutional neural networks. The training described above may be referred to as pretraining. The data-driven model may be configured for performing few shot or even zero shot predictions with respect to a plurality of use cases after pretraining. The performance of the data-driven model may be increased further by additional training referred to as finetuning.

[0128] FIG. 8 illustrates an embodiment of input embedding. Where the sequence of elements associated with the input data, preferably comprised in the input data, may be of one type, the input embedding 202, 220, 252, 266 as described within the context of FIG. 4 - FIG. 6 may be used. For example, a type of input data may be text where the elements may be associated with at least a part of a word, a punctuation character, a start token specifying the beginning of one or more sequences associated with the input data and / or the end token. In another example, the input data may be at least partially numerical. Hence, the input data may comprise a plurality of numbers. Numerical input data may be for example tabular data. Tabular data may specify one or more rows and / or one or more columns. Hence, the tabular data may comprise one or more cells, wherein the cells may be associated with one or more numerical values.

[0129] Numerical input data may require a different embedding than text input data. Input embeddings for numerical input data may comprise a token embedding, a positional embedding, a column embedding, a row embedding or a combination thereof.

[0130] Applying a token embedding to one or more elements, in particular tokens associated with the input data may result in a machine-processable representation associated with the one or more elements, in particular tokens. Applying the token embedding to one or more elements may refer to passing the one or more elements through the embedding layer, eg as described within the context of FIG. 3. Hence, token embeddings may specify the one or more elements, in particular tokens in a machine-processable representation. For example, the token embedding may transform a numerical value into a vector. This is advantageous since this representation can be enriched by further information such as the position of the token within the sequence and / or within a table associated with the sequence of tokens. The positional embedding may be analogous to the positional embedding as described within the context of FIG. 3, FIG. 4 - FIG. 6. Where the input data may be tabular data, column embedding may be applied. Applying a column embedding to one or more elements, in particular tokens associated with the input data may result in a machine-processable representation specifying the location of the one or more elements within a table 402, preferably within the columns of the table 402. Applying the column embedding may refer to adding a column factor to the input data embedded via token embeddings, in particular the embedded input data. The column factor may be the same for elements associated with the same column and / or may differ between two or more elements associated with different columns. Analogous, row embeddings may be applied where the input data may be tabular data. Applying a row embedding to one or more elements, in particular tokens associated with the input data may result in a machine-processable representation specifying the location of the one or more elements within a table 402, preferably within the rows of the table 402. Applying the row embedding may refer to adding a column factor to the input data embedded via token embeddings, in particular the embedded input data. The row factor may be the same for elements associated with the same row and / or may differ between two or more elements associated with different rows.

[0131] In an embodiment, input data may be at least partially numerical and at least partially text. Hence, the input data may comprise two or more types of data. A type of data may refer to a modality. Followingly, different embeddings may be applied to the input data. To parts of the input data comprising text the input embedding referred to in FIG. 3, FIG. 4 - FIG. 6 may be applied. To parts of the input data being numerical token embeddings, positional embeddings, column embeddings and row embeddings may be applied. Further, segment embeddings may be applied to the input data independent of the type of input data. The segment embedding may specify the type of input data one or more elements may be associated to. For example, if the input data comprises of text and numbers, the input data may comprise of two types of input data. Applying the segment embedding to the input data may refer to adding a segment factor to the input data, preferably the embedded input data and / or the input data after having applied the token embedding. The segment factor may specify the type of data associated with the one or more elements. The segment factor may be the same for one or more elements associated with the same type of input data and / or may differ between two or more elements associated with different types of input data.

[0132] Applying the token embedding, the positional embedding, the segment embedding, the column embedding, the row embedding or a combination thereof may result in embedded input data and / or may be the output of any one of the encoder input 278, 284, 288 or decoder input 284, 294. The data obtained by applying the token embedding, the positional embedding, the segment embedding, the column embedding, the row embedding or a combination thereof may be processed by the encoder block 274, 286, decoder block 280, 290, encoder output 276, decoder output 292, 282.

[0133] FIG. 9 illustrates an embodiment of input embedding.

[0134] Input data to the data-driven model, in particularto the encoder input and / orthe decoder input as described in the context of FIG. 4 - FIG. 6, may comprise image data. The data-driven model may be parametrized to receive image data. For processing image data as input data, the data-driven model may comprise one or more encoder blocks and / or one or more decoder blocks and / or one or more encoder outputs and / or one or more decoder outputs as described within the context of FIG. 4 - FIG. 6. FIG. 9 may show an embodiment of an encoder input and / or a decoder input. When processing image data, the encoder input and / or the decoder input of the data-driven model may be as described within the context of FIG. 9. The encoder input and / or decoder input may comprise one or more linear projection layers 514 for a linear projection of one or more images, preferably one or more partial images, more preferably a sequence of two or more partial images. The one or more linear projection layers 514 may be suitable for changing the dimension of the one or more received images, preferably one or more partial images, preferably passing the one or more images, preferably partial images, through the one or more linear projection layers 514 may result in applying image embedding, preferably partial image embedding to the one or more images and / or partial images.

[0135] Furthermore, when a sequence of two or more images and / or partial images may be received, positional embedding may be applied to the sequence, preferably by passing the sequence of one or more images and / or partial images through the one or more linear projection layers 514. Applying positional embedding may refer to adding a positional factor. The positional factor may be different depending on the position of the image and / or the partial image within the sequence. In particular, the positional factor added to a first element of the sequence may be different to the positional factor added to a second element of the sequence. The first element of the sequence may be a first image and / or first partial image. The second element of the sequence may be a second image and / or a second partial image.

[0136] The representation of the one or more images, preferably one or more partial images, may be obtained based on the following equation: wherexdass is the image class embedding 528 ,XNis the n-th image, in particular partial image in the sequence,z<i is the representation of the one or more images, preferably one or more partial images, (H,W) are the resolution of the image, in particular the image the partial images are generated on, C is the number of channels associated with the one or more image, in particular the one or more partial images and D is the dimension of the representation of the one or more images, preferably one or more partial images. Applying the partial image embedding may refer to forming the product ofXNpwith E above-described equation. Applying the positional embedding may refer to adding the factor -^pos according to the above-described equation. By doing so, text-based data, numerical data, tabular data, image data or the like may be processed by one data-driven model.

[0137] The present disclosure has been described in conjunction with preferred embodiments and examples as well. However, other variations can be understood and effected by those persons skilled in the art and practicing the claimed invention, from the studies of the drawings, this disclosure and the claims. Notably, in particular, any steps presented can be performed in any order, i.e. the present invention is not limited to a specific order of these steps. Moreover, it is also not required that the different steps are performed at a certain place or at one node of a distributed system, i.e. each of the steps may be performed at different nodes using different equipment / data processing.

[0138] As used herein ..determining" also includes ..initiating or causing to determine", “generating" also includes ..initiating and / or causing to generate" and “providing” also includes “initiating or causing to determine, generate, select, send and / or receive”. “Initiating or causing to perform an action” includes any processing signal that triggers a computing node or device to perform the respective action.

[0139] In the claims as well as in the description the word “comprising” does not exclude other elements or steps. The indefinite article “a” or “an” and the definite article “the” does not exclude a plurality. In particular, indefinite article “a” or “an” may be replaced with one or more and the definite article “the” may be replaced with the one or more. A single element or other unit may fulfill the functions of several entities or items recited in the claims. The mere fact that certain measures are recited in the mutual different dependent claims does not indicate that a combination of these measures cannot be used in an advantageous implementation.

[0140] Procedures like the receiving of a product request, the acquiring of the retrieval request, the retrieving of measurement data and the analysis thereof, etc., performed by one or several units or devices can be performed by any other number of units or devices. These procedures can be implemented as program code means of a computer program and / or as dedicated hardware. A computer program product may be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium, supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems. Any disclosure and embodiments described herein relate to the methods, the systems, devices, any computer program element lined out above and vice versa. Advantageously, the benefits provided by any of the embodiments and examples equally apply to all other embodiments and examples and vice versa. Any reference signs in the claims should not be construed as limiting the scope.

[0141] The invention relates to a method for determining a property of a chemical product, including a) receiving a product request fordetermining a target property value associated with a target property of a target product, and b) acquiring a retrieval request for retrieving, from a data management system, measurement data associated with the target product and suitable for determining based thereon the target property value. The retrieval request is acquired by providing the product request to a data-driven model trained to relate product requests to retrieval requests. Furthermore, the method includes c) retrieving measurement data associated with the target product from the data management system using the acquired retrieval request, and d) determining the target property value based on the retrieved measurement data. The invention allows to make more efficient use of measurement data associated with chemical products.

Claims

Claims:1 . A method (100) for determining one or more properties of one or more chemical products, the method (100) including:- receiving (101) a product request for determining one or more target property values associated with a target property of a target product,- acquiring (102) a retrieval request for retrieving, from a data management system (20), measurement data associated with the target product and suitable for determining based thereon the one or more target property values, the retrieval request being acquired by providing the received product request to a data-driven model (10) trained to relate product requests to retrieval requests,- retrieving (103) measurement data associated with the target product from the data management system (20) using the acquired retrieval request, and- determining (104a, 104b) the one or more target property values based on the retrieved measurement data.

2. The method (100) as defined in claim 1 , wherein the one or more target property values are determined based on the retrieved measurement data by:- acquiring (104a) analysis instructions for analyzing the retrieved measurement data with respect to the target property using an analysis engine (40), the analysis instructions being acquired by providing the retrieved measurement data to a data- driven model (30) trained to relate measurement data to analysis instructions, and- instructing (104b) the analysis engine (40) to analyze the retrieved measurement data with respect to the target property according to the acquired analysis instructions.

3. The method (100) as defined in claim 2, wherein a single data-driven model (10, 30) is used for acquiring the retrieval request and for acquiring the analysis instructions.

4. The method (100) as defined in any of claims 2 and 3, wherein the analysis engine (40) is a single-purpose analysis engine for carrying out a specific kind of analysis on a specific type of measurement data.

5. The method (100) as defined in any of claims 2 and 3, wherein the analysis engine (40) is a multi-purpose analysis engine for carrying out a selectable one of a plurality of possible types of analyses on a plurality of types of measurement data.

6. The method (100) as defined in any of the preceding claims, wherein the product request is indicative of a chemical reaction to which the target property relates.

7. The method (100) as defined in any of the preceding claims, wherein the data management system (20) includes a plurality of databases, wherein the model (10) trained to relate product requests to retrieval requests is trained such that it relates product requests to one or more retrieval requests for one or more selected databases of the data management system (20)8. The method (100) as defined in any of the preceding claims, wherein the data management system (20) includes a plurality of databases, wherein the retrieval request for retrieving the measurement data associated with the target product is acquired by providing the received product request to the data-driven model (10) trained to relate product requests to retrieval requests in an adapted form, wherein the adapted form of the received product request is indicative of the product request and data retrieval specifications for the plurality of databases.

9. The method (100) as defined in claim 2 or, as far as they depend on claim 2, any of claims 3 to 8, wherein the analysis engine (40) is one of a plurality of analysis typespecific analysis engines (40) for carrying out respective specific types of analyses, wherein the analysis instructions are indicative of a respective analysis engine (40) to be used for analyzing the retrieved measurement data.

10. The method (100) as defined in claim 9, wherein the analysis instructions are acquired by providing the retrieved measurement data to the data-driven model (30) trained to relate measurement data to analysis instructions in an adapted form, wherein the adapted form of the measurement data is indicative of the product request and analysis specifications for the plurality of analysis engines.

11. The method (100) as defined in claim 9 or 10, wherein analysis instructions are acquired for several of the analysis type-specific analysis engines (40), and wherein the one or more target property values are determined by analyzing the retrieved measurement data with the several analysis type-specific analysis engines (40).

12. The method (100) as defined in claim 11 , further including:- acquiring (105) a response to the product request by providing the analysis results of the several analysis engines (40) to a data-driven model (50) trained to relate analysis results from analysis engines (40) to responses to product requests.

13. The method (100) as defined in claim 12, wherein the response to the product request is acquired by providing the analysis results of the several analysis engines (40) to the data-driven model (50) trained to relate analysis results from analysis engines to responses to product requests in an adapted form, wherein the adapted form of the analysis results is indicative of the product request and output specifications for the response to the product request.

14. A data processing system (200) comprising a processor configured to carry out the steps of any of the methods as defined in claims 1 to 13.

15. A use of any of the methods as defined in the claims 1 to 13 and / or the data processing system as defined in claim 14 for determining one or more properties of one or more chemical products.

Citation Information

Patent Citations

  • Methods and systems that provide unified bills of material

    US20100030767A1

  • Distributed industrial performance monitoring and analytics platform

    US20190369607A1

  • Methods and apparatuses for characterizing chemical substances, measuring physicochemical properties and generating control data for synthesizing chemical substances

    WO2023198927A1