Sustainability in performing chemical reactions and measurements

A data-driven model optimizes chemical reactions and measurements by identifying matching historical data, reducing redundancy and enhancing sustainability by ensuring new useful results in chemical processes.

WO2026017911A1PCT designated stage Publication Date: 2026-01-22BASF SE
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/070870
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-19
Filing Date
2025-07-21
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Chemical reactions and measurements of physiochemical properties often result in redundant or unsuccessful outcomes, leading to waste of material, time, and energy, as they are planned without considering past experiments and may not deliver new useful results.

Method used

A method utilizing a data-driven model to analyze historical and target chemical instructions, determining matching or non-matching conditions to optimize chemical reactions and measurements, thereby increasing the fraction of new useful results.

Benefits of technology

Enhances the sustainability of chemical processes by reducing redundant efforts and improving the efficiency of chemical reactions and measurements through informed decision-making based on historical data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025070870_22012026_PF_FP_ABST
    Figure EP2025070870_22012026_PF_FP_ABST
Patent Text Reader

Abstract

A method for performing a target chemical reaction and / or a target measurement of a physiochemical property is presented. The method includes a) receiving (102) target chemical instructions associated with the target chemical reaction and / or the target measurement of the physiochemical property, b) retrieving (104) historical chemical instructions associated with one or more historical chemical reactions and / or historical measurements of physio-chemical properties that were performed in the past, c) providing (106), to a data-driven model, task instructions for determining an indication of whether the historical chemical instructions and the target chemical instructions match, and d) providing (108), based on an output generated by the data-driven model in response, an indication for performing the target chemical reaction and / or the target measurement associated with the target chemical instructions. The presented method allows to increase the fraction of chemical reactions and / or measurements of physiochemical properties being performed that deliver new useful results.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Sustainability in performing chemical reactions and measurements

[0002] FIELD OF THE INVENTION

[0003] The present disclosure relates to a method and a system for performing chemical reactions and measurements of physiochemical properties, to a training of a data-driven model to be used in the method and / or the system, and to a use of any of the foregoing, particularly of a result of the method or an output of the system or model. The disclosure also relates to sustainability in performing chemical reactions and measurements.

[0004] BACKGROUND OF THE INVENTION

[0005] In the chemical industry, chemical reactions and measurements of physiochemical properties are usually planned based on an analysis of chemical reactions and measurements of physiochemical properties that have been performed in the past. Nevertheless, it happens that chemical reactions and measurements of physiochemical properties are performed unsuccessfully, i.e., without delivering useful results. Likewise, it can happen that chemical reactions and measurements of physiochemical properties are performed even though they have been performed almost identically already in the past, thereby delivering redundant results. Both cases are a waste of material, time and energy. Furthermore, a plurality of aspects can be analyzed based on a measurement in relation to a chemical product and / or chemical reaction. Former experiments may focus on a part of aspects. Thus, measurement data may be reused in a different context than the previous one. Hence, in order to make the chemical industry economically and environmentally more sustainable, there remains a desire to increase the fraction of chemical reactions and / or measurements of physiochemical properties being performed that deliver new useful results.

[0006] SUMMARY OF THE INVENTION

[0007] It is an object addressed by the present disclosure to increase the fraction of chemical reactions and / or measurements of physiochemical properties being performed that deliver new useful results. In a first aspect, a method for performing a target chemical reaction and / or a target measurement of a physiochemical property is disclosed. The method includes a) receiving target chemical instructions associated with the target chemical reaction and / or the target measurement of the physiochemical property, and b) retrieving historical chemical instructions associated with one or more historical chemical reactions and / or historical measurements of physiochemical properties that were performed in the past. Furthermore, the method includes c) providing, to a data-driven model, task instructions for determining an indication of whether the historical chemical instructions and the target chemical instructions match, wherein the data-driven model is configured to follow task instructions and the provided task instructions include the historical chemical instructions and the target chemical instructions, and d) providing, based on an output generated by the data-driven model in response to being provided with the task instructions, an indication for performing the target chemical reaction and / or the target measurement associated with the target chemical instructions.

[0008] Since it is generated by the data-driven model in response to corresponding task instructions, the output generated by the data-driven model may be understood as an indication of whether the historical chemical instructions and the target chemical instructions match. An indication of whether target chemical instructions match historical experimental instructions can have a significant value for a sustainable design of chemical reactions and / or measurements of physiochemical properties. In particular, if it is known that a planned chemical reaction and / or a planned measurement of a physiochemical property was already performed in the past similarly as planned, it might not need to be performed anymore. Alternatively, the chemical reaction and / or planned measurement might be replanned, such as by adapting and / or completing the target chemical instructions. Both can increase the fraction of new useful chemical reactions and / or measurements of physiochemical properties being performed.

[0009] The term “chemical reaction” as used herein shall also include biological reactions, such as enzymatic reactions, for instance. The term “chemical” as used herein shall insofar also cover biological attributes.

[0010] Measurements of a physiochemical property are understood herein as including analyses of educts, intermediate products, transition states and products. Intermediate products may be understood as products obtained from educts reacting towards a product. A transition state may be understood as a non-stable state of a chemical compound occurring during a reaction from an educt.to a product. Accordingly, a physiochemical property as understood herein may be associated with a substance and / or a process. Moreover, physiochemical properties are understood herein as including physical properties and chemical properties, wherein, as indicated above, chemical properties may include biological properties. Physical properties are not understood herein as being limited to “purely” physical properties, and chemical properties are not understood herein as being limited to “purely” chemical properties. Accordingly, physiochemical properties include, but are not limited to, properties that are both physical and chemical.

[0011] For instance, a physiochemical property associated with a product may be a physical or chemical property associated with a chemical reaction for which the product is used or from which it arises. In particular, a physiochemical property associated with a product may be a physical or chemical property associated with a production (e.g., a synthesis) of the product. One possible physiochemical property would be, for instance, a conversion ratio associated with the product, such as a conversion ratio achievable in producing the product under various production conditions.

[0012] Insofar as the target chemical instructions are associated with a target chemical reaction, they may be associated with one or more educts intended for the target chemical reaction, with one or more products intended for the target chemical reaction, with reaction conditions intended for the target chemical reaction, with a solvent intended for the target chemical reaction, with a catalyst intended for the target chemical reaction, with laboratory equipment intended for performing the target chemical reaction, or with a combination of two or more of the foregoing.

[0013] Insofar as the target chemical instructions are associated with a target measurement of a physiochemical property, they may be associated with the property to be measured, with one or more chemicals intended to be used for performing the measurement, with one or more measurement devices intended to be used for performing the measurement, or with a combination of two or more of the foregoing.

[0014] More generally, the target chemical instructions may refer to instructions for performing the target chemical reaction and / or the target measurement, respectively. Hence, the target chemical instructions may be understood as a particular kind of chemical data. In particular, the target chemical instructions may specify educts and / or products, a quantity associated with educts and / or products, and / or reaction conditions. The target chemical instructions may also correspond to instructions for processing the educts towards the products. Correspondingly, the historical chemical instructions may generally refer to instructions according to which a respective historical chemical reaction and / or historical measurement was performed. Hence, also the historical chemical instructions may be understood as a particular kind of chemical data.

[0015] In particular, insofar as the historical chemical instructions are associated with a historical chemical reaction, they may be associated with one or more educts intended for or used in the historical chemical reaction (e.g., a type of quantity of a respective educt), with one or more products intended to be produced or actually produced in the historical chemical reaction (e.g., a type of quantity of a respective product), with instructions for processing the educts towards the products, with reaction conditions that were intended for or were actually present in the historical chemical reaction, with a solvent intended for or actually used in the historical chemical reaction, with a catalyst intended for or actually used in the historical chemical reaction, with laboratory equipment intended for or actually used for performing the historical chemical reaction, or with a combination of two or more of the foregoing.

[0016] Insofar as the historical chemical instructions are associated with a historical measurement of a physiochemical property, they may particularly be associated with the measured physiochemical property, with one or more chemicals used in the measurement, with one or more measurement devices used in the measurement, or with a combination of two or more of the foregoing.

[0017] Historical and target chemical instructions can be said to match if they at least partially correspond to each other. Historical chemical instructions at least partially corresponding to target chemical instructions could also be referred to as historical chemical reactions being relevant to the target chemical instructions. In that case, only historical chemical instructions that do not correspond even partially to the target chemical instructions could be considered irrelevant to the target chemical instructions. A matching between the target chemical instructions and the plurality of historical chemical reactions, as may be carried out using the data-driven model, could be based on a relevancy measure indicating a matching degree for any pair of a) target chemical instructions and b) historical chemical instructions.

[0018] For instance, the data-driven model may be trained and / or configured to provide, in response to receiving a pair of a) target chemical instructions and b) historical chemical instructions as input, as output a score indicative of a relevancy of the historical chemical instructions for the target chemical instructions. The input may furthermore indicate to the model the task, namely that a score indicative of a relevancy of the historical chemical instructions for the target chemical instructions is to be provided. The score may be such that it measures a degree to which and / or how many parameters in the historical and the target chemical instructions match. Based on the score and a threshold value for the score, it might then be determined whether the historical chemical instructions match the target chemical instructions or not. In this case, the data-driven model may be configured to provide a score, i.e. a relevancy score, per set of a plurality of retrieved sets of historical chemical instructions, wherein, for the determination regarding the matching based on the scores, the data-driven model may not be needed. For instance, for determining whether historical chemical instructions match target chemical instructions, a relevancy score determined by the data-driven model based on the two kinds of chemical instructions may be provided as input to a classifier for determining, based on the score, whether the chemical instructions match. In that case, the classifier may implement the threshold value for the score.

[0019] Additionally or alternatively, the one or more historical chemical reactions may be retrieved in sets, wherein per set a respective historical chemical reaction or a respective historical measurement of one or more physiochemical properties may be associated (i.e., with the respective set), and wherein the data-driven model may be configured to encode a) the target chemical instructions, which may also form one or more sets, and b) the one or more sets of historical chemical instructions received as input into a form by which the target chemical instructions can be compared to the one or more sets of historical chemical instructions to determine a respective similarity. Based on the encoding, the data-driven model may be configured to carry out a similarity search to find matching historical chemical instructions.

[0020] The output of the data-driven model may already comprise the matching historical chemical instructions or not. It is also possible that the data-driven model just provides the encoding of the matching chemical instructions as output. The encoding can be a vector encoding, for instance, in which case a similarity between the target chemical instructions and a set of historical chemical instructions may be determined based on respective vectors corresponding to the respective encoded versions of the target chemical instructions and the set of historical chemical instructions. For instance, a scalar product between the vectors may be determined and used as indicator for the similarity. Thus, the similarity search may include determining one or more products between encoded data, wherein the products may be products between vectors corresponding to the encoded data. Furthermore, the similarity search may include determining whether the one or more products lie within a predefined range. If so, the respective data pair for which the product has been determined may be considered as matching. For instance, the respective historical chemical instructions may be selected.

[0021] Optionally, the output of the data-driven model, the matching historical chemical instructions and / or the target chemical instructions may be provided to an apparatus such as a chemical production system for performing the target chemical reaction and / or a measurement system for performing the target chemical measurement based on the output, the matching historical chemical instructions and / or the target chemical instructions, respectively.

[0022] In an embodiment, the provided indication for performing the target chemical reaction and / or the target measurement associated with the target chemical instructions may include one or both of the following: a) a trigger for performing the target chemical reaction and / or the target measurement, wherein the trigger is provided upon determining a difference between the historical chemical instructions and the target chemical instructions, and b) historical chemical instructions, wherein the historical chemical instructions are provided upon determining a match between at least a part of the historical chemical instructions and the target chemical instructions. A combination of the two options would be present, for instance, if the trigger is a trigger for providing further target chemical instructions based on the historical chemical instructions and the target chemical instructions.

[0023] For instance, a summary of the matching historical chemical instructions and / or data that was acquired in the historical chemical reaction and / or the historical measurement associated with the matching historical chemical instructions can be provided. The summary may be displayed to a user.

[0024] Also, the method may include determining a difference between the target chemical instructions and matching historical chemical instructions. The difference can be determined, for instance, by providing the target chemical instructions and the matching historical chemical instructions as input to a data-driven model, wherein the data-driven model is instructed to determine the difference in response. Hence, the task to be followed by the data-driven model would insofar be to determine the difference. The same data-driven model as used for determining the matching historical chemical instructions might be used for determining the difference.

[0025] The indication for performing the target chemical reaction and / or the target measurement associated with the target chemical instructions may be provided based on the determined difference. For instance, the indication may be provided so as to depend on and / or include the difference.

[0026] In an embodiment, the data-driven model may be further configured to identify, in response to receiving the task instructions, whether the historical chemical instructions lack chemical data that match the target chemical instructions. In particular, the chemical data that lack in the historical chemical instructions but match the target chemical instructions may be identified. Knowing that the historical chemical instruction lack chemical data matching the target chemical instructions, and particularly knowing which or which kind of chemical data matching the target chemical instructions lack in the historical chemical instructions, can be a valuable information in planning a target chemical reaction and / or target measurement, since it may, for instance, avoid that the historical chemical instructions are considered irrelevant prematurely, or allow to discover needs for further measurements, for further calculations, simulations or other kinds of numerical derivations or for an access of further sources of chemical data.

[0027] A matching between chemical data and chemical instructions like the target chemical instructions may be defined in the same was as described above for the matching between historical chemical instructions and target chemical instructions. In other words, the options described herein for the matching between different sets of chemical instructions may be extended to the matching between other kinds of chemical data and chemical instructions. Thus, in particular, the data-driven model of the above embodiment may be configured to identify, in response to receiving the task instructions, whether the historical chemical instructions lack chemical data relevant to the target chemical instructions. The matching between the chemical data lacking in the historical chemical instructions and the target chemical instructions may thus refer to a relevancy of the lacking chemical data for the target chemical instructions.

[0028] A matching score may be defined for any pair of chemical data and target chemical instructions, wherein the matching score may indicate a degree of matching between the chemical data and the target chemical instructions, or a degree of relevancy of the chemical data to the target chemical instructions. For instance, chemical data may be considered as matching or as being relevant to target chemical instructions if a matching score determined for the chemical data and the target chemical instructions is equal to or above a predefined threshold, and / or if it is within a predefined match range. Conversely, chemical data may be considered as not matching or as being irrelevant to target chemical instructions if a matching score determined for the chemical data and the target chemical instructions is below a predefined threshold and / or within a predefined no-match range. It shall be understood that the lacking of chemical data, in the historical chemical instructions, that match or are relevant for the target chemical instructions does not necessarily mean a complete lack of such chemical data in the historical chemical instructions. In contrast, even historical chemical instructions which are determined to match the target chemical instructions to a relatively high degree might be determined to lack chemical data, i.e., further chemical data, that match, i.e., would match, the target chemical instructions. In particular, in historical chemical instructions partially matching the target chemical instructions, a lack of further chemical data partially matching the target chemical instructions may be identified by the data-driven model.

[0029] In an embodiment, the method may further include a step of retrieving the chemical data that lack in the historical chemical instructions but match or are relevant to the target chemical instructions. If these chemical data have previously already been identified by the data- driven model, they may be retrieved based on this previous identification, and preferably provided as part of the indication for performing the target chemical reaction and / or the target measurement associated with the target chemical instruction. However, it is also possible that previously chemical data are retrieved, which may then be regarded as chemical data that potentially lack in the historical chemical reactions but match or are relevant to the target chemical instructions, wherein for these chemical data it may be determined by the data-driven model whether they indeed lack in the historical chemical instructions and match or are relevant to the target chemical instructions. If so, they may be provided as part of the indication for performing the target chemical reaction and / or the target measurement associated with the target chemical instruction. The retrieval of the chemical data that lack or potentially lack in the historical chemical instructions may be a retrieval from any of a plurality of data sources, which may even include publicly available data sources.

[0030] In another embodiment, particularly where no chemical data is accessible that lack in the historical chemical instructions but would match or be relevant for the target chemical instructions, the method may include a step of providing an indication to perform a measurement to acquire chemical data that lack in the historical chemical instructions but match or are relevant to the target chemical instructions. The indication may be based on the identification, by the data-driven model, that the historical chemical instructions lack chemical data matching or being relevant to the target chemical instructions. The indication may identify what kind of chemical data matching or being relevant to the target chemical instructions lack in the historical chemical instructions, but may not include these chemical data. Additionally or alternatively to retrieving or measuring the chemical data lacking or potentially lacking in the historical chemical instructions but matching or being relevant to the target chemical instructions, these chemical data may also be calculated (e.g., by a quantum mechanical calculation), determined from a simulation (e.g., a Monte Carlo simulation), or otherwise derived numerically (e.g., from one or more formulas known from textbooks or other public sources). Thus, a step of the method presented herein may also include providing an indication to perform a numerical derivation of the chemical data that lack in the historical chemical instructions but match or are relevant to the target chemical instructions. The numerical derivation may include calculations or simulations based on known equations, numerical algorithms and / or analytic algorithms. The numerical derivation may be performed by a numerical derivation engine upon receiving the indication. The numerical derivation engine may include one or more known software applications.

[0031] In an embodiment, the method may further include checking an authorization of a user associated with the target chemical instructions, wherein the indication for performing the target chemical reaction and / or the target measurement associated with the target chemical instructions may be provided only if it has been verified in the check that the user is authorized.

[0032] In further embodiments, the target chemical instructions may be received via a user interface and / or the target chemical instructions may be received from a database. Hence, the target chemical instructions may be received via a user interface, from a database or both. In the latter case, for instance, the target chemical instructions may be determined based on a combination of a user input provided via the user interface and a retrieval of at least a part of the target chemical instructions from the database. The database may be a knowledge database with which a user might interact to plan a target chemical reaction and / or a target measurement of a physiochemical property. The knowledge database may store data associated with, for instance, chemicals, laboratory equipment and measurement devices.

[0033] In a particular example, a user may provide, as input via the user interface, an indication of the target chemical reaction associated with the target chemical instructions. Receiving the target chemical instructions may then correspond to retrieving the target chemical instructions (e.g., from a database), wherein the retrieving may include providing a request (e.g., to the database) for receiving the target chemical instructions associated with the target chemical reaction. The request may be obtained from the indication of the target chemical reaction provided by the user. Moreover, in an embodiment, the historical chemical instructions may be received from a database. The database from which the historical chemical instructions may be received may or may not be different from the database from which the target chemical instructions may be received.

[0034] In an embodiment, the target chemical instructions and / or the historical chemical instructions may comprise one or more of numerical data, string data and image data. More particularly, the target chemical instructions and / or each set of historical chemical instructions associated with a respective historical chemical reaction and / or a historical measurement may comprise at least two of numerical data, string data and image data. The data-driven model may in that case be configured to internally represent the respective instructions by mapping the at least two datatypes to a common representation. For instance, data of two different types may be mapped to a single second-rank tensor.

[0035] Moreover, the data-driven model may be a pre-trained model. Furthermore, the data driven model may be a fine-tuned model. Thus, for instance, the data driven model may be a pretrained model that is optionally also fine-tuned.

[0036] The pre-trained model may be parameterized and / or trained based on data with a plurality of contexts and / or unstructured data, in particular text data and optionally numerical data such as tabular data or image data. The pre-trained model may be configured to perform a plurality of tasks and / or to process data of a plurality of contexts. The pre-trained model may be configured to perform a task according to a provided task instruction. Hence, the pre-trained model may be configured to be provided with a plurality of different task instructions and / or provide a plurality of different types of output data upon receiving different task instructions.

[0037] A fine-tuned model may be obtained by training a pre-trained model configured to perform a plurality of tasks according to a plurality of task instructions. The fine-tuned model may be trained additionally on a training data set comprising a plurality of task instructions of one type and corresponding output data. The fine-tuned model may be trained additionally to provide output data of a predefined type according to a training data set. The fine-tuned model may be configured to be provided with a plurality of different task instructions and / or provide a plurality of different types of output data upon receiving different types of task instructions. Further, the fine-tuned model may be configured for providing one type of output data upon receiving one type of task instruction with a higher accuracy than providing other types of output data upon receiving other types of task instructions. In an embodiment, as already indicated above, the output provided by the data-driven model may be indicative of whether the received target chemical instructions and the retrieved historical chemical instructions correspond at least partially to each other. A partial correspondence may particularly be determined per data type of the chemical instructions, or across the data types.

[0038] For instance, the data-driven model may originally have been a pre-trained model, wherein the pre-trained model may then have been subject to a supplementary training in which it has been trained to provide, in response to receiving target chemical instructions and historical chemical instructions as input, an output indicative of whether the received target chemical instructions and historical chemical instructions at least partially correspond to each other. However, it is also possible that, for instance, this training is carried out as original training, i.e., not using a model that has been pre-trained already.

[0039] Also, in an embodiment, the target chemical instructions and the historical chemical instructions may be provided as input to the data-driven model together with model instructions specifying a type of processing of the target chemical instructions and the historical chemical instructions by the data-driven model. Model instructions are also being referred to as task instructions herein.

[0040] For instance, the data-driven model may be a large language model (LLM), in which case the model instructions may refer to text being input via a user interface and specifying matching criteria for the target chemical instructions and the historical chemical instructions. The matching criteria may indicate, for instance, which parts of the respective instructions should have a higher significance for the matching than others. Matching criteria may be provided via a user interface, e.g. based on human-expert input received. This allows for control of matching the historical and target instructions by the human-expert, i.e. the benchmark in the field of experimental planning and analysis.

[0041] In a particular embodiment, the data-driven model may be a fine-tuned data-driven model, wherein the fine-tuned data-driven model is trained based on a) one or more sets of task instructions for determining whether the historical chemical instructions and the target chemical instructions match and b) corresponding indications. The one or more sets of task instructions may include, per set, two sets of chemical instructions. The corresponding indications may indicate, per set of task instructions, whether the two sets of chemical instructions included therein match or not. The indications may be verified, such as by a human experts, for instance. The one or more sets of task instructions and the corresponding indications may be regarded as training data for the fine-tuning of the data-driven model, particularly as training input and output data, respectively.

[0042] In a further aspect, a system for performing a target chemical reaction and / or a target measurement of a physiochemical property is disclosed. The system comprises a) a target chemical instructions receiver configured to receive target chemical instructions associated with the target chemical reaction and / or the target measurement of the physiochemical property, and b) a historical chemical instructions retriever configured to retrieve historical chemical instructions associated with one or more historical chemical reactions and / or historical measurements of physiochemical properties that were performed in the past. Furthermore, the system comprises c) a data-driven model instructor configured to provide, to a data-driven model, task instructions for determining an indication of whetherthe historical chemical instructions and the target chemical instructions match, wherein the data-driven model is configured to follow task instructions and the provided task instructions include the historical chemical instructions and the target chemical instructions, and d) an indication generator configured to provide, based on an output generated by the data-driven model in response to being provided with the task instructions, an indication for performing the target chemical reaction and / or the target measurement associated with the target chemical instructions.

[0043] Another of the aspects disclosed relates to a use of a system as defined above for performing a target chemical reaction and / or a target measurement of a physiochemical property.

[0044] Similarly, in an aspect, a use of a data-driven model for determining whether historical chemical instructions and target chemical instructions match is disclosed, wherein the data- driven model is trained based on a) one or more sets of task instructions for determining whether the historical chemical instructions and the target chemical instructions match and b) corresponding indications.

[0045] A further aspect of the present disclosure relates to a use of the output of the data-driven model for performing a target chemical reaction and / or a target measurement of a physiochemical property. Where the output of the data-driven model corresponds to matching historical chemical instructions, for instance, the matching historical chemical instructions may be used for performing the target chemical reaction and / or the target measurement of a physiochemical property. The present disclosure also relates, in an aspect, to a method for further training a pretrained data-driven model, wherein the method includes a) providing the pre-trained data- driven model, b) providing training data including pairs of training input data and training output data, wherein the training input data, i.e. the training input data in a pair of training input data and training output data, correspond to two sets of chemical instructions associated with respective chemical reactions and / or target measurements of a physiochemical property, and the training output data, i.e. the training output data in a pair of training input data and training output data, correspond to verified indications of whether the two sets of chemical instructions in the respective training input data match, and c) training the pretrained data-driven model further using the training data such that the further trained data- driven model is trained and / or parameterized to provide, upon receiving two sets of chemical instructions as input, as output an indication of whether the two sets match. During the training, i.e. the further training, task instructions may be used for causing the model to generate the outputs in response to the training inputs. These task instructions may be the same or different from the task instructions provided to the trained data-driven model, i.e. the task-instructions used after deployment.

[0046] A related aspect disclosed herewith concerns a system for further training a pre-trained data-driven model, wherein the system comprises a providing unit for a) providing the pretrained data-driven model, and b) providing training data including pairs of training input data and training output data, wherein the training input data, i.e. the training input data in a pair of training input data and training output data, correspond to two sets of chemical instructions associated with respective chemical reactions and / or target measurements of a physiochemical property, and the training output data, i.e. the training output data in a pair of training input data and training output data, correspond to verified indications of whether the two sets of chemical instructions in the respective training input data match. Further, the system comprises a training unit configured for c) training the pre-trained data- driven model further using the training data such that the further trained data-driven model is trained and / or parameterized to provide, upon receiving two sets of chemical instructions as input, as output an indication of whether the two sets match. The training unit may use task instructions for causing the model to generate the outputs in response to the training inputs. These task instructions may be the same or different from the task instructions provided to the trained data-driven model, i.e. the task-instructions used after deployment.

[0047] The further training of a pre-trained data-driven model may also be understood as a fine- tuning of the pre-trained data-driven model. The above method may hence also be viewed as a method for fine-tuning a pre-trained data-driven model. In case it is intended to use the fine-tuned data-driven model with particular kinds of historical and target chemical instructions, wherein the kinds are such that the historical chemical instructions differ structurally from the target chemical instructions, the training data used in the fine-tuning can be adapted accordingly. In particular, among the training input data, it may be distinguished between historical chemical instructions and target chemical instructions, wherein the model being trained may receive one of each kind at a time also in the training. That is to say, the training input data in a pair of training input data and training output may include a set of historical chemical instructions and a set of target chemical instructions.

[0048] It shall be understood that the methods, the systems and the uses according to the aspects disclosed herein have similar and / or identical preferred embodiments, as defined in the dependent claims. In particular, it will be understood that the methods may be implemented - and hence, e.g., performed, or carried out - by the respective systems and that the systems may be or comprise respective computers or computer systems. Thus, it will also be understood that the methods may be computer-implemented. Similarly, it shall be understood that the uses disclosed herein may involve a computer or computer system and may therefore be considered computer-implementable.

[0049] It shall be understood that a preferred embodiment of the invention can also be any combination of the dependent claims with the respective independent claim.

[0050] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter.

[0051] BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In the following, the present disclosure is further described with reference to the enclosed figures. The same reference numbers in the drawings and this disclosure are intended to refer to the same or like elements, components, and / or parts.

[0053] FIG. 1 illustrates an embodiment of a system for performing a target chemical reaction and / or a target measurement of a physiochemical property, including a use thereof for chemical production,

[0054] FIG. 2A illustrates an embodiment of a method for performing a target chemical reaction and / or a target measurement of a physiochemical property. FIG. 2B illustrates an embodiment of a system for performing a target chemical reaction and / or a target measurement of a physiochemical property.

[0055] FIG. 2C illustrates a target chemical reaction in an embodiment.

[0056] FIG. 3 illustrates an embodiment of a training of an embedding layer.

[0057] FIG. 4A illustrates an embodiment of a transformer encoder architecture.

[0058] FIG. 4B illustrates an embodiment of a transformer decoder architecture.

[0059] FIG. 4C illustrates an embodiment of a transformer encoder-decoder architecture.

[0060] FIG. 5 illustrates an embodiment of training and / or deploying the transformer encoder, the transformer decoder and / or the transformer encoder-decoder.

[0061] FIG. 6 illustrates an embodiment of input embedding.

[0062] FIG. 7 illustrates a further embodiment of input embedding.

[0063] DETAILED DESCRIPTION OF EMBODIMENTS

[0064] The following embodiments are mere examples for implementing the methods and systems disclosed herein and shall not be considered limiting.

[0065] FIG. 1 shows schematically and exemplarily an operating system 10 for one or more chemical reaction and / or measurement facilities 20, in which chemical reactions and / or measurements of physiochemical properties can be performed. The one or more chemical reactions may involve one or more educts 30 reacting to one ore more products 40, wherein, for performing the one or more chemical reactions, corresponding chemical apparatuses may be used. The one or more chemical reactions may be monitored based on one or more measurements of physiochemical properties associated with the one or more educts 30, reaction conditions, and / or the one or more products 40. The one or more production facilities 20 may be operated, particularly regarding a control of the one or more chemical reactions and the one or more measurements, by the operating system 10. As illustrated, the operating system 10 may comprise one or more data sources 3. The data retrievable from the one or more data sources 3 may include a plurality of historical chemical instructions associated with historical chemical reactions and / or historical measurements of physiochemical properties that were performed in the past at the one or more reaction and / or measurement facilities 20.

[0066] As also illustrated, the operating system 10 may further comprise an intake interface 2. The intake interface may be a user interface (as in the case of FIG. 2B), or may be an interface for data exchange with another operating system, for instance. Via the intake interface 2, target chemical instructions may be received, wherein the target chemical instructions may be associated with targets chemical reactions and / or measurements of physiochemical properties that are planned to be performed at the one or more reaction and / or measurement facilities 20.

[0067] Furthermore, the operating system 10 may comprise a system in the form of a model engine 1 that is configured to receive and / or retrieve, from the intake interface 2 and the one or more data sources 3, respectively, a) target chemical instructions associated with one or more target chemical reactions and / or target measurements of physiochemical properties to be performed in the one or more chemical production facilities 20, and b) historical chemical instructions associated with one or more historical chemical reactions and / or historical measurements of physiochemical properties that were performed in the past. Based on these two types of chemical instructions, the model engine 1 may be configured to provide, to a data-driven model configured to follow task instructions, task instructions for determining an indication of whether the two types of chemical instructions match. The task instructions provided by the model engine 1 to the data-driven model may include the historical chemical instructions and the target chemical instructions. Based on an output generated by the data-driven model in response to being provided with the task instructions, the model engine 1 may provide an indication for performing the one or more target chemical reactions and / or target measurements associated with the target chemical instructions. A particular implementation of the model engine 1 will be described with reference to FIG. 2B.

[0068] In the example shown in FIG. 1 , the indication provided by the model engine 1 may, as illustrated, be forwarded to a control engine 4 of the operating system 10. The control engine 4 may be configured to control one or more apparatuses in the one or more reaction and / or measurement facilities 20. In particular, the control engine 4 may, based on the indication from the model engine 1 , cause one or more chemical reactions and / or measurements of physiochemical properties to be started in the one or more facilities 20. FIG. 2A shows schematically and exemplarily a method for performing a target chemical reaction and / or a target measurement of a physiochemical property. The method could also be referred to as a sustainable experimental design method. It may be carried out by processing units such as the model engine 1 .

[0069] In a first step 102 of the method, target chemical instructions associated with the target chemical reaction and / or the target measurement of the physiochemical property are being received. In an example, the target chemical instructions may be received via a user interface 201 of a system as shown schematically and exemplarily in FIG. 2B. In other examples, the target chemical instructions may be received differently, such as via another kind of intake interface or from a data source like a database as shown in FIG. 1 .

[0070] In a second step 104 of the method, historical chemical instructions associated with one or more historical chemical reactions and / or historical measurements of physiochemical properties that were performed in the past are being retrieved. The historical chemical instructions may, in an example, be retrieved from a database 203 of a system as shown schematically and exemplarily in FIG. 2B. In the example of FIG. 1 , the historical chemical instructions could be retrieved from one of the data sources 3.

[0071] The target chemical instructions and the historical chemical instructions may comprise one or more of numerical data, string data and image data. The target chemical instructions and the historical chemical instructions may be of a same or of a different data type. In an illustrative example, a user may provide an input via the user interface 201 that comprises a text and / or numerical input, such as generated via keyboard, and a further input, such as via selection of a corresponding data file by use of a computer mouse or a touchscreen or by use of a camera, that comprises image data. The historical chemical instructions may have been stored in the database 203 based on similar interactions of the user, and optionally a plurality of further users, with the interface 201 or similar interfaces of the system shown in FIG. 2B. The user inputs leading to the historical chemical instructions may be historical, i.e., past, inputs, provided in relation to respective historical chemical reactions and / or historical measurements of physiochemical properties, such as before, while and / or after they were performed.

[0072] Optionally, the historical chemical instructions may be retrieved by a first entity, whereas the target chemical instructions may be received from a second entity different from the first entity. The entities may refer, for instance, to different chemical companies. For instance, the first entity may be a data owner of the data corresponding to the historical chemical instructions. Access to the database 203 may therefore, in an example, be only granted to the first entity, and / or by the first entity to other entities. In this way, the first entity may remain in full control of whether and which historical chemical instructions are shared.

[0073] The step of receiving target chemical instructions may also be viewed as, or may go along with, a receipt of a query or request to provide an indication for performing the target chemical reaction as that of the target measurement. Upon receiving such a query or request (e.g., only then), the steps of retrieving the historical chemical instructions may be carried out. The query or request may be issued by the second entity mentioned above, and may be received by the first entity, such as via an interface like user interface 201 shown in FIG. 2B.

[0074] The chemical instructions, i.e. the historical and / or the target chemical instructions, may refer to one or more educts of a respective historical or target chemical reaction, one or more respective products, respective reaction conditions (e.g., a temperature, a pressure, a composition of surrounding educts and / or products such as a composition of air, in particular whether an inert gas, natural air or a higher oxygen composition is used), a solvent, a catalyst (e.g., a composition of a catalyst material such as whether it comprises a substrate, an active component, etc.), laboratory equipment such as a type of distillation column, or any combination of the foregoing. Educts and products may be characterized in terms of any of the following data types, for instance: SMILES, SMARTS, graph, image, tabular data listing type of atoms and coordinates and / or bonds between atoms, name of compound as string (name according to UIPAC nomenclature and / or trivial name).

[0075] Hence, the target and historical chemical instructions may be unstructured data. Moreover, they may be sequential. Data-driven models such as large language models may receive sequential data such as text token by token, wherein a token may represent at least a part of a word. The chemical instructions may comprise two or more elements associated with at least one of a respective product, one or more respective educts, reaction conditions such as the particular ones mentioned above, a solvent, a catalyst such as a catalyst composition as indicated above, laboratory equipment like a type of distillation column, or any combination of the foregoing.

[0076] In some cases, the historical chemical instructions may contain more information than the target chemical instructions. For instance, the historical chemical instructions may comprise a set of instructions suitable for synthesizing a historical product. A historical product may be a product having already been synthesized, in particular according to the historical chemical instructions. Historical chemical instructions may also comprise sensor data, hence data measured by sensors e.g. installed in an environment for performing the one or more historical chemical reactions.

[0077] In a third step 106 of the method illustrated in FIG. 2A, the target chemical instructions and the historical chemical instructions are being provided as input to a data-driven model for determining, as output of the data-driven model, an indication of whether the historical chemical instructions and the target chemical instructions match. The input provided to the data-driven model may be provided in terms of task instructions including the historical chemical instructions and the target chemical instructions. The task instructions may indicate to the data-driven model that its output is expected to indicate whether the historical chemical instructions and the target chemical instructions match.

[0078] The indication of whether the historical chemical instructions and the target chemical instructions match may refer to an indication of whether any of the historical chemical instructions match the target chemical instructions at least partially, and if so, which of the historical chemical instructions match the target chemical instructions. The indication of whether chemical instructions match each other can be based on a predefined matching rule, which may use a relevancy measure or score, for instance. The matching rule may be predefined based on a predetermined representation of the chemical instructions. For instance, as outlined further below respect to FIG. 3, the chemical instructions may be embedded as vectors in a vector space. For determining whether two sets of chemical instructions match, particularly whether a set of target chemical instructions match a set of historical chemical instructions, the corresponding vectors may be compared. The relevancy measure or score may be determined based on this comparison.

[0079] As indicated above, the target chemical instructions and the historical chemical instructions may be provided to the data-driven model together with model instructions indicating to the data-driven model that the model is supposed to generate an output indicative of whether the historical chemical instructions and the target chemical instructions match. For instance, the input provided to the data-driven model may be a prompt including the model instructions, wherein the prompt may furthermore include or refer to the target chemical instructions and the historical chemical instructions. The model instructions may be predefined, such as in terms of a prompt template. More generally, the model instructions may specify a type of processing of the target chemical instructions and the historical chemical instructions by the data-driven model. The type of processing may also be characterized, for instance, by numerical values for model parameters of the data-driven model. The data-driven model may be a pre-trained and / or fine-tuned data driven model, as outlined in more detail below with respect to FIG. 3 to FIG. 7. The model instructions may be predefined based on a type of the data-driven model.

[0080] The data-driven model may be configured / parametrized for receiving sequential data, in particular iteratively, and for generating sequential data, in particular iteratively, in response to receiving sequential data. The data-driven model may be trained with unstructured data, in particular unstructured data comprising one or more sequences of elements. An element may comprise one or more data points. The unstructured data may be independent of chemical instructions. The data-driven model may be parametrized and / or trained for mapping the target chemical instructions and the historical chemical instructions into a machine- processable representation of the target experimental instructions and the historical experimental instructions, such as into vectors. The data-driven model may process the machine- processable representation to an indication indicative of whether at least a part of the historical experimental instructions correspond to the target experimental instructions. This may be referred to as matching.

[0081] In a fourth step 108 of the method, an indication for performing the target chemical reaction and / or the target measurement is being provided based on the output of the data-driven model, i.e., based on the indication of whether the historical chemical instructions and the target chemical instructions match. The indication may be provided back to the user, such as in terms of a visual representation on a display of the user interface 201 . However, it may also be preferred that the indication is provided to an apparatus to be used in the target chemical reaction and / or the target measurement, such as in terms of an interface to the apparatus and / or a control engine 4 as illustrated in FIG. 1. Irrespective of how, or where to, the indication is provided, it may include one or both of a) a trigger for performing the target chemical reaction and / or the target measurement, wherein the trigger is provided upon determining a difference between the historical chemical instructions and the target chemical instructions, and b) historical chemical instructions, wherein the historical chemical structures are provided upon determining a match between at least a part of the historical chemical instructions and the target chemical instructions.

[0082] For instance, if the output of the data-driven model indicates that the target chemical instructions do not match any of the historical chemical instructions, this may be signaled to the user. The user may then decide that the target chemical reaction and / or the target measurement should be carried out, since it has not been carried out before. In a variant, if the output of the data-driven model indicates that there is a difference between the target chemical instructions and an identified set of historical chemical instructions, this may trigger a change in operation mode of an apparatus used for carrying out the target chemical reaction and / or the target measurement.

[0083] Also, if the data-driven model is configured to provide matching historical chemical instructions as output, i.e. historical chemical instructions for which a match with the target chemical instructions has been determined, then these matching historical chemical instructions may be displayed to the user, whereupon the user may decide, based on the matching historical chemical instructions, whether the target chemical reaction and / or the target measurement should be carried out or not. Additionally or alternatively, the matching historical chemical instructions may be provided as control input to an apparatus carrying out the target chemical reaction and / or the target measurement. As illustrated by FIG. 1 , such control input may be forwarded to the apparatus by a control engine, for instance. The control engine may control the apparatus. In particular, based on the output of the data- driven model, the control engine may cause the apparatus to start the target chemical reaction and / or the target measurement.

[0084] Thus, the data-driven model may not only be configured to determine whether two sets of chemical instructions match, but also to determine a difference between the two sets of chemical instructions. Based on a difference between a set of historical chemical instructions and a set of target chemical instructions, which may form part of the output of the data-driven model, in step 108 an indication may be determined for performing the target chemical reaction and / or the target measurement associated with the target chemical instructions. The indication may be directed to a user and / or an apparatus carrying out the target chemical reaction and / or the target measurement. Moreover, the indication may be an indication of a) whether or not and b) how the target chemical reaction and / or the target measurement should be carried out.

[0085] Optionally, at least a part of one or more sets of historical chemical instructions associated with at least a part of the one or more historical chemical reactions may be provided in step 108, in response to determining in step 106 that at least this part of the historical experimental instructions corresponds to the target experimental instructions. In particular, providing at least the part of the historical chemical instructions may comprise determining if a user associated with the target chemical instructions, such as a requesting second entity as indicated above, may be allowed to receive at least the part of the historical chemical instructions. Hence, receiving the target chemical instruction may include receiving encrypted target chemical instructions and decrypting the encrypted target chemical instructions, such as by using public and private keys. Hence, it may be determined if the user may be an authorized user. Furthermore, the historical chemical instructions may be associated with metadata indicative of whether at least a part of the historical chemical instructions may be provided to the user.

[0086] FIG. 2B shows schematically and exemplarily a system for performing a target chemical reaction and / or target measurement of a physiochemical property as it may be used in the method shown in FIG. 2A.

[0087] The system comprises a target chemical instructions receiver 202 that is configured to receive target chemical instructions associated with the target chemical reaction and / or the target measurement of the physiochemical property, and a historical chemical instructions retriever 204 that is configured to retrieve historical chemical instructions associated with one or more historical chemical reactions and / or historical measurements of physiochemical properties that were performed in the past. In the example shown, the target chemical instructions receiver 202 receives the target chemical instructions via a user interface 201 , whereas the historical chemical instructions retriever retrieves the historical chemical instructions from a database 203 of the system. However, as explained above, the sources of the chemical instructions may be different in other examples.

[0088] Furthermore, the system comprises a data-driven model instructor 206 configured to provide task instructions including the target chemical instructions and the historical chemical instructions as input to a data-driven model for determining, as output of the data-driven model, an indication of whether the historical chemical instructions and the target chemical instructions match. The data-driven model is not shown in FIG. 2B, as it may be hosted remotely and / or in a distributed manner and accessed by the data-driven model instructor 206 via known communication means. However, it is also not excluded that the data-driven model is hosted locally.

[0089] The system further comprises an indication generator 208 that is configured to provide, based on the output of the data-driven model, an indication for performing the target chemical reaction and / or the target measurement associated with the target chemical instructions. The providing of the indication may also be referred to as a generating of the indication. As explained in more detail above with reference to the method of FIG. 2A, the indication may be provided to a user via the user interface 201 , although other ways of providing the indication are equally possible, such as via an interface to an apparatus for performing the target chemical reaction and / or the target measurement. FIG. 2C shows a concrete example of a target chemical reaction. In this example, the target chemical reaction comprises 3-OI-propanoic acid reacting under basic conditions of pH = 11 and with ethanol as solvent to Ethoxypropanoic acid. Hence, in this case, the target chemical instructions could specify the educts to be 3-OI-propanoic acid and ethanol, and the product to be Ethoxypropanoic acid. Furthermore, the target chemical instructions may in this case specify the reaction conditions to be basic with a pH value equal to 1 1 . Partially matching historical chemical instructions may then, for instance, specify the same educt and the same product, and also still basic reaction conditions, but a pH value slightly different from 11 .

[0090] In an embodiment, the data-driven model may be further configured to identify, in response to receiving the task instructions, whether the historical chemical instructions lack chemical data matching or being relevant to the target chemical instructions. An indication thereof may be generated as part of an output of the data-driven model. In particular, the data- driven model may be configured such that, upon identifying that the historical chemical instructions lack further chemical data that would match or be relevant to the target chemical instructions, an indication of the lacking data is generated as part of the output of the data-driven model. That is to say, the output generated by the data-driven model may not only indicate that, in the historical chemical instructions, there is a lack of chemical data matching or being relevant to the target chemical instructions. Rather, the output generated by the data-driven model may be indicative of the lacking chemical data, i.e., it may comprise additional chemical data or at least information based on which additional chemical data can be retrieved. Accordingly, the data-driven model may be configured to reveal “hidden information” in the historical chemical instructions. This “hidden information” may be used as a basis for further target chemical instructions.

[0091] The data-driven model may be configured to identify the lack of chemical data in the historical chemical instructions, i.e., the lack of (e.g., further) chemical data matching or relevant to the target chemical instructions, based on a further set of chemical data. For instance, a plurality of sets of further chemical data may be retrieved and provided, with the historical chemical instructions, as input to the data-driven model. The data-driven model may be instructed to determine, from the plurality of sets of further chemical data, chemical data that match the historical chemical instructions. If the matching further chemical data comprise data not included in the historical chemical instructions, i.e., additional information, and if this additional data, or information, matches or is relevant to the target chemical instructions, this may be provided by the data-driven model as indication that the historical chemical instructions lack chemical data matching or being relevant to the received target chemical instructions. Whether chemical data match or are relevant to the target chemical instruction may again be determined based on a matching as explained above with respect to the historical chemical instructions and the target chemical instructions.

[0092] In an example, the received target chemical instructions may indicate that a chemical reaction is to be found by which a product P is produced from educts E1 and E2, and the provided historical chemical instructions may indicate reaction parameters for an attempted historical reaction in which the educts E1 and E2 did not react with each other. Just this may be indicated by the historical chemical instructions, and not, for instance, any further information about the educts E1 and E2 beyond the information necessary for identifying them, such as their name or chemical structure. These historical chemical instructions may be considered as matching the received target chemical instructions by the data-driven model because they relate to the same educts E1 and E2. When taken on their own, the matching historical chemical instructions may result in an indication that the target chemical reaction cannot be carried out and therefore other educts should be considered for producing the product P. However, if further chemical data are retrieved that indicate that, for instance, educt E1 is inert, i.e. does not react, below some temperature threshold, wherein the reaction parameters for the attempted historical reaction of E1 and E2 indicated a reaction temperature below this threshold, this may cause the data-driven model to identify a lack of chemical data in the historical chemical instructions that is relevant to the target chemical construction. Then, the data-driven model may in particular provide the historical chemical instructions as output together with the threshold temperature below which E1 is inert. A target chemical reaction may then be initiated with the educts E1 and E2 using the reaction parameters indicated by the historical chemical reaction, except that a reaction temperature above the threshold may be chosen.

[0093] In the example above, the further chemical data, which indicate a temperature threshold below which E1 is inert, are not only relevant to the target chemical instructions, but also match (i.e., partially match) the target chemical instructions, since they relate to one of the target educts, namely E1. If the target chemical instructions already include an indication of a target reaction temperature, the further chemical data, which by the temperature threshold also relate to a reaction temperature, would also match the target chemical instructions for this reason.

[0094] The further chemical data based on which the data-driven model may be configured to identify the lack of relevant chemical data in the historical chemical instructions may, as in the example above, refer to one or more properties of a substance referred to by the historical chemical instructions. If the historical chemical instructions refer to a device, such as a measurement device used for a historical measurement or a device used to carry out a historical chemical reaction, the further chemical data may also refer to this device. The one or more properties, such as the temperature threshold in the above example, may be retrieved from generally accessible data sources such as digital versions of textbooks, scientific articles, device manuals and / or the Internet, or from confidential digital documents such as internal historical reports. Thus, for instance, the one or more data sources 3 in FIG. 1 may include data sources of chemical data potentially lacking in the historical chemical instructions but matching or being relevant to the target chemical instructions. Accordingly, in FIG. 2B, a data source in addition to the database 203 may be indicated, wherein also a chemical data retriever apart from the historical chemical instructions retriever 204 may be added that would retrieve chemical data from the additional data source.

[0095] Thus, the data-driven model may not only be configured to follow task instructions for determining an indication of whether the historical chemical instructions match target chemical instructions, but also to follow task instructions, i.e. further task instructions, for identifying whether historical chemical instructions lack (e.g., further) chemical data matching or being relevant to target chemical instructions. To follow this further kind of task instructions, the data-driven model may be configured to determine whether historical chemical instructions and / or target chemical instructions match further chemical data. The further task instructions may include the further chemical data or a reference to one or more data sources of further chemical data. The further chemical data may be considered as chemical data that is potentially lacking in or matching / relevant to the historical chemical instructions, and potentially matching / relevant to the target chemical instructions.

[0096] Moreover, the method shown in FIG. 2A may include a further step (not shown) of retrieving chemical data associated with one or more entities referred to in the historical chemical instructions and the target chemical instructions. The one or more entities may include, for instance, chemical substances like educts and products, and / or devices like measurement devices and devices used to carry out chemical reactions. The retrieved chemical data may be provided, such as in step 106, to the data-driven model as part of task instructions for identifying whether the historical chemical instructions lack chemical data that match or are relevant to the target chemical instructions. If, then, a lack of chemical data matching or being relevant to the target chemical instructions is identified in the historical chemical instructions by the data-driven model, the indication for performing the target chemical reaction and / or the target measurement associated with the target chemical instructions, which is provided in step 108 based on the output generated by the data-driven model, may include the lacking chemical data matching or being relevant to the target chemical instructions, which may be extracted from one or more sets of the retrieved chemical data that have been determined by the data-driven model to match the historical chemical instructions and the target chemical instructions.

[0097] In a variant, if chemical data are identified that match the historical chemical instructions and the target chemical instructions, but it is further determined by the data-driven model that even with these chemical data the historical chemical instructions still lack chemical data relevant to the target chemical instructions, an indication to perform a measurement for acquiring the still lacking chemical data may be provided, such as in step 108. In a further variant, lacking or potentially lacking chemical data matching or relevant to the target chemical instructions may be derived numerically, such as by providing a corresponding indication to a dedicated engine that is configured to perform quantum mechanical calculations or Monte Carlo simulations, or to apply other known numerical or analytic schemes.

[0098] In the following, particular features of possible data-driven models and their training as considered herein will be described with reference to FIG. 3 to FIG. 7.

[0099] FIG. 3 illustrates an embodiment of obtaining an embedding layer usable in a data-driven model. The embedding layer may be obtained by training for example a continuous bag of words model (CBOW) or a skip-gram model. The embedding layer may be suitable for generating embedded input data based on input data. Generating embedded input data may refer to embedding input data.

[0100] As is the case throughout the subsequent description of FIG. 3 to FIG. 7, the input data may be unstructured or structured data. For instance, the input data may be or comprise general text, numerical and / or image data. However, more particularly, the input data throughout the subsequent description of FIG. 3 to FIG. 7 may also refer to target chemical instructions, historical chemical instructions and / or further chemical data. Such particular input data may comprise text, numerical and / or image data structured in a particular form.

[0101] Embedding input data may result in a representation associated with the input data. Thus, the embedded input 314 may be the representation associated with the input data. The input data may comprise one or more elements. The one or more elements may be represented by the input vector 306. In particular, the embedded input 314 and / or the input vector 306 may be machine-readable and / or processable by a processor. For this purpose, the embedded input 314 and / or the input vector 306 may be a tensor, in particular a first- rank tensor. Specifically, the input vector 306 may be a one-hot vector or a summation of a plurality of one-hot vectors. A one-hot vector may be a vector with one entry unequal to zero. Examples for one-hot vectors may be 308, 310 and 312. The entries unequal to zero in the one-hot vector and / or in the input vector 306 may indicate the element. For example, a lookup table may define the relation between the position of the entries unequal to zero and the element indicated by the one-hot vector. The lookup table may specify a plurality of different elements. The number of different elements may be equal to the number of entries in the one-hot vector. The number of different elements may be referred to as vocabulary size. In an example, the elements may be represented by tokens and a sequence of elements may refer to at least a part of a sentence. The at least a part of the sentence may be represented by a plurality of tokens. A token may represent at least a part of the element and / or word. For example, where one element would be associated with only one word, words such as “embeddings", “embedding” or “embed” would constitute different elements. A first token may represent the stem “embed” and the endings, typically appearing in a plurality of words, may be represented by a second token, a third token and a fourth token. The second token, the third token and the fourth token may be used for representing other words such as “look”, “looking” or the like, preferably together with a fifth token representing the stem “look”. Ultimately, this tokenization of elements associated with a plurality of stems and a plurality of endings results in less tokens to be used for representing a plurality of elements and thus, uses less computational resources.

[0102] A lookup table specifying a subset of the vocabulary size e.g. of the English language may comprise 10,000 words or more. The embedded input 314 may be a lower-dimensional representation than the input vector 306. For example, typical embedded inputs 314 may comprise some hundreds of different entries. Followingly, the embedded inputs 314 constitute a densified representation of one or more elements using less computational resources. More than that, the embedded input 314 may represent a relation between two or more elements. For example, the words “Italy” and “Germany” may be similar or may be more closely related since they both define European countries, whereas the word “embodiment” may be very different from the two respective words. The smaller the dot product between two embedded inputs 314 may be, the more similar the two elements associated with the embedded inputs 314 may be. Hence, the embedded inputs 314 may represent one or more elements accurately and lead to accurate results based on processing the embedded inputs 314.

[0103] For transforming the input vector 306 into the embedded input 314, the embedding layer may comprise a number of neurons equal to the number of entries in the embedded input 314. Based on the embedded inputs 314, the output layer may generate the output vector 316. The output vector may be a vector and / or may indicate one or more elements. The output vector 316 may indicate one or more elements different from the input vector 306 and / or the one-hot vectors associated with the input vector 306. For this purpose, the output layer may comprise a number of neurons equal to the number of entries of the input vector 306 and / or the output vector 316. The output layer may apply a softmax function to the embedded inputs 314. By doing so, the output vector may comprise the probabilities associated with the elements associated with the entries of the output vector 316 unequal to zero. Hence, from the output vector 316 one or more elements may be obtained with a corresponding probability. Where the input vector 306 may specify one or more sequence^) of elements, the output vector 316 may specify one or more elements corresponding to the sequence(s) of elements specified by the input vector 306. In the example of FIG. 3, the element associated with vector 318 may correspond to the input vector with a probability of 71 %. Additional or alternative elements may correspond to the input vector as indicated by the output vector with lower probability. By defining a threshold to which the probability may be compared, the selection of the corresponding elements may be tailored to the needs of the user. The elements generated by the model comprising the embedding layer 302 and the output layer 304 may refer to the most probable elements indicated by the output vector 316. Hence, the model depicted in FIG. 3 may generate the element associated with the vector 318 with a confidence score of 71 %.

[0104] The model of FIG. 3 may be a continuous bag of words (CBOW) model. The CBOW model may be trained based on a training data set comprising a plurality of input vectors and corresponding output vectors. As the training data set may not be labeled, the training of the CBOW model may be referred to as self-supervised. Before training of the CBOW model, the CBOW model may be initialized with random values assigned to the weights of the neurons. During the training of the CBOW model, the input vectors may be passed through the initialized embedding layer and the output layer and a loss may be determined by comparing the output vector obtained by passing the input vector 306 through the model to the output vector corresponding to the input vector 306 as specified by the training data set. Based on the determined loss, backpropagation may be applied to determine the gradients associated with the neurons of the embedding layer 302 and the output layer 304 to lower the loss. According to the determined gradients, the weights of the neurons may be updated by using a gradient descent algorithm. If a predetermined loss may be achieved by the CBOW model, the training may be terminated and a trained CBOW model may be obtained. From the trained CBOW model, the embedding layer 302 may be suitable for embedding input data comprising one or more elements. This embedding layer 302 may be used in other machine-learning architectures requiring an embedding layer 302 such as a transformer encoder, transformer decoder or transformer encoder decoder architecture as described within the context of FIG. 4A, FIG. 4B and FIG. 4C, all of which are possible architecture for the data-driven model considered herein. For training these architectures, a trained embedding layer 302 may be required. Hence, a model such as a CBOW model may be trained prior to training the transformer encoder, transformer decoder or transformer encoder decoder architecture.

[0105] FIG. 4A illustrates an embodiment of a transformer encoder architecture. The transformer encoder comprises an encoder input 478, one or more encoder blocks 474, 414 and an encoder output. The transformer encoder architecture may be derived from the transformer encoder-decoder architecture as known in the art and shown in FIG. 4C. In particular, the transformer encoder may be referred to as X-former. The transformer encoder architecture may correspond to the encoder architecture associated with the transformer encoder-decoder architecture with an additional encoder output instead of connecting the encoder block directly to the decoder of the transformer encoder-decoder architecture. A plurality of transformer encoder architectures are available in the art, such as the bi-directional encoder representations from transformers (BERT).

[0106] The input data may be received at the encoder input 478. The encoder input 478 may apply an input embedding 402. Applying the input embedding 402 may refer to passing the input data through an embedding layer, e.g. as described within the context of FIG. 3. Further, the encoder input 478 may apply positional encoding 404. Applying positional encoding 404 may refer to adding a positional factor to the embedded input obtained via input embedding. Preferably, the input data may specify a sequence of elements. The positional factor pposmay be indicative of the position of the elements within the sequence. For example, the positional factor pposmay be obtained based on the following equation: where pos may refer to the position of the element within the sequence, i may refer to the dimension associated with the input embedding and d may refer to the dimension of the model, e.g. transformer decoder, transformer encoder or transformer encoder-decoder. This may be referred to as absolute positional embeddings. Alternatively, the positional encoding may be based on rotary positional embeddings (RoPE). Positional encoding is beneficial since it enables the processing of sequential data without requiring further dimensions indicating the position of each element. Followingly, the positional encoding 404 reduces the computational resources needed for embedding the input data. By passing the input data through the encoder input, the input data may be transformed into a second- rank tensor representing the sequence of elements. This second-rank tensor may be referred to as embedded input data. The embedded input data may be processed by the encoder block. The embedded input data may be provided to the layer normalization 408 by a residual connection. Multi-head self-attention 406 may be applied to the embedded input data. Multi-head self-attention 406 may comprise the two components multi-head and self-attention. Self-attention may be understood as being a filter applied to the embedded input data. By applying the filter to the embedded input data, the elements associated with the embedded input data contributing to the to be generated output data may be identified for generating the output data. Hence, the filter may represent the degree of contributing to the to be generated output data by the elements associated with the embedded input data. Applying the filter may be referred to as weighting the elements associated with the embedded input data. This is advantageous specifically regarding long sequences of elements. The filter may be learned and improved during the training by learning to identify the contribution of elements associated with the embedded input data. For example, in the partial sentence “I went to the bakery to buy a” the last word may be generated by the data- driven model such as the transformer encoder. The self-attention may focus the transformer encoder to attend to the word “bakery” and “buy” mostly to generate the word “bread”. Self-attention may refer to attention generated based on the input data. Hence, the filter may be determined based on the input data, preferably the embedded input data. The embedded input data may serve as query Q, key K and value V with respect to the self-attention operation. The self-attention may refer to attention based on the received input data. Hence, the filter may be calculated based on the following formula by inserting the respective tensors based on the embedded input data: where dkcorresponds to the dimension of the key.

[0107] For improving the efficiency of the transformer encoder further, the multiple heads are used to apply the filter resulting in the multi-head self-attention 406. Multi-head self-attention 406 may comprise applying the filter to two or more parts of the embedded input data. Hence, the tensor may be split into two or more parts and the filter may be applied to the two or more parts separately by two or more heads according to the following equation: headt= Attention^QW*2, KW ,VW ) with parameter matrices where i may refer to the number of heads, dv, dKand dQmay refer to the dimensions of the value, key and query.

[0108] The result of the two or more heads may be concatenated according to the following equation:

[0109] MultiHead(Q, K, V) = Concat (head, , ... , headh)W(>where l¥0e IRftdyXd,anc| h may refer to the number of heads.

[0110] The embedded input data may be transformed via the multi-head self-attention 406 into a context tensor. The context tensor may represent the sequence of elements and the relation between two or more elements of the input data. The context tensor may be a second rank tensor and / or may comprise one or more first rank tensor(s). After the multi-head selfattention 406, layer normalization 408 may be applied based on the context tensor and / or the embedded input data from the residual connection. Applying layer normalization 408 may refer to normalizing the context tensor. Normalizing the context tensor may lower the values of the entries of the context tensor. This reduces the computational cost associated with processing the context tensor. Further, it improves the training by contributing the loss to converge and preventing instabilities.

[0111] Layer normalization 408 may be followed by passing the context tensor to a feed-forward layer 410, again followed by layer normalization 412 based on the residual connection to the context tensor and / or the output of the feed-forward layer 410. The feed-forward layer 410 may be a feed-forward neural network. The feed-forward neural network may comprise of a plurality of fully connected neurons. Passing the context tensor through the feed-forward neural network may result in transforming the context tensor linearly. Additionally or alternatively, the neural network may comprise one or more activation functions such as a rectified linear unit (ReLU). Hence, the neural network may be configured for performing one or more non-linear operations to the context tensor and / or transforming the context tensor non-linearly. After the context tensor has been transformed and / or normalized by the feed-forward layer 410 and the layer normalization 412, the context tensor may be provided to one or more further encoder blocks 414. Having passed the context tensor through the feed-forward layer 410 may adapt the context tensor for the processing by a further attention layer of the one or more further encoder blocks 414 for applying a selfattention filter, preferably multi-head self-attention 406. The context vector after being transformed by the layer normalization 412 and the feed-forward layer 410 may be referred to as hidden state.

[0112] The encoder output 476 comprises of a linear layer 416 and a softmax layer 418. The linear layer 416 may transform the context vector into a logits vector. The linear layer may be fully-connected. The logits vector obtained by passing the context tensor through the linear layer 416 may be passed through the softmax layer 418. Passing the logits vector through the softmax layer 418 may refer to applying the softmax function to the logits vector. Applying the softmax function to the logits vector may result in a probability distribution of one or more elements corresponding to the sequence of elements in the input data. From the probability distribution based on predefined selection criteria, one or more elements may be chosen. The one or more chosen elements may be referred to as the one or more elements generated by the transformer encoder. The one or more generated elements may be provided to the encoder input for generating further one or more elements corresponding to the sequence of the input data and the one or more elements generated by the transformer encoder as described within the context of FIG. 7.

[0113] FIG. 4B illustrates an embodiment of a transformer decoder architecture.

[0114] The transformer decoder comprises a decoder input 484, one or more decoder blocks 480, 432 and a decoder output 492. The transformer decoder architecture may be derived from the transformer encoder-decoder architecture as known in the art and shown in FIG. 4C. The transformer decoder may be referred to as X-former. The transformer decoder architecture may correspond to the decoder architecture associated with the transformer encoder-decoder architecture independent of receiving one or more hidden states from the encoder of the transformer encoder-decoder. A plurality of transformer decoder architectures are available in the art, such as the generative pretrained transformers (GPT).

[0115] The decoder input 484 may apply input embedding 420 and positional encoding 422 analogous to the input embedding 402 and the positional encoding 404 as described within the context of FIG. 4A.

[0116] The decoder block 480 may comprise the layer normalizations 426, the masked multi-head self-attention 424, the feed-forward layers 428 and / or the layer normalization 430. The embedded input data resulting from passing the input data through the decoder input 484 may be provided to the layer normalization 426 via a residual connection. Further, masked multihead self-attention 424 may be applied to the embedded input data. Masked multi-head self-attention 424 corresponds to the multi-head self-attention 406 as described within the context of FIG. 4A with additionally masking a part of the embedded input data associated with elements later in the sequence than the element to be generated. Additionally or alternatively, the part of the input data associated with elements later in the sequence than the element to be generated may not be received and / or transformed into the embedded input data. Thus, the transformer decoder may be suitable for generating a subsequent element to a sequence, whereas the transformer encoder may be suitable for generating a missing element in within one sequence and / or between two or more sequences. Therefore, the transformer encoder may be configured for classification tasks. The transformer decoder may be configured for text generation.

[0117] Similar to the transformer encoder as described within the context of FIG. 4A, a context tensor may be generated by applying the masked multi-head self-attention 424 and the layer normalization 426. The context tensor may be provided to the layer normalization 430 via a residual connection. Further, the feed-forward layer 428 and the layer normalization 430 may be analogous to the feed-forward layer 410 and the layer normalization 412 as described within the context of FIG. 4A. The context tensor may be provided to one or more further decoder blocks 432.

[0118] The decoder output 492 may comprise a linear layer 434 and a softmax layer 436. The linear layer 434 and the softmax layer 436 may be analogous to the linear layer 416 and the softmax layer 418 as described within the context of FIG. 4A.

[0119] FIG. 4C illustrates an embodiment of a transformer encoder-decoder architecture. The transformer encoder-decoder may comprise the encoder input 488, the one or more encoder blocks 486, 464, the decoder input 494, the decoder block 490 and the decoder output 492. The encoder input 488 may correspond to the encoder input 478 of FIG. 4A. The one or more encoder block(s) 486, 464 may correspond to the one or more encoder blocks 474, 414 of FIG. 4A. The decoder input 494 may correspond to the decoder input 484 of FIG. 4B.

[0120] The decoder block 490 may comprise a masked multi-head self-attention 470, a layer normalization 472, a feed-forward layer 438 and a layer normalization 440 analogous to the masked multi-head self-attention 424, the layer normalization 426, the feed-forward layer 428 and the layer normalization 430 as described within the context of FIG. 4B. The de- coder block 490 may further comprise a multi-head self-attention 450 and a layer normalization 448. Analogous to the description of FIG. 4B, the context tensor may be obtained from the masked multi-head self-attention 470 and the layer normalization 472. Multi-head self-attention 450 analogous to the multi-head self-attention 406 of FIG. 4A may be applied to the context vector obtained from the layer normalization 472 and the hidden states of the one or more encoder blocks 486, 464. Layer normalization 448 may be applied to the context vector obtained from the multi-head self-attention 450 and the context vector obtained from the layer normalization 472 provided via a residual connection. The context vector resulting from the layer normalization 448 may be processed via the feed-forward layer 438 and the layer normalization 440 analogous to the description of FIG. 4B. The context vector resulting from the layer normalization 440 may be provided to further decoder blocks 442 analogous to the decoder block 490. The context vector obtained from the one or more decoder blocks 490, 442 may be provided to the decoder output 492. The decoder output 492 may correspond to the decoder output 482 of FIG. 4B.

[0121] With the above-described architecture, the transformer encoder-decoder may receive and process input data at the encoder input 488 and the one or more encoder blocks 486, 464 and the decoder block 490 and the decoder output 492. Based on the input data, the transformer encoder-decoder may generate output data part by part or sequentially. The sequentially generated output data may be provided to and / or may be processed by the decoder input 494, the one or more decoder blocks 490, 442 and the decoder output 492. Preferably, a sequence may be provided to the encoder input 488 and after having generated at least a part of the output data, the decoder input 494 may be provided with at least the part of the elements of the output data already generated. By doing so, the next elements of the output data may be generated with a higher accuracy by taking the input data and the generated output data into account since more data may be received by the transformer encoder-decoder over time.

[0122] Because of the transformer encoder-decoder architecture, the transformer encoder-decoder may be configured for transforming a sequence into another representation of the sequence. An example for transforming one sequence into another representation may be translation of one sentence into another language. A plurality of transformer encoder-decoders are available in the art, such as BART, T5 or the like.

[0123] In an embodiment, the layer normalization 408, 412 may be applied prior to the masked multi-head self-attention 424, multi-head self-attention 406 and / or the feed-forward layer 410 in the transformer decoder, the transformer encoder and / or the transformer encoder- decoder. By doing so, the computational resources for applying the multi-head self-attention 406 and / or the feed-forward layer 410 to the embedded input data and / or the context tensor may be decreased as the entries of the respective tensors may be lower after normalization.

[0124] In an embodiment, the decoder output 492 may comprise a classification neural network, further feedforward layers, convolutional layers, fully connected layers or the like. For example, the transformer encoder-decoder may be configured for choosing between a plurality of options. For this purpose, the transformer encoder-decoder may be provided with three different input data sets and may classify the context vectors obtained from the one or more decoder blocks 490 via one or more linear layers. Followingly, the architecture may be extended depending on the use case to be solved.

[0125] FIG. 5 illustrates an embodiment of training and / or deploying the transformer encoder, the transformer decoder and / or the transformer encoder-decoder.

[0126] The encoder / decoder / encoder-decoder architecture 502 may correspond to the transformer decoder, the transformer encoder and / or the transformer encoder-decoder as described within the context of FIG. 4A - FIG. 4C.

[0127] The output data generated by the encoder / decoder / encoder-decoder architecture 502 may comprise one or more elements, in particular a sequence of elements. The previously generated elements of the output data may be provided as input for generating the next element in the sequence of the output data.

[0128] If, for instance, the input data did correspond to target chemical instructions and historical chemical instructions, then the output data may correspond to an indication of whether the historical chemical structures and the target chemical instructions match, i.e. , each other. Similarly, if, for instance, the input data did correspond to a) target chemical instructions and / historical chemical instructions and b) further chemical data, then the output data may correspond to an indication of whether the historical chemical structures and / or the target chemical instructions match the further chemical data, particularly to an indication of whether the further chemical data comprise parts lacking in the historical chemical instructions but matching or being relevant to the target chemical instructions. In the example of FIG. 5, the input data may comprise N elements, in particular input tokens. For instance, any target and / or historical chemical instructions and / or further chemical data, which may initially be input in a format comprising text and / or numerical data, may be tokenized, thereby converting it into a sequence of tokens. An input token may be a token dedicated to be inputted into a data-driven model such as the transformer decoder, the transformer encoder or the transformer encoder-decoder. The output data to be generated may comprise M elements. The encoder / decoder / encoder-decoder architecture 502 may generate one element of the output data based on receiving the input data and optionally previously generated elements of the output data at a timestep. Hence, for generating M elements M time steps are required. A time step comprises providing input 510, 512, 514 to the encoder / decoder / encoder-decoder architecture 502 and receiving output data 504, 508, 506 from the encoder / decoder / encoder-decoder architecture 502. In a first timestep, the input 510 may comprise of N input tokens. The N input tokens may be associated e.g. with N words, stems or endings. Preferably, the N input tokens may specify target chemical instructions and historical chemical instructions, and optionally further chemical data potentially lacking in the historical chemical instructions but matching or being relevant to the target chemical instructions. Optionally the input tokens may further specify a question relating to the target and historical chemical structures, such as whether the target and historical chemical instructions match or whether the further chemical data are indeed lacking in the historical chemical instructions but match or are relevant to the target chemical instructions. One or more input tokens may specify the beginning of the sequence of tokens and / or the end of the sequence of tokens. The input 510 may be processed by the encoder / decoder / encoder-decoder architecture 502. Based on the input 510 at least a part of the output data 504 may be generated. The at least a part of the output data may comprise a first output token. In the next timestep, the generated first output token may be provided together with the input 512. Specifically, where the input 512 may be received by a transformer encoder-decoder the input tokens may be received at the encoder input 488 and the first output token may be received at the decoder input 494. Where the input 512 may be received by the transformer encoder, the input 512 may be received by the encoder input 478 and analogously regarding the transformer decoder and the decoder input 484. Based on the input 512, the output data 508 comprising the first output token and a second output token may be generated. Generating the output data 508 based on the input 512 may refer to generating the second token based on the first token and the N input tokens, wherein the first token may have been generated based on the N input tokens. This process may be repeated until the last token in the sequence of the output data 506 may be generated. Preferably, the last token may be an end token. The end token may terminate the generation of a further output token. Similarly, to the data processing during deployment of the encoder / decoder / encoder-de- coder architecture 502, the encoder / decoder / encoder-decoder architecture 502 may be trained. The training data set may comprise a plurality of sequences comprising a plurality of elements. The sequences may be associated with the input data and / or the output data. Additionally or alternatively, the sequences may be independent of the input data and / or the output data. For example, where the input data and the output data may refer to chemical compositions represented via text, the training data set may comprise sequential text data independent of chemical compositions. In this example, the training data set may comprise sequences of words originating from a conversation. In an embodiment, the training data set may comprise at least partially input data sets and / or output data sets.

[0129] The training may be initialized by initializing the encoder / decoder / encoder-decoder architecture 502. In an embodiment, the parameters associated with the encoder / decoder / en- coder-decoder architecture 502 may be initialized randomly. Additionally or alternatively, the input embedding of the encoder / decoder / encoder-decoder architecture 502 may be obtained by training a CBOW model or a skip gram model as described within the context of FIG. 3. The trained embedding layer may be used during training. The parameters associated with the embedding layer may be kept constant and / or may be updated after a predefined number of training epochs. By doing so, the number of parameters to be updated is lower, enabling a faster and less computational resources-consuming training. Further, the accuracy associated with the embedding layer may be constant and / or may be increased by avoiding error compensation in relation to the just initialized encoder / de- coder / encoder-decoder architecture 502.

[0130] During the training of the encoder / decoder / encoder-decoder architecture 502, at least a part of the sequences of the training data set may be provided to the encoder / decoder / en- coder-decoder architecture 502 one by another and one or more elements may be generated based on the sequences of the training data set one by another. The elements generated based on the sequences may follow the elements of the parts of sequences the encoder / decoder / encoder-decoder architecture 502 may have been provided with. The generated one or more elements may be compared to the one or more elements following the at least a part of the sequences provided to the encoder / decoder / encoder-decoder architecture 502 as specified by the training data set. Hence, during the training the en- coder / decoder / encoder-decoder architecture 502 may generate a guess on the next element and the guess on the next element in a sequence may be compared to the ground truth specifying the actual next element according to the training data set. Based on the guess on the next element and the ground truth a loss may be determined. The loss may define the similarity between the guess on the next element and the ground truth. The loss may be determined by forming a vector dot product between the token associated with the one or more elements and the token associated with the ground truth. A loss unequal to zero may result in updating the parameters associated with encoder / decoder / encoder-de- coder architecture 502. Preferably the parameters associated with the encoder / de- coder / encoder-decoder architecture 502 may be independent of the embedding layer. For example, the parameters associated with the encoder / decoder / encoder-decoder architecture 502 may be weights of the neurons of the encoder / decoder / encoder-decoder architecture 502.

[0131] Based on the determined loss, backpropagation may be applied to determine the gradients associated with the parameters of the parameters associated with encoder / decoder / en- coder-decoder architecture 502 to lower the loss. According to the determined gradients, the parameters associated with the encoder / decoder / encoder-decoder architecture 502, preferably the weights of the neurons associated with the encoder / decoder / encoder-de- coder architecture 502, may be updated by using a gradient descent algorithm.

[0132] The training data set may be unlabeled. The sequences of elements within the training data set may inherently comprise the ground truth for determining the loss with respect to the one or more elements generated during the training of the encoder / decoder / encoder-de- coder architecture 502. Hence, the encoder / decoder / encoder-decoder architecture 502 may be trained self-supervised. This is advantageous since time and resources for creating a labeled training data set may be saved. Furthermore, this enables the usage of large training data sets associated with a size of several terabytes. Consequently, the data- driven model may be accurate in generating elements of a sequence. In addition, the large training data set enables few shot predictions or even zero shot predictions. Hence, the data-driven models trained as described above are versatile contributing to saving resources needed for training and / or hosting a plurality of purpose-driven models such as convolutional neural networks. The training described above may be referred to as pretraining. The data-driven model may be configured for performing few shot or even zero shot predictions with respect to a plurality of use cases after pretraining. The performance of the data-driven model may be increased further by additional training referred to as fine- tuning. The training data used for fine-tuning may comprise pairs of training input data and training output data. For fine-tuning the model to generate, upon being provided with target chemical instructions and historical chemical instructions as input, as output an indication of whether the target chemical instructions and the historical chemical instructions match, the training input data may comprise pairs of historical chemical instructions, and the training output data may comprise verified indications of whether the pairs are matching pairs. Similarly, for fine-tuning the model to generate, upon being provided with chemical instructions and further chemical data as input, as output an indication of whether the chemical instructions match the further chemical data, the training input data may comprise pairs of a) chemical instructions and b) further chemical data, and the training output data may comprise verified indications of whether the pairs are matching pairs.

[0133] FIG. 6 illustrates an embodiment of input embedding. Where the sequence of elements associated with the input data, preferably comprised in the input data, may be of one type, the input embedding 402, 420, 452, 466 as described within the context of FIG. 4A - FIG. 4C may be used. For example, a type of input data may be text where the elements may be associated with at least a part of a word, a punctuation character, a start token specifying the beginning of one or more sequences associated with the input data and / or the end token. In another example, the input data may be at least partially numerical. Hence, the input data may comprise a plurality of numbers. A set of target chemical instructions, historical chemical instructions or further chemical data, for instance, may comprise both text and numerical data. Numerical input data may be for example tabular data. Tabular data may specify one or more rows and / or one or more columns. Hence, the tabular data may comprise one or more cells, wherein the cells may be associated with one or more numerical values.

[0134] Numerical input data may require a different embedding than text input data. Input embeddings for numerical input data may comprise a token embedding, a positional embedding, a column embedding, a row embedding or a combination thereof.

[0135] Applying a token embedding to one or more elements, in particular tokens associated with the input data may result in a machine-processable representation associated with the one or more elements, in particular tokens. Applying the token embedding to one or more elements may refer to passing the one or more elements through the embedding layer, e.g. as described within the context of FIG. 3. Hence, token embeddings may specify the one or more elements, in particular tokens in a machine-processable representation. For example, the token embedding may transform a numerical value into a vector. This is advantageous since this representation can be enriched by further information such as the position of the token within the sequence and / or within a table associated with the sequence of tokens. The positional embedding may be analogous to the positional embedding as described within the context of FIG. 3, FIG. 4A - FIG. 4C. Where the input data may be tabular data, column embedding may be applied. Applying a column embedding to one or more elements, in particular tokens associated with the input data may result in a machine-processable representation specifying the location of the one or more elements within a table 602, preferably within the columns of the table 602. Applying the column embedding may refer to adding a column factor to the input data embedded via token embeddings, in particular the embedded input data. The column factor may be the same for elements associated with the same column and / or may differ between two or more elements associated with different columns. Analogous, row embeddings may be applied where the input data may be tabular data. Applying a row embedding to one or more elements, in particular tokens associated with the input data may result in a machine-processable representation specifying the location of the one or more elements within a table 602, preferably within the rows of the table 602. Applying the row embedding may refer to adding a column factor to the input data embedded via token embeddings, in particular the embedded input data. The row factor may be the same for elements associated with the same row and / or may differ between two or more elements associated with different rows.

[0136] In an embodiment, input data may be at least partially numerical and at least partially text. As indicated above, this may be the case for historical as well as target chemical instructions and further chemical data. Hence, the input data may comprise two or more types of data. A type of data may refer to a modality. Followingly, different embeddings may be applied to the input data. To parts of the input data comprising text the input embedding referred to in FIG. 3, FIG. 4A - FIG. 4C may be applied. To parts of the input data being numerical token embeddings, positional embeddings, column embeddings and row embeddings may be applied. Further, segment embeddings may be applied to the input data independent of the type of input data. The segment embedding may specify the type of input data one or more elements may be associated to. For example, if the input data comprises of text and numbers, the input data may comprise of two types of input data. Applying the segment embedding to the input data may refer to adding a segment factor to the input data, preferably the embedded input data and / or the input data after having applied the token embedding. The segment factor may specify the type of data associated with the one or more elements. The segment factor may be the same for one or more elements associated with the same type of input data and / or may differ between two or more elements associated with different types of input data.

[0137] Applying the token embedding, the positional embedding, the segment embedding, the column embedding, the row embedding or a combination thereof may result in embedded input data and / or may be the output of any one of the encoder input 478, 484, 488 or decoder input 484, 494. The data obtained by applying the token embedding, the positional embedding, the segment embedding, the column embedding, the row embedding or a combination thereof may be processed by the encoder block 474, 486, decoder block 480, 490, encoder output 476, decoder output 492, 482. FIG. 7 illustrates a further embodiment of input embedding.

[0138] Input data to the data-driven model, in particular to the encoder input and / or the decoder input as described in the context of FIG. 4A - FIG. 4C, may comprise image data. Also chemical instructions and further chemical data may comprise image data, such as in the form of an image, particularly a graph, associated with a chemical reaction and / or measurement to which they refer. The data-driven model may be parametrized to receive image data. For processing image data as input data, the data-driven model may comprise one or more encoder blocks and / or one or more decoder blocks and / or one or more encoder outputs and / or one or more decoder outputs as described within the context of FIG. 4A - FIG. 4C. FIG. 7 may show an embodiment of an encoder input and / or a decoder input. When processing image data, the encoder input and / or the decoder input of the data-driven model may be as described within the context of FIG. 7. The encoder input and / or decoder input may comprise one or more linear projection layers 714 for a linear projection of one or more images, preferably one or more partial images, more preferably a sequence of two or more partial images. The one or more linear projection layers 714 may be suitable for changing the dimension of the one or more received images, preferably one or more partial images, preferably passing the one or more images, preferably partial images, through the one or more linear projection layers 714 may result in applying image embedding, preferably partial image embedding to the one or more images and / or partial images.

[0139] Furthermore, when a sequence of two or more images and / or partial images may be received, positional embedding may be applied to the sequence, preferably by passing the sequence of one or more images and / or partial images through the one or more linear projection layers 714. Applying positional embedding may refer to adding a positional factor. The positional factor may be different depending on the position of the image and / or the partial image within the sequence. In particular, the positional factor added to a first element of the sequence may be different to the positional factor added to a second element of the sequence. The first element of the sequence may be a first image and / or first partial image. The second element of the sequence may be a second image and / or a second partial image.

[0140] The representation of the one or more images, preferably one or more partial images, may be obtained based on the following equation: where xciassis the image class embedding 728 , xNpis the n-th image, in particular partial image in the sequence, z0is the representation of the one or more images, preferably one or more partial images, (H,W) are the resolution of the image, in particular the image the partial images are generated on, C is the number of channels associated with the one or more image, in particular the one or more partial images and D is the dimension of the representation of the one or more images, preferably one or more partial images. Applying the partial image embedding may refer to forming the product of xNpwith E above-described equation. Applying the positional embedding may referto adding the factor Eposaccording to the above-described equation.

[0141] By doing so, text-based data, numerical data, tabular data, image data or the like may be processed by one data-driven model.

[0142] The present disclosure has been described in conjunction with preferred embodiments and examples as well. However, other variations can be understood and effected by those persons skilled in the art and practicing the claimed invention, from the studies of the drawings, this disclosure and the claims. Notably, in particular, any steps presented can be performed in any order, i.e. the present invention is not limited to a specific order of these steps. Moreover, it is also not required that the different steps are performed at a certain place or at one node of a distributed system, i.e. each of the steps may be performed at different nodes using different equipment / data processing.

[0143] As used herein ..determining" also includes ..initiating or causing to determine", “generating" also includes ..initiating and / or causing to generate" and “providing” also includes “initiating or causing to determine, generate, select, send and / or receive”. “Initiating or causing to perform an action” includes any processing signal that triggers a computing node or device to perform the respective action.

[0144] In the claims as well as in the description the word “comprising” does not exclude other elements or steps. The indefinite article “a” or “an” and the definite article “the” does not exclude a plurality. In particular, indefinite article “a” or “an” may be replaced with one or more and the definite article “the” may be replaced with the one or more. A single element or other unit may fulfill the functions of several entities or items recited in the claims. The mere fact that certain measures are recited in the mutual different dependent claims does not indicate that a combination of these measures cannot be used in an advantageous implementation. Procedures like the receiving of target chemical instructions, the retrieving of historical chemical instructions, the retrieving of further chemical data, the providing of the chemical instructions and optionally further chemical data as input to the data-driven model, the providing of the indication for performing the target chemical reaction and / or the target measurement, etc., performed by one or several units or devices, can be performed by any other number of units or devices. These procedures can be implemented as program code means of a computer program and / or as dedicated hardware. In other words, the methods disclosed herein can be computer-implemented. Also the uses referred to herein may be at least partially computer-implemented. A computer program product may be stored / dis- tributed on a suitable medium, such as an optical storage medium or a solid-state medium, supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems.

[0145] Any disclosure and embodiments described herein relate to the methods, the systems, devices, any computer program element lined out above and vice versa. Advantageously, the benefits provided by any of the embodiments and examples equally apply to all other embodiments and examples and vice versa.

[0146] Any reference signs in the claims should not be construed as limiting the scope.

[0147] A method for performing a target chemical reaction and / or a target measurement of a physiochemical property is presented. The method includes a) receiving target chemical instructions associated with the target chemical reaction and / or the target measurement of the physiochemical property, b) retrieving historical chemical instructions associated with one or more historical chemical reactions and / or historical measurements of physiochemical properties that were performed in the past, c) providing, to a data-driven model, task instructions for determining an indication of whether the historical chemical instructions and the target chemical instructions match, and d) providing, based on an output generated by the data-driven model in response, an indication for performing the target chemical reaction and / or the target measurement associated with the target chemical instructions. The presented method allows to increase the fraction of chemical reactions and / or measurements of physiochemical properties being performed that deliver new useful results.

Claims

CLAIMS1 . A method for performing a target chemical reaction and / or a target measurement of a physiochemical property, wherein the method includes:Receiving (102) target chemical instructions associated with the target chemical reaction and / or the target measurement of the physiochemical property,Retrieving (104) historical chemical instructions associated with one or more historical chemical reactions and / or historical measurements of physiochemical properties that were performed in the past,Providing (106), to a data-driven model, task instructions for determining an indication of whether the historical chemical instructions and the target chemical instructions match, wherein the data-driven model is configured to follow task instructions and the provided task instructions include the historical chemical instructions and the target chemical instructions, andProviding (108), based on an output generated by the data-driven model in response to being provided with the task instructions, an indication for performing the target chemical reaction and / or the target measurement associated with the target chemical instructions.

2. The method as defined in claim 1 , wherein the provided indication for performing the target chemical reaction and / or the target measurement associated with the target chemical instructions includes one or both of the following: a trigger for performing the target chemical reaction and / or the target measurement, wherein the trigger is provided upon determining a difference between the historical chemical instructions and the target chemical instructions, and historical chemical instructions, wherein the historical chemical instructions are provided upon determining a match between at least a part of the historical chemical instructions and the target chemical instructions.

3. The method as defined in claim 1 or 2, wherein the data-driven model is further configured to identify, in response to receiving the task instructions, whether the historicalchemical instructions lack chemical data that are relevant to the target chemical instructions.

4. The method as defined in claim 3, further including a step of a) retrieving the chemical data that lack in the historical chemical instructions but are relevant to the target chemical instructions, b) providing an indication to perform a measurement to acquire the chemical data that lack in the historical chemical instructions but are relevant to the target chemical instructions, and / or c) providing an indication to perform a numerical derivation of the chemical data that lack in the historical chemical instructions but are relevant to the target chemical instructions.

5. The method as defined in any preceding claim, further including checking an authorization of a user associated with the target chemical instructions, wherein the indication for performing the target chemical reaction and / or the target measurement associated with the target chemical instructions is provided only if it has been verified in the check that the user is authorized.

6. The method as defined in any preceding claim, wherein the target chemical instructions are received via a user interface (201).

7. The method as defined in any preceding claim, wherein the target chemical instructions are received from a database.

8. The method as defined in any preceding claims, wherein the historical chemical instructions are received from a database (203).

9. The method as defined in any preceding claim, wherein the target chemical instructions and / or the historical chemical instructions comprise one or more of numerical data, string data and image data.

10. The method as defined in any preceding claim, wherein the data-driven model is a pre-trained model.

11. The method as defined in any preceding claim, wherein the data-driven model is a fine-tuned data-driven model, and wherein the fine-tuned data-driven model is trained based on a) one or more sets of task instructions for determining whether the historicalchemical instructions and the target chemical instructions match and b) corresponding indications.

12. The method as defined in any preceding claim, wherein the output provided by the data-driven model is indicative of whether the received target chemical instructions and the retrieved historical chemical instructions correspond at least partially to each other.

13. The method as defined in any preceding claim, wherein the target chemical instructions and the historical chemical instructions are provided as input to the data-driven model together with model instructions specifying a type of processing of the target chemical instructions and the historical chemical instructions by the data-driven model.

14. A system for performing a target chemical reaction and / or a target measurement of a physiochemical property, wherein the system comprises: a target chemical instructions receiver (202) configured to receive target chemical instructions associated with the target chemical reaction and / or the target measurement of the physiochemical property, a historical chemical instructions retriever (204) configured to retrieve historical chemical instructions associated with one or more historical chemical reactions and / or historical measurements of physiochemical properties that were performed in the past, a data-driven model instructor (206) configured to provide, to a data-driven model, task instructions for determining an indication of whether the historical chemical instructions and the target chemical instructions match, wherein the data-driven model is configured to follow task instructions and the provided task instructions include the historical chemical instructions and the target chemical instructions, and an indication generator (208) configured to provide, based on an output generated by the data-driven model in response to being provided with the task instructions, an indication for performing the target chemical reaction and / or the target measurement associated with the target chemical instructions.

15. Use of a data-driven model for determining whether historical chemical instructions and target chemical instructions match, wherein the data-driven model is trained based ona) one or more sets of task instructions for determining whether the historical chemical instructions and the target chemical instructions match and b) corresponding indications.

Citation Information

Patent Citations

  • Systems and methods for efficient workflow similarity detection

    US20140129285A1

  • Machine learning system with two encoder towers for semantic matching

    US20230420085A1

  • Formulation generation

    US20240047014A1