Method for anomaly mitigation in a production cycle of a distributed chemical production

EP4751148A1Pending Publication Date: 2026-06-03BASF SE

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
BASF SE
Filing Date
2024-07-22
Publication Date
2026-06-03

AI Technical Summary

Technical Problem

In distributed chemical production environments, anomalies often go undetected until the end of a production cycle, leading to inferior product quality and waste.

Method used

A method utilizing a generative data-driven model, trained on reference and current production data, to detect anomalies in real-time and provide mitigation instructions, thereby preventing degradation of final product quality.

Benefits of technology

Enables early detection and mitigation of anomalies, reducing waste and saving resources by maintaining product quality throughout the production cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024070729_30012025_PF_FP_ABST
    Figure EP2024070729_30012025_PF_FP_ABST
Patent Text Reader

Abstract

The disclosure may relate to climate change mitigation through advanced manufacturing in an industrial environment of chemical production. The disclosure further relates to early detection of an anomaly in a production cycle and removing the source of the anomaly in a production cycle for restoring the quality of a product in a production cycle. Said detection and removal of the anomaly may be provided by using one or more pre-trained transformer-based models trained using one or more training data sets associated with one or more production operations of distributed chemical production. Said removal of the anomaly source in a production cycle may allow for reducing waste of a product and production resources associate with producing the product.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD FOR ANOMALY MITIGATION IN A PRODUCTION CYCLE OF A DISTRIBUTED CHEMICAL PRODUCTIONTECHNICAL FIELDThis disclosure relates a method for anomaly mitigation in a production cycle of a distributed chemical production.BACKGROUND ARTA distributed production environment of e.g. a chemical plant is highly complex, where various chemical processes take place controlled by a variety of chemical apparatuses. These processes may involve the handling, storage, and transformation of different chemicals (which may include hazardous, flammable and / or toxic chemicals). Completing a production cycle of chemical products in such an environment may take days to weeks, e.g. performed in a batch-wise manner. When an anomalous behavior occurs in a production plant (such as, for example, one or more plants discussed in W02020165045 (A1), WO2021 116123 (A1) and WO2021156157 (A1 )), said anomalous behavior may be detected at the end of the production cycle when the final product quality is measured.SummaryAccording to a first aspect a method for anomaly mitigation in a production cycle of a distributed production environment is disclosed, comprising:Receiving an analyze instruction related to detecting one or more anomalies in plant based-data, wherein the analyze instruction comprises a reference indicator indicating at least one reference data set and an analysis indicator indicating at least one data set to be analyzed;Obtaining, based on the analyze instruction, the at least one reference data set and the at least one data set to be analyzed;Determining at least one anomaly mitigation instruction, the determining the at least one anomaly detection result comprising: providing a task instruction, based on the analyze instruction, to at least one generative data-driven model, the at least one generative data-driven model having been trained to generate the at least one anomaly mitigation instruction based on the at least one reference data set and the at least one data set to be analyzed in response to receiving the task instruction;Providing the at least one anomaly mitigation instruction.According to further aspects, respective apparatus, system, and use are disclosed.EMBODIMENTSA distributed production environment of e.g. a chemical plant is highly complex, where various chemical processes take place controlled by a variety of chemical apparatuses. These processes may involve the handling, storage,and transformation of different chemicals (which may include hazardous, flammable and / or toxic chemicals). Production of chemical products in such an environment may take days to weeks, e.g. in a batch-wise manner. The aspects, embodiments, and examples provided in this disclosure may allow for early discovery of anomalies in the production environment, that may e.g. lead to inferior product quality. The aspects, embodiments, and examples provided in this disclosure may allow for identifying anomalous behavior and its sources in a production cycle in (near) real-time (e.g., within hours, before the end of the affected production cycle). This may allow mitigating degradation of final product quality in an affected production cycle and in next consecutive cycles. Hence, waste production, e.g., production of products with unacceptable quality, may be reduced. In addition to reducing product waste, the resources spent on producing said product may be saved.The aspects, embodiments, and examples provided in this disclosure may allow for improving the efficiency of chemical production in view of reducing waste caused by unacceptable quality of produced products for mitigating climate change.According to a first aspect a (in particular computer-implemented) method for anomaly mitigation in a production cycle of a distributed chemical production is disclosed comprising:Receiving (e.g. from a plant operator) an analyze instruction related to detecting one or more anomalies in plant based-data, wherein the analyze instruction comprises a reference indicator indicating at least one reference data set and an analysis indicator indicating at least one data set to be analyzed;Obtaining, based on the analyze instruction, the at least one reference data set and the at least one data set to be analyzed (e.g. from a database);Determining at least one anomaly mitigation instruction (e.g. for removing one or more sources of the one or more anomalies), the determining the at least one anomaly detection result comprising: providing a task instruction, based on the analyze instruction, to at least one generative data-driven model (e.g. a transformer based model), the at least one generative data-driven model having been trained to generate the at least one anomaly mitigation instruction based on the at least one reference data set and the at least one data set to be analyzed in response to receiving the task instruction;Providing the at least one anomaly mitigation instruction (e.g. to a plant operator for mitigating any identified anomaly in a (ongoing) production process).A reference indicator and / or an analysis indicator may allow to identify a specific data set e.g. a reference data set or data set to be analyzed comprising production data of a specific production line over a specific time range. For instance, it may be a given data range (e.g. from date 1 to date 2) and be associated with a designation of a plant, production line or apparatus. A reference indicator and / or an analysis indicator may also comprise the respective data set, so that the respective data set may not needed to be retrieved and may be obtained directly from the analyze instruction. A reference data set may be a data set comprising data of historical or past operation of a plantor production line and may be used as a training data set for fine-tuning a generative data driven model. Reference data sets may be stored in a database and may be (automatically) retrieved using a (database) query identifying the data set. A data set to be analyzed may comprise data of a current or ongoing production cycle, e.g. measured data at the end of a certain production line of the production cycle. A data set to be analyzed may also be stored on a database or another computer storage such as a memory. It may be data measured and (automatically) stored during the production cycle.Obtaining (e.g. from a database) the at least one reference data set and the at least one data set to be analyzed may comprise generating, e.g. by the at least one generative data-driven model or another generative data-driven model, a retrieval query for a database, and retrieving the data set based on the retrieval query from the database.A task instruction may be determined based on the analyze instruction, e.g. by including at least part of the analyze instruction along with e.g. further data, e.g. a context, or the task instruction may be generated by the at least one generative data-driven model or another generative data-driven model.According to an example embodiment of any aspect, the task instructing comprises at least a part of the at least one reference data set and / or at least a part of the at least one data set to be analyzed.For instance, the task instruction may comprise at least a part of the analyze instruction, at least a part of the at least one reference data set and at least a part of the at least one data set to be analyzed (e.g. as context). Including the part of the at least one reference data set in the task instruction may serve as one-shot training of the generative data-driven model.According to an example embodiment of any aspect, the generative data-driven model is fine-tuned or re-trained using the at least one reference data set (e.g. on the fly) and the fine-tuned or re-trained generative data-driven model is used as the generative data-driven model in the determining the at least one anomaly mitigation instruction.When using a fine-tuned generative data-driven model a task instruction for the fine-tuned generative model may comprise at least a part of the data set to be analyzed, but e.g. not the at least one reference data set.According to an example embodiment of any aspect, obtaining the at least one reference data set and / or the at least one data set to be analyzed comprises determining whether the at least one reference data set and / or the at least one data set to be analyzed are available in a data base;Upon determining that the at least one reference data set and / or the at least one data set to be analyzed is available in the data base: retrieving the at least one reference data set and / or the at least one data set to be analyzed from the data base.For example, upon determining that the at least one reference data set is not available in the data base: requesting a user to provide the at least one reference data set; and receiving at least one reference data set via a computer interface.For instance, upon determining that the at least one data set to be analyzed is not available in the data base: requesting a user to provide the at least one data set to be analyzed; and receiving at least one data set to be analyzed via a computer interface.According to an example embodiment of the method according to the first aspect, the method further comprising: Operating and / or controlling the production cycle of the distributed chemical production based on the at least one anomaly mitigation instruction, in particular to produce a chemical product. An operator may influence the parameters of a production line based on the at least one anomaly mitigation instruction, so that the quality of the produced product is increased compared to the production line being operated with unchanged parameters.According to an example embodiment of any aspect, the task instructing comprises an instruction to generate a technical report comprising a description indicating presence of the one or more anomalies or indicating absence of an anomaly, wherein, in case the one or more anomalies is detected, optionally, the technical report further comprising an indication of one or more sources of the one or more anomalies.According to an example embodiment of any aspect, the generative data-driven model is fine-tuned or re-trained using a technical training dataset (e.g. first training data set) comprising one or more plant technical documents and wherein the fine-tuned or re-trained generative data-driven model is used as the generative data-driven model in the determining the at least one anomaly mitigation instruction. For example, this fine-tuning may be in addition to fine-tuning using the at least one reference data set (e.g. as a second training dataset) comprising plant-based data associated with one or more production operations of one or more production lines of the distributed chemical production. For example, the generative data-driven model may be fine-tuned in a first training cycle using the technical training data set and this first cycle fine-tuned generative data-driven model may then be further fine-tuned using the reference data set (in a second cycle).According to an example embodiment of any aspect, the technical training dataset comprises datapoints associated with one or more production operations and associated with standard operating conditions of a production line, wherein the standard operating conditions are defined by a predefined range of one or more parameters associated with the one or more production operation of the production line, wherein optionally the one or more parameters is any one of a time period, a plant identification, a production line identification, a production cycle identification, one or more operating parameters associated with one or more production operation, an identification of a product, aplant, one or more parameters associated with one or more properties of a product, data type such as temperature, pressure, flow, mass.According to an example embodiment of the method according to the first aspect, the method further comprising:Determining (e.g. identifying) whether access and / or authorization credentials to access one or more databases comprising the at least one reference data set, the at least one data set to be analyzed, and / or the technical data set,Upon determining that access and / or authorization credentials are required, generating (e.g. by the generative data-driven model) a request for a user to provide an authorization and / or access credentials to access the one or more databases, retrieving (e.g. from a database) the at least one reference data set (e.g. by querying a database), the at least one data set to be analyzed, and / or the technical data set and optionally, using the at least one reference data set and / or the technical data set to fine-tune or re-train the generative data-driven model; and wherein the finetuned or re-trained generative data-driven model is used as the generative data-driven model in the determining the at least one anomaly mitigation instruction.Whether to use the at least one reference data set and / or the technical data set to fine-tune or re-train the generative data-driven model may be decided by an operator via a computer interface or it may be performed automatically after retrieving the at least one reference data set and / or the technical data set.According to an example embodiment of the method according to the first aspect, the method further comprising:Providing, via the computer interface, access to the at least one generative data-driven model, wherein the determining the at least one anomaly mitigation instruction is carried out using the accessed at least one generative data-driven model.According to an example embodiment of any aspect, obtaining the at least one reference data set, the at least one data set to be analyzed, and / or the technical data set comprises:Generating by the at least one generative data-driven model a retrieval query for obtaining the at least one reference data set, the at least one data set to be analyzed, and / or the technical data set;Providing the generated retrieval query to a database comprising the at least one reference data set, the at least one data set to be analyzed, and / or the technical data set for obtaining the at least one reference data set, the at least one data set to be analyzed, and / or the technical data set;Obtaining, from the database, the at least one reference data set, the at least one data set to be analyzed, and / or the technical data set based in the provided retrieval query.According to a second aspect an apparatus is disclosed, the apparatus comprising respective means for carrying out or performing the steps of the method according to the first aspect or comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to carry out the steps of the method according to first aspect.According to a further aspect a use of an anomaly mitigation instruction generated according to the methods according to the first aspect, or by the apparatus according to the second aspect is disclosed for displaying the anomaly mitigation instruction to an operator of the chemical plant and / or for producing a chemical product.According to a further example aspect, a system for operating a chemical plant (e.g. to produce a chemical product) is disclosed, the system comprising an apparatus according to any aspect, a database (server) providing the at least one reference data set, the at least one data set to be analyzed, and / or the technical data set, the database being communicatively coupled to the apparatus, a server providing the at least one generative data-driven model, the server being communicatively coupled to the apparatus, together performing or carrying out at least the steps of the method according to the first aspect (and / or any embodiment or example and combinations thereof of the method).According to a further example aspect, a computer element is disclosed, the computer element comprising instructions, which when executed by a processor or a computing apparatus perform or carry out the steps according to the methods or as defined by the apparatuses disclosed herein.According to a further example aspect, a computer program or computer program product is disclosed, the computer program or computer program product when executed by a processor causing an apparatus, for instance a server, to perform and / or control the actions of the method according the any aspect.According to a further example aspect, a (e.g. tangible and / or non-transitory) computer readable storage medium is disclosed, the computer readable storage medium comprising a computer program, the computer program when executed by a processor causing an apparatus, for instance a server, to perform and / or control the actions of the method according the any aspect.A generative data-driven model, e.g. a transformer-based model, may be trained on big data (e.g., several GBs of purpose unspecific, unstructured text and / or image data). Trained generative data-driven models such as trained transformer-based models may have improved capabilities for predicting data patterns such as patterns in natural languages. The improved capabilities may be attributed to large number of parameters obtained by said training. For example, transformer-based models such as OpenAI GPT may include 1 17 million parameters (GPT-1), 1 .5 billion parameters, 175 billion parameters (GPT-3), 170 trillion parameters (GPT-4). Said parameters allow said GPT models generating improved data output ascompared to other models such as recurrent neural networks (RNNs) or long short-term memory (LSTM) networks that do not comprise a transformer component."Attention Is All You Need" by Vaswani et al., 31 st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA (6 Dec 2017, arXiv:1706.03762v5) may describe a mechanism in machine learning comprising a transformer component (a transformer based model), which is hereby incorporated by reference.ISO / I EC 23053:2022(en), ISO / IEC TR 24372:2021 (en), ISO / IEC 22989, ISO / I EC 23053 may define standards in the field of Artificial Intelligence, Al and machine learning, ML. Big data may be specified in ISO / IEC 20546:2019(en) Information technology — Big data. Data quality may, for example, be specified in ISO / IEC 20546:2019(en), ISO 8000- 66:2021 (en) / Data quality, ISO / IEC DIS 5259-1 (en)Generative data-driven models such as transformer-based architectures may allow to capture long-range dependencies and parallelization of computation. Furthermore, generative data-driven models such as transformer-based architectures may be pre-trained on larger text-based datasets and fine-tuned (or re-trained) for specific tasks with smaller (labeled) datasets. Fine-tuning may be a process of taking a pre-trained generative data-driven model, e.g. trained on a large dataset, and further training it on a smaller, specific dataset, which may allow to transfer knowledge learned by the pretrained model to the specific task. During fine-tuning, the model's weights may be updated based on the provided specific dataset, wherein the pre-trained weights serve as a starting point, and e.g. only a small number of additional training steps are performed.Said fine-tunning or re-training may allow using the outstanding analytic capabilities of pre-trained transformer-based models for analyzing data patterns other than patterns in natural languages.This disclosure may relate to using generative transformer based models in analyzing data patters using a pre-trained transformer-based model for distributed production such as chemical production. Data suitable for re-training or fine-tuning a pre-trained transformer-based model for said production may be obtained by providing historic production data accumulated over more than 150 years.The amount of training data used to fine-tune (or re-train) a pre-trained transformer-based model may depend on the specific task, the complexity of the model, and the desired performance level. In many cases, transformer models can be fine-tuned (or re-trained) with smaller amounts of purpose-specific training data as compared to the size of pre-training dataset. However, if the specific application is very different from the pre-training data, re-training or fine-tuning a pretrained transformer-based model may require larger datasets as compared to the scenario when the specific application is similar to the pre-training data. For example, text classification or sentiment analysis may require few hundreds of MBs toa few GBs of labeled data. Machine translation may require tens to hundreds of GBs of text data for re-training or fine- tuning. Question-answering may require several GBs of training data for re-training a pre-trained transformer-based model.With better quality of training data, the amount of data needed for training may be reduced (the definition of the term “data quality” is, for example, given in ISO / IEC 20546:2019(en), ISO 8000-66:2021 (en) / Data quality, ISO / IEC DIS 5259-1 (en)). Pre-processing of data for generating high quality training data may involve labeling data, removing noise, irrelevant information and / or alike. Thus, pre-processing data for generating high quality training data may be beneficial for more efficient re-training or fine-tuning.Larger amounts of training data may allow fine-tuning or re-training pre-trained transformer-based models with more parameters. For example, in order to fine-tune Distil BERT model less training data may be required as compared to the amount of data necessary for fine-tunning GPT-4 model. Larger amount of said parameters may provide improved capabilities in data analysis (e.g., GPT-4 is more powerful in analyzing data patterns than GPT-3).A generative data-driven model may be a transformer-based model, such as TinyBERT, DistilBERT, Llama 7B, Mistral 7B, GPT-Neo, or a larger GPT variant. Further, pre-trained transformer based models may be, for example, ChatGPT (GPT-2, GPT-3, GPT-3-5-turbo, GPT-4), Davinci, BERT (Bidirectional Encoder Representations from Transformers), DistilBERT, Transformer-XL, XLNet (extreme Language understanding Network), T5 (Text-to-Text Transfer Transformer), RoBERTa (Robustly Optimized BERT approach), ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacements Accurately), Reformer, Longformer, DeBERTa (Decoding-enhanced BERT with disentangled attention). The properties, and thus, output data of said transformer-based models may vary for the same input data due to differences in architectures and / or pre-training datasets. Thus, one or more of said transformer-based models (GPT-2, GPT-3, GPT-3-5-turbo, GPT-4, Davinci, BERT, DistilBERT, Transformer-XL, XLNet, T5, RoBERTa, ELECTRA, Reformer, Longformer, DeBERTa) may be used as one alternative to a pre-trained transformer-based model for a technical purpose or one or more of said models may be used in a combination for a technical purpose to provide a plurality of data outputs for complimentary data analysis.Time used for fine-tuning or re-training may depend on the size of the generative data-driven model as well as on the size of the data set used for fine-tuning or re-training. Fine-tuning or re-training of generative data-driven models having a million to up to 2 billion parameters (such as TinyBERT with 4.4 million parameters, DistilBERT with 66 million parameters, GPT-Neo with 125 million parameters) may be performed in a few hours (e.g. 1 - 2 hours) using a data set of a few GB size (e.g. corresponding to 10 years of production), so below the time a full production cycle may take (e.g. 15 days or more) and may allow for fast adaptation of any parameters downstream in the production cycle, e.g. parameters effecting a subsequent production line. Larger generative data-driven models, e.g. having up to 10 billion parameters (such as Llama 7B, Mistral 7B), may take longer on the same size data set, e.g. 1 to 2 days, to fine-tune.A generative data-driven model, e.g. a transformer-based model, (generative artificial intelligence, Al model) in the context of the current application may refer to a foundation model (a machine learning, ML model) that may comprise a transformer component such as described in FIGs. 9-12.Generative artificial intelligence (Al) may refer to a computer program that may generate output as, for example, described in ISO / IEC 23053:2022(en), ISO / IEC 23053:2022(en), ISO / IEC TR 24372:2021 (en), ISO / IEC 22989, ISO / IEC 23053, ISO / IEC DIS 5259-1 (en), ISO / I EC 24661 :2023(en). A generative Al program may comprise a ML model such as a transformer-based model (generative pretrained transformer, GPT model). The generative pretrained transformer model (or simply, the transformer based model) may also be referred to as a foundation model.The definition of big data is given in ISO / IEC 20546:2019(en) Information technology — Big data. The definition of “data quality” is, for example, given in ISO / IEC 20546:2019(en), ISO 8000-66:2021 (en) / Data quality, ISO / IEC DIS 5259-1(en)."Attention Is All You Need" by Vaswani et al., 31 st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA (6 Dec 2017, arXiv:1706.03762v5) may describe a mechanism in machine learning comprising a transformer component (a transformer based model), which is hereby incorporated by reference.By using a pre-trained transformer based model retrained on production data, detecting and removing an anomaly early in a production cycle may allow for recovering the quality of products. In particular, early detection of an anomaly in an affected production cycle, wherein the anomaly occurs, may allow for restoring the quality of produced intermediate product(s) already during the affected production cycle or during a next production cycle resulting in an acceptable quality of final product at the end of the affected production cycle or at the end of the next production cycle.Due to the speed and data analysis properties of the proposed generative artificial intelligence, Al systems to determine deviations of data from data patterns, generative Al systems may be used for anomaly detection in plant production data to detect deviations from standard (reference) data patterns obtained during standard operation of a plant.In addition to early detection of anomaly in an affected production cycle, the aspects, embodiments and examples of this disclosure may also provide a more user friendly interface that is easy-to-operate for detecting anomalies. Even more, by combining both the plant’s operational documents and production data the efficiency of anomaly detection may be even further increased. The aspects, embodiments and examples of this disclosure may serve different, multiple plants in a distributed production environment and may be used by the plant operators without data science knowledge for finding an anomaly in plant operation. Thus, the aspects, embodiments and examples of this disclosure may provide improved method and system with improved user interface for efficiently guiding auser (i.e., operator of a plant) through the process of identifying and removing an anomaly associated with an operation of a production cycle.An optional continuous monitoring and notification system may provide further advantages in improving monitoring of production.The term „real time” for production may mean a time interval of up to a few hours e.g. depending on volume of training data and the size of the chosen model. For example, analyzing 10 years of production data resulting in e.g., hundreds of GBs, might take hours. In this example, it may be possible to detect anomaly within a few hours. In another example, analyzing 10 days of data resulting in, e.g., few GBs, might take minutes. In said example, detecting anomaly may take a few minutes. Data volume may depend on the complexity and size of the plant or part of the plant to be analyzed.DESCRIPTION OF THE DRAWINGSIn the following, embodiments of the present disclosure will be outlined by ways of examples. It is to be understood that the present disclosure is not limited to said embodiments and / or examples. The present disclosure has been described in conjunction with preferred embodiments and examples as well. However, other variations can be understood and effected by those persons skilled in the art and practicing the claimed invention, from the studies of the drawings, this disclosure and the claims.FIG. 1 illustrates an embodiment of detecting an anomaly in a production cycle.FIG. 2 illustrates an embodiment of a production line.FIG. 3A illustrates an operating system of the production line of FIG. 1 .FIG. 3B illustrates training a pre trained transformer-based model on production data generated by the production line of FIG. 1.FIG. 3C illustrates an embodiment of using a trained transformer-based model for detecting an anomaly in a production cycle and removing the anomaly in a production cycle.FIG. 4 illustrates a pre-processing engine of the operating system shown in FIG. 3.FIG. 5 illustrates production data associated with one or more production operations of a production line of a distributed production environment.FIG. 6 illustrates an embodiment of a user interface.FIG. 7 illustrates further aspects of training a pre-trained transformer-based model.FIG. 8 illustrates contextualization of prompts.FIG. 9 illustrates an embodiment of training an embedding layer.FIG. 10A illustrates an embodiment of a transformer encoder architecture.FIG. 10B illustrates an embodiment of a transformer decoder architecture.FIG. 10C illustrates an embodiment of a transformer encoder-decoder architecture.FIG. 11 illustrates an embodiment of training and / or deploying the transformer encoder, the transformer decoder and / or the transformer encoder-decoder.FIG. 12 illustrates an embodiment of input embedding.DETAILED DESCRIPTIONThe following embodiments are mere examples for implementing the method, the system, the apparatus or application device disclosed herein and shall not be considered limiting. The following description serves to deepen the understanding and shall be understood to complement and be read together with the description as provided in the above summary and embodiment sections of this specification. Some aspects may have a different terminology than e.g. provided in the description above. The skilled person will nevertheless understand that those terms refer to the same subject-matter, e.g. by being more specific.Any steps presented herein can be performed in any order, unless a specific order is disclosed. The methods disclosed herein are not limited to a specific order of these steps, unless the specific order is disclosed (e.g., pretraining of a transformer-based model precedes training; training precedes retraining; obtaining data for training precedes training a model; and / or alike that is directly and unambiguously disclosed in the application).It is also not required that the different steps are performed at a certain place or in a certain computing engine of a distributed system, i.e. each of the steps may be performed at different computing engines using different equipment / data processing.As used herein ..determining" also includes ..initiating or causing to determine", “generating" also includes ..initiating and / or causing to generate" and “providing” also includes “initiating or causing to determine, generate, select, send and / or receive”. “Initiating or causing to perform an action” includes any processing signal that triggers a computing node or device to perform the respective action.In the claims as well as in the description the word “comprising”, “including” or similar wording does not exclude other elements or steps and the indefinite article “a” or “an” does not exclude a plurality. A single element or other unit may fulfill the functions of several entities or items recited in the claims. The mere fact that certain measures are recited in the mutual different dependent claims does not indicate that a combination of these measures cannot be used in an advantageous implementation.Any disclosure and embodiments described herein relate to the methods, the systems, devices, the computer program element lined out above and vice versa. Advantageously, the benefits provided by any of the embodiments and examples equally apply to all other embodiments and examples and vice versa.All terms and definitions used herein are understood broadly and have their general meaning.An “engine” in the context of FIGs. 2-8 comprises at least one computer processor.A “computer interface” in the context of the current disclosure may be, e.g., a graphical user interface, an application programming interface, a web-based interface.For instance, a pre-trained transformer-based model may be pre-trained for a first purpose (e.g., analyzing unspecific text-based data) and re-trained / fine-tuned for a second purpose (e.g., production). A trained transformerbased model suitable for the second purpose may be further re-trained / fine-tunned for the second purpose in order to improve output data provided by the model."Distributed production environment" or“plant(s)” may refer, without limitation, to any technical infrastructure that is used for an industrial purpose of manufacturing, producing or processing of one or more products, i.e., a manufacturing or production process or a processing performed by the distributed production environment. The distributed production environment may be one “plant” (infrastructure) having distributed units for production.The distributed production environment may be one technical infrastructure (plant) comprising one or more production lines. Said one or more production lines may comprise one or more production operations. Said one or more production operations may be distributed.The distributed production environment may be more than one plant distributed in space and / or directed at distributed operations. The distributed production environment may be one or more of a chemical plant, a process plant, a pharmaceutical plant, a fossil fuel processing facility such as an oil and / or a natural gas well, a refinery, a petrochemical plant, a cracking plant, and the like. The distributed production environment may even be any of a distillery, a treatment plant, or a recycling plant. The distributed production environment may be a combination of any of the examples given above or their likes.The “product” produced by the one or more production lines of the distributed production environment by the end of one cycle of the production may, for example, be any physical product, such as a chemical, a biological, a pharmaceutical, a food, nutritional, a beverage, a textile, a metal, a plastic, a semiconductor, cosmetic or even any of their combination. Additionally, or alternatively, the product may be a service product, for example, recovery or waste treatment such as recycling, chemical treatment such as breakdown or dissolution into one or more chemical products. Some non-limiting examples of the chemical product are, organic or inorganic compositions, monomers, polymers, foams, pesticides, herbicides, fertilizers, feed, nutrition products, precursors, pharmaceuticals or treatment products, or any one or more of their components or active ingredients. In some cases, the chemical product may be a product usable by an end-user or consumer, for example, a cosmetic or pharmaceutical composition. The chemical product may be a product that is usable for making further one or more products, forexample, the chemical product may be a synthetic foam usable for manufacturing soles for shoes, or a coating usable for automobile exterior. The chemical product may be in any form, for example, in the form of solid, semisolid, paste, liquid, emulsion, solution, pellets, granules, powder.One or more production lines of the distributed production environment may comprise equipment or process units such as any one or more of a heat exchanger, a column such as a fractionating column, a furnace, a reaction chamber, a cracking unit, a storage tank, an extruder, a pelletizer, a precipitator, a blender, a mixer, a cutter, a curing tube, a vaporizer, a filter, a sieve, a pipeline, a stack, a filter, a valve, an actuator, a mill, a transformer, a conveying system, a circuit breaker, a machinery e.g., a heavy duty rotating equipment such as a turbine, a generator, a pulverizer, a compressor, an industrial fan, a pump, a transport element such as a conveyor system, a motor, etc.Further, the one or more production lines of the distributed production environment may typically comprise a plurality of sensors and at least one control system for controlling at least one parameter related to production process, or process parameter, in the plant. Such control functions are usually performed by the control system or controller in response to at least one measurement signal from at least one of the sensors. The controller or control system of the plant may be implemented as a distributed control system, DCS, and / or a programmable logic controller, PLC. The plurality of sensors may be distributed in the distributed production environment for monitoring and / or controlling purposes. Such sensors may generate a large amount of data. The sensors may or may not be considered a part of the equipment. Thus, production, such as chemical and / or service production, may be a data heavy environment. A distributed production environment may produce a large amount of process related data.Said sensors may be used for measuring one or more process parameters and / orfor measuring operating conditions of said equipment or parameters related to the equipment or the process units. For example, the sensors may be used for measuring a process parameter such as a flowrate within a pipeline, a level inside a tank, a temperature of a furnace, a chemical composition of a gas, etc., and some sensors can be used for measuring vibration of a pulverizer, a speed of a fan, an opening of a valve, a corrosion of a pipeline, a voltage across a transformer, etc. The difference between these sensors cannot only be based on the parameter that they sense, but it may even be the sensing principle that the respective sensor uses. Some examples of sensors based on the parameter that they sense may comprise: temperature sensors, pressure sensors, radiation sensors such as light sensors, flow sensors, vibration sensors, displacement sensors and chemical sensors, such as those for detecting a specific matter such as a gas. Examples of sensors that differ in terms of the sensing principle that they employ may for example be: piezoelectric sensors, piezoresistive sensors, thermocouples, impedance sensors such as capacitive sensors and resistive sensors, and so forth.The distributed production environment (comprising one or more production lines) may be a plurality of distributed production environments. The plurality of distributed production environments may be coupled such that the distributed production environments forming the plurality of distributed production environments may share one or more of their value chains, educts and / or products. The plurality of distributed production environments may also be referred to as a compound, a compound site, a Verbund or a Verbund site. Such Verbund sites or chemical parks may be or may comprise one or more distributed production environments, where products manufactured in the at least one distributed production environment may serve as a feedstock for another distributed production environment."Production" refers to any industrial process which when, used on, or applied to an input component provides an output product different from the input product. The production may thus be any manufacturing or treatment process or a combination of a plurality of processes that are used for obtaining the product as defined above. The production process may even include packaging and / or stacking of one or more of the products.The production process may be continuous, in campaigns, for example, when based on catalysts which require recovery, it may be a batch chemical production process. One main difference between these production types is in the frequencies occurring in the data that is generated during production. For example, in a batch process the production data extends from start of the production process to the last batch over different batches that have been produced in that run. In a continues setting, the data is more continuous with potential shifts in operation of the production and / or with maintenance driven down times. Thus, the required data analysis may be different based on the differences in the data flow, batch or continuous. For example, scheduled re-training or fine-tuning of a trained transformer-based model may be advantageous for batch data flows, while continuous re-training or fine-tuning of a trained transformer-based model may be advantageous for continuous data flows.The terms “production data", “plant data” or “plant-based data” may be used interchangeably and may relate to product data (e.g., properties of a product), process data (process parameters), operating conditions. Plant based data may refer to data comprising values, for example, numerical or binary signal values, measured during the production process, for example, via the one or more sensors. The process data may be time-series data of one or more of the process parameters and / or the equipment operating conditions. Typically, the plant-based data may comprise temporal information of the process parameters and / or the equipment operating conditions, e.g., the data contains time stamps for at least some of the data points related to the process parameters and / or the equipment operating conditions. The plant-based data may comprise time-space data, i.e., temporal data and the location or data related to one or more equipment zones that are located physically apart, such that time-space relationship can be derived from the data."Process parameters" may refer to any of the production process related variables, for example any one or more of temperature, pressure, time, level, etc. relevant for producing the product as defined above.The above definitions of a distributed production environment, products produced by the distributed production environment, production processes, the data generated by the production environment and the control of the production are mere examples and should not be construed as limiting. It may be understood that the system and the method of the claimed invention may apply to any kind of production producing a product and generating multiparameter data flows related to production. Any type of plant-based data may be pre-processed by a pre-processing unit to generate the required plant-based training data (labeled data, (pre)structured data, filtered data, data in the format of numbers and / or text, etc.) or plant-based input data suitable for re-training or fine-tuning a pre-trained transformer-based model or using a trained transformer-based model for production, respectively. Thus, the distributed production environment should be construed broadly as a technical environment producing a product (a physical product and / or a service associated with a product; a product may even be data product) and while producing said product generating, by said technical environment, multi-parameter production related data flows (plant-based data).An intermediate product of one production line may be used as an input product to the following production line of a production cycle. The final product produced at the end of one production cycle may be used as an input product for a different production process and / or plant producing a different product. Thus, the production processes and production lines of a distributed production facility are mutually interrelated and impact the final product, setting up challenging requirements for controlling and monitoring distributed production.FIG. 1 illustrates an embodiment of detecting an anomaly in a production cycle.As illustrated in FIG. 1 , chemical production may comprise one or more consecutive production cycles 1 , 2, ...., m. At the end of said one or more production cycles a final product may be produced. Quality of an intermediate / final product may be monitored to meet standard quality requirements. The quality control (quality check) of a product may be met when one or more quality parameters (e.g., color, strength, and / or alike) of the product is within an acceptable range (a predefined range for one or more parameters). Quality of a product may not be acceptable when one or more product parameters are outside of said predefined range for said one or more parameters. When quality of the product (intermediate or final) is not acceptable, the product may need to be recycled or discarded thereby generating waste and, in addition to that, resulting in a waste of production resources for producing said product.Said one or more productions cycles may comprise one or more of interconnected production lines 1 , 2, ..., n. Front end of said one or more production lines may receive an intermediate product from a back end of said one or more production lines. The front end of the first production line may receive raw material. At the end of each of the one or more production line quality of an output product may be measured following the same principles as describedabove. At the end of each of the one or more production lines 1 , 2, .... (n-1), an intermediate product may be produced that may become an input product to a consecutive production line. At the end of a production cycle m that ends at the last production line n, a final product may be produced. Quality of the final product may be measured or quantified following the principles described above.When an anomaly occurs in a production cycle (e.g., affected production line 2 of affected production cycle 1 ), the quality of a product (e.g., intermediate product, quality check 2) may not be acceptable. If unacceptable quality is not corrected, supplying the intermediate product from the back end of the affected production line (production line 2 of production cycle 1) may result in that quality of the final product is also unacceptable (e.g., final product of production line n of production cycle 1 ). If the source of anomaly is not identified and removed by the start of a consecutive production cycle (e.g., production cycle 2 in this example), the quality of an intermediate product or a final product of the consecutive production cycle (cycle 2), may be also compromised resulting in accumulation of product waste and wasting production resources until the source of the anomaly is finally corrected.During production cycles, operational parameters of each production line 1 , 2, ...n, may be measured. Based on the measurement, the operation of each production line 1 , 2, ... n may be monitored and / or controlled. In other words, based on the measurements, the operational parameters of each production line 1 , 2, ... n may be adjusted.The measured production data (plant data) may be used for analysis of typical data patterns associated with standard operational conditions of production lines.To ensure and maintain high quality production cycles, plant operators (users, operators of production lines) can utilize their operational knowledge associated with one or more production operations of a production line. The plant operators may further analyze the measured production data (plant data) to identify anomalies in data patterns (e.g., a deviation of one or more data points from a predefined range, e.g., standard / normal range, for one or more parameters). Monitoring anomaly in data may reduce risks of having reduced quality products in affected and consecutive production cycles.A production process of a final chemical product may be typically conducted in several production lines 1, 2, ... , n (dividing the whole process into sub-processes). At the end of each production line, there is a quality check performed. The quality check may mean that a quality of a product at the back end of each production line producing the product is withing an acceptable range for one or more parameters measured on the produced product, e.g., strength, color and other properties. In addition, a final quality check (such as strength, color and other properties) may be conducted at the end of the production cycle for the final product.If an anomalous behavior in any of one or more production lines 1 , 2, ..., n occurs, due to the features of the claimed invention (i.e., using a trained transformer-based model for analyzing data patterns in production data), said anomaly may be resolved at the time it occurs without waiting until the full cycle ends. For example, an anomaly may be identified until affected cycle 1 of FIG. 1 ends. Cycle 1 may comprise interconnected one or more production lines 1 , 2, ..., n, wherein each of said one or more production lines may comprise one or more interconnected production operations.Said interconnected production lines comprising interconnected production operations result in a complex chemical production. Said interconnected lines and operations of chemical production influence each other affecting the quality of the final product thereby resulting in an entangled processes setting up high requirements for data analysis, monitoring and control of chemical production.For example, in one production line (e.g., affected production line 2), the temperature of a certain sensor may be heated up, e.g., 10 degrees Celsius (°C) higher than it should be heated up under standard operating conditions. In this example, an operator may adjust the operating conditions in consecutive steps such that the quality of the final product or intermediate product in the affected or consecutive production cycle is not compromised. Detecting this anomalous rise in temperature and notifying this anomaly to an operator early (e.g., before the affected cycle ends or a next cycle begins) may result in that the operator may adjust the operating conditions thereby restoring the quality of the final product of the affected cycle or a next production cycle.Another example may be that in one production line, the quality check of one or more product parameters is not up to the standard (one or more product parameters may fall outside of a predefined range). Once this anomalous deviation of the one or more product parameters is detected and notified to an operator of a plant, the operator may adjust one or more production operations. For example, instead of adding the output product from this affected production line to the whole process (e.g. purified / distilled / synthesized raw material or intermediary component), the operator may decide what needs to be adjusted. For example, the operator may rectify and recover the quality to an acceptable quality or if it is not recoverable, the operator may decide whether they should continue with the cycle or not. Thereby the waste of resources may be reduced.The aspects, embodiments and examples provided herein may be directed at providing a system and a method applicable for different plants without the need of setting up separate investigations on a case basis. By using the system and method of the claimed invention the operator may not need to log the case of anomaly in plant data. The operator may not need to provide the case to a team of data scientist for investigation how to remove the source of anomaly.By the features of the current disclosure, an operator (e.g., human operator) may be provided with a user friendly computer interface to trigger an analysis of production data. The aspects, embodiments and examples of this disclosure may allow to restore the quality of production by detecting anomaly and identifying the source of the anomaly.The user interface may comprise interface elements providing functionality of translating natural language into database query for retrieving the production data. Said user interface may be integrated in one or more plants of the distributed production environment. Said user interface may provide transferrable services across variety of production lines, production operations, production facilities producing a variety of (interrelated) products, thereby improving the integration of the anomaly detection system in a distributed production environment.The aspects, embodiments and examples of this disclosure may allow for detecting and identifying a source of an anomaly in a production cycle fast, thereby allowing for restoring the quality of an intermediate / final product in aproduction cycle. By removing the source of anomaly early, less of products and production resources may be wasted. The advanced manufacturing technology, as described in the current application, may enable eliminating time-consuming efforts for removal of the anomaly and reducing waste. Thus, described advanced manufacturing technology of the current application may provide a climate mitigation technology for chemical production.FIG. 2 illustrates an embodiment of a production line.A distributed production environment may comprise equipment 1002 and sensors 1003 generating one or more sensor related data flows. The distributed production environment may produce one or more products as defined above, wherein properties of said one or more products may be measured, extracted or calculated generating one or more product related data flows (e.g., color, strength, and / or alike, as described in the context of FIG. 1 ). Plant data 1005 (plant-based data or production data) may comprise data obtained from each of said one or more data flows.Equipment 1002 may be any equipment of a distributed production environment such as pumps, heat exchangers, valves, reaction tanks, separation chambers and / or alike.Sensors 1003 may be any kind of sensors of a distributed production environment such temperature sensor, flow sensor, pressure sensor and / or alike.One or more products produced by one or more production lines may be any type of an intermediate or a final product, as described in the general description above. Properties of the products may be measured by, for example, gas chromatography.Plant data (production data) 1005 may be stored in a database, e.g., as historic production data. Plant data 1005 may be provided to the pre-processing engine that pre-process the data and provides plant-based input data to an analytics engine 1004 that may analyze the plant-based input data and generate machine readable instructions 1007 for control and / or monitoring engine 1006. Generation of the machine-readable instructions 1007 may be automatic (i.e., without involving an operator) by the analytics engine 1004 based on the analysis of the input plant data. For example, the analytics engine 1004 may continuously receive input plant data and may analyze said data in a continuous mode. When an anomaly occurs, the analytics engine may generate machine readable instructions 1007 for the control and / or monitoring engine to remove the anomaly. The analytics engine may identify solutions for improving production efficiency by analyzing plant-based input data on the background of production processes and may send machine readable instructions 1007 to the control and / or monitoring engine for improving production (e.g., machine readable 1007 instructions for how to remove the source of anomaly). The control and / or monitoring engine 1004 may display a push notification to an operator to review the instructions 1007 and based on the review may generate machine readable instructions 1008. The control system, based on the machine-readable instructions 1008, may change the operating parameters of one or more pieces of equipment 1002. Reviewing said machine readable instructions may allow for safe integration of the trained transformer-based model in the distributed production environment such as chemical production.Alternatively, or in addition to the automatic generation of machine-readable instructions 1007 by the analytics engine, an operator may prompt the analytics engine 1004 to provide machine-readable instruction 1007 based on a prompt.Machine readable instructions 1007 may be used by an operator for controlling and / or monitoring of one or more production operations of the distributed production environment.Machine readable instructions 1007 may comprise operating instructions for production such as machine-readable instructions for controlling equipment 1002 and / or machine-readable instructions for monitoring equipment 1002 and / or sensors 1003. Machine readable instructions 1007 may comprise operating instructions for an operator for controlling and / or monitoring the distributed production environment, wherein the operator may be a human based operator, computer-based operating system or a hybrid system comprising a human operator and a computer-based assistance system.The control system of the distributed production environment may comprise one or more computing units that may be able to manipulate one or more process parameters related to the production process by controlling one or more of the actuators or switches and / or end effector units, for example via manipulating one or more of the equipment operating conditions. The controlling is typically done in response to the one or more signals retrieved from the equipment.The control and monitoring engine may comprise one or more computer processors for revising machine-readable instructions 1007 and generating machine readable instructions 1008 for the control system of the distributed production. The control system of the distributed production, based on the machine-readable instructions 1008, may adjust the equipment operating conditions such that the adjusted process parameters and / or equipment operating conditions result in a controlled product (such as a chemical product) that has one or more required or predetermined properties or performance parameters. Production can thus be controlled on-the-fly whilst ensuring that the equipment operating conditions are adapted to undesired variations in the process parameters.It may be understood that control and monitoring of distributed production environment in general relates to controlling equipment and / or production lines for producing a product by sending machine readable instructions to the production environment. A product should be construed broadly as described above.The transformer-based model (first ML model) operated by the Al engine as described in the context of FIGs. 3a, 3b, 9-12 may be integrated via a computer interface with another computer program or a second ML model. The second ML may comprise a different algorithm to the first ML model. For example, the second ML may be a classical ML not based on a transformer architecture; a variation of the transformer based architecture of the first model or alike. Alternatively, the second ML may comprise the same architecture as the first ML model but the second ML model may be trained on a different dataset as compared to the first ML model. Training on a different dataset mayprovide different properties to the first and the second ML models. Different properties may allow to use the first and the second ML models in combination or as alternatives.For example, a second ML model may be a data driven model such as the one disclosed in WO2021156157 (A1 ) that may be integrated with the first model (transformer based model) via a computer interface. The second ML model may be used for pre-processing of raw plant-based data to generate plant based training data. Pre-processing of the raw data may involve removing noise, filtering data, labeling data, sorting data, converting raw data into a different format, converting operator data into a different format better suitable for the first ML model (transformer based model), and or alike. For example, pre-processing may comprise converting training data into text and / or numbers and usings said text and / or numbers for training as described in the context of FIGs. 9-12. Alternatively, instead of second ML model or in addition to it, pre-processing may be carried out by classical computer algorithms not involving ML (e.g., data labeling may not require ML). Pre-processing of raw plant-based data may improve the quality of training or input data based on the raw data and may reduce the computing power required for training / using the first ML model (transformer-based model).FIG. 3A illustrates an operating system of the production line of FIGs. 1 and 2.The operating system comprises (re)training or fine-tuning a pre-trained transformer-based model and using a trained transformer-based model for controlling and / or monitoring production such as chemical production. Training a pre-trained transformer-based model is further described in the context of FIG: 3B. Using the purpose trained transformer-based model is further described in the context of FIG. 3C.FIG. 3B illustrates training a pre trained transformer-based model on production data generated by a production line of FIG. 1 and 2.Training (fine-tunning, re-training) of a pre-trained transformer-based model, such as described in the context of FIGs. 9-12, comprises accessing a pretrained transformer-based model via a computer interface (GUI, API, webbased interface) and receiving, via the interface, plant historic data 1010 (as an example of a reference data set) from pre-processing engine 1009. An operator re-training or fine-tuning the pre-trained transformer-based model may prompt the pre-trained transformer-based model to access the plant historic data via the computer interface, or alternatively, the operator may upload, via the computer interface, the historic plant-based data from a database to the Al engine. Al engine may comprise at least one processor operating the pre-trained transformer-based model. The operator re-training or fine-tuning the pre-trained model may be a human operator, an automated operating system comprising a computer processor, or a hybrid operating system comprising a human operator and a computer assisted operating system prompting the pre-trained model to train, fine-tune or re-train on plant-based data.The plant historic data 1010 may be based on plant data 1005. Plant data 1005 and plant historic data may be stored in one or more databases (e.g. provided by one or more servers). Plant data 1005 may comprise a plurality of production data as described, for example, in the context of FIG. 5. During the training, the plant historic datamay be embedded via an embedding layer as described within the context of FIG. 9. Embedding the plant data may result in embedded plant data.The above examples of plant data are mere examples illustrating possible ways of implementing the aspects and embodiments of this disclosure. Plant data should be construed broadly and should be understood as any kind of data associated with producing a product by a technical infrastructure. Any type or format of data associated with production may be pre-processed by the pre-processing engine into data suitable for training or using the transformer based model for production, e.g., into text and / or numbers.At the end of a training cycle, the analytics engine 1004 may output (release, provide, generate) a trained transformer-based model suitable for use in a distributed production environment such as chemical production. The released trained model may be stored in a database for, e.g., version control of released models. The released trained model may be a computer program product. An access to the released trained model may be provided to a user as a data service for assisting production.Training plant-based data (e.g. in a reference data set or a technical training dataset) may be plant-based data in one or more languages. Training the pre-trained transformer-based models on training plant-based data in one or more languages may allow for enlarging plant-based training dataset, and additionally, may allow for providing operating instructions for production in one or more languages. Providing instructions in one or more languages may improve user interaction with the trained transformer-based model.FIG. 3C illustrates an embodiment of using a trained transformer-based model for detecting an anomaly in a production cycle and removing the source of the anomaly in a production cycle.The plant data 1005 may be provided to the pre-processing engine 1004 for pre-processing. Pre-processing the data 1005 may comprise steps as, for example, described in the context of FIG. 4. Pre-processing may also involve other steps such as labeling data, removing noise, structuring unstructured data, converting data into different formats, and / or alike required to provide training / input data based on plant data suitable for training or using a transformer-based model for production (e.g., text / numbers).The pre-processed data from the pre-processing engine 1004 may be used as plant-based input data for the trained transformer-based model for production.The trained transformer-based model may receive the plant-based input data and may, for example, predict anomaly in the plant-based input data. The trained transformer based model may identify the source of anomaly. The trained transformer based model may notify an operator. The notification may comprise instructions to an operator. Additionally or alternatively, the analytics engine 1004 may generate machine readable instructions 1007 for control and / or monitoring engine 1006.The analytics engine 1004 operating the pre-trained transformer based model may have at least one computer interface for operating the transformer based model (uploading data, providing prompts such as prompts, means for reviewing output generated by the model “Response”, and / or alike). The transformer based model (untrained, pre-trained, trained) may be stored in a database or a cloud. The model may be operated by the Al engine via the at least one computer interface (e.g., API; GUI).An operator of the trained transformer-based model may provide an input via a computer interface to the trained transformer-based model such as a text and / or audio query. The operator of the trained transformer-based model may additionally provide extracted data points, extracted, for example, from the data provided by the pre-processing engine. The extracted data points may, for example, relate to anomaly data points, typical or optimal operating parameters of the distributed production environment. The extracted data points together with the operator’s query may form a part of the input data to the trained transformer-based model.The trained transformer-based model may process the input data and provide a solution to the user query, for example, in the form of machine-readable instructions 1007.The control and / or monitoring engine may review the machine-readable instructions 1007.After the review, the control and / or monitoring engine 1006 may generate machine readable instructions 1008 that may be identical to machine readable instructions 1007, may be in part based on machine readable instructions 1007 or may be different from machine readable instructions 1007. An operator of the control and / or monitoring engine may use machine readable instructions 1007 merely for monitoring the production or may forward machine readable instructions 1007 as machine readable instructions 1008 for controlling the production. The operator may be a human operator and / or an operating system comprising a processor, and optionally, a human operator.The control and / or monitoring engine may trigger re-training or fine-tuning of the trained transformer-based model. The re-training or fine-tuning may be triggered by an operator such as a human operator or by an automated system comprising a computer processor. Alternatively or additionally, re- training or fine-tuning may be scheduled or continuous.The trained generative data-driven model (e.g. transformer based model) may be trained on training plant-based data in one or more languages. The trained generative data-driven model (e.g. transformer based model) may provide operating instructions for production in one or more languages. Providing operating instructions in one or more languages may be advantageous, for example, in that training data may be more available in one language than in another language. The trained generative data-driven model (e.g. transformer based model) trained on plant based training data in one or more languages may be set to provide operating instructions in a language in which the most training data is available. Additionally, or alternatively, the trained generative data-driven model (e.g. transformer based model) may be set to provide operating instructions for production in a language of choice of a user making the operation of the model more user friendly. Additionally, or alternatively, the trained generative data- driven model (e.g. transformer based model) may be requested to provide operating instructions for production in more than one language for cross checking operating instructions and selecting the most suitable instructions for improved production.The trained generative data-driven model (e.g. transformer based model) may be used for a continuous monitoring and notification system. In case an anomaly is detected automatically by the analytics engine operating the trained generative data-driven model (e.g. transformer based model), the analytics engine may send a notification to an operator notifying the operator of the detection of the anomaly.FIG. 4 illustrates a pre-processing engine of the operating system as shown in FIG. 3.Raw plant data 1005 as described in the context of FIGs. 4 and 5 may be provided to pre-processing engine 1009. The pre-processing engine 1009 may pre-process raw plant-based data to provide plant-based training data and / or plant-based input data to the generative data-driven model (e.g. transformer based model). Pre-processing steps may comprise selecting required parameters, merging / aggregating, calculating plant-based training data, e.g., calculating derived parameters, remove outliers and alike. Pre-processing may comprise filtering the data, removing noise, labeling of data, sorting data, converting data from formats that are not suitable for training / using the model into formats that are suitable for training / using the model and / or alike. The output data of the pre-processing engine may be stored in a database and used as plant-based training data for re-training or fine-tuning a pre-trained a generative data-driven model (e.g. transformer based model) as described in the context of FIGs. 3, 7 and 9 to 12. The output data of the pre-processing engine may also be used as input data for the trained generative data-driven model (e.g. transformer based model).FIG. 5 illustrates production data associated with one or more production operations of a production line of a distributed production environment.Plant data may be received from the distributed production environment via a computer interface such as a user interface as described in the context of FIG. 6. Alternatively, an operator may prompt the model to access the data via an alternative computer interface, such as an API or a web-based interface.The plant data may comprise different categories of plant data, e.g., sensor data, operating data, plant metadata, analytical data.Sensor data may relate to measured quantities available in production plants by means of installed sensors, e.g. temperature sensors, pressure sensors, flow rate sensors, etc.Analytical data may relate to quantities provided from analytics measurements of samples extracted at any point from a production plant such as a composition of a reactant, starting material, a product and / or a side product as determined e.g. via gas chromatography from samples extracted during production at different stages of the production process, e.g. before or after catalytical reactor(s).Operating data may relate to raw data (basic, non-processed analytical and / or sensor data), or processed or derived parameters (directly or indirectly derived from raw data).Plant metadata may indicate a physical plant layout and may include plant-specific quantities that describe, e.g., the properties of reactor(s), which are pre-defined by a physical plant lay out and may be relevant to the plant or reactor performance.The plant data 1005 may comprise text and / or numbers (structured data). Plant data may also be unstructured. The unstructured data (such as, for example, scans, datasheets comprising images, QR codes, and alike) may be pre- processed by the pre-processing engine and converted into text and / or numbers suitable for training of a pre-trained transformer-based model as described in the context of FIGS. 9 to 12.FIG. 6 illustrates an embodiment of a user interface.A user interface for operating a method according to the first aspects (e.g. comprising operating a transformer based model) may comprise elements of the user interface providing such functionality for a user as: selecting / upload ing training data (as an example of a reference data set); selecting / upload ing new data; upload engineering documents; detecting anomalies; generating machine readable instructions for monitoring and / or removal of anomalies (Response); outputting Response as text. Selecting / uploading training / new data and / or engineering documentation (as an example of a technical training dataset) may be implemented by one or more user prompts prompting a transformer-based model to select / upload training / new data and / or engineering documentation. The transformer based model may be prompted to analyse new data (as an example of a data set to be analyzed) and provide a Response to user query (as an example of an analyze instruction), e.g., whether there is / was anomaly in the new data, and if so, how to remove the source of the anomaly in a production cycle. In response to user prompt, the transformer-based model may provide Response comprising identification of one or more anomalies, identification of one or more sources of the one or more anomalies and / or machine readable instructions for removing the source of said one or more anomalies. Said machine readable instructions may comprise text and / or control instructions for monitoring and / or controlling equipment and / or sensors of a one or more production lines. The control instructions may comprise machine readable instructions 1007 (operating instructions), for example, for changing operating parameters of equipment 1002, predictive maintenance (e.g., replacing sensors 1003) and / or alike. The text may comprise operating instructions for an operator, e.g., to review an anomaly in data, review one or more sources of anomaly, instructions for how to remove the source of anomaly and / or alike.FIG. 7 illustrates further aspects of training a pre-trained generative data-driven model (e.g. transformer based model).A pre-trained generative data-driven model (e.g. transformer based model) model pre-trained on text and / or numbers (such as illustrated in FIGs. 9 to 12) may be trained for production based on training data associated with production. The training data based on plant historic data (as an example of a reference data set) may be converted into text / numbers by a computer processor, e.g., the pre-processing engine.(Re)training or fine-tuning of a pre-trained generative data-driven model may comprise one or more training cycles. The one or more training cycles may use one or more training datasets. As illustrated in FIG. 7, the training data associated with production may comprise training data 10131 for a first training cycle and training data 10132 for a second training cycle. Training data 10131 may comprise plant documentation data (e.g., engineeringdocumentation, operating manuals, technical reports, definition of plant’s normal condition, e.g, ranges of one or more parameters associated with one or more production operation as examples of technical training data sets). Training data 10132 (as an example of a reference data set) may comprise plant historic data (e.g., production data 1005 as described in the context of FIG. 5).For the first and the second training cycles, the steps as described within the context of FIG. 2b and FIG. 9-12 may be followed. In this case, generic text / number data used for training a generative data-driven model described in the context of FIGs. 9-12 may be replaced by the training data 10131 and / or 1032 (that is convertible into text and / or number formats). During the training (first and / or second training cycles), training data 10131 and / or 10132 may be embedded in the embedding layer as described in the context of FIG. 9.Trained during the first training cycle, the trained generative data-driven model may become trained (parametrized) to analyze plant data context, e.g., normal / standard operating condition of a plant.The trained generative data-driven model trained during the first training cycle may be further trained in a second training cycle using training data 10132. The training data 10132 may comprise data when the plant operates in normal conditions (condition comprising operating parameters without a predefined range). At the end of the second training cycle a trained generative data-driven model may be generated, wherein said model becomes suitable for analyzing production data patterns under normal / standard operating conditions (i.e., conditions comprising one or more parameters related to one or more production operation, wherein said one or more parameters is (are) withing a predetermined range, e.g., standard / normal range).Example 1 (Natural language prompt translated into a database query).A pre-trained transformer model as described in the context of FIGs. 9-12 may be trained on a set of database queries in one or more database query languages for providing a trained transformer based model suitable to generate database queries upon request.A plant operator training a pre-trained transformer-based model in a first and / or second training cycles may select training data 10131 , 10132 via a natural language prompt. The natural language of the operator may be interpreted by the trained transformer based model (trained using the set of database queries) into a database query to retrieve training data 10131, 10132 from one or more databases.A distributed production facility, such as chemical production, may comprise one or more databases associated with one or more database systems (e.g., SQL databases, EDL - Enterprise Data Lake and / or alike). Said one or more database systems may be connectable to connectors / APIs to retrieve data from one or more databases. Optionally, access to one or more databases may comprise an authorization procedure requesting a user / operator to authorize. An operator or user may require access rights to access said one or more databases.The transformer based model may be trained to interpret natural language instructions such as, for example, “Get production data between time X and Y for plant area key Z”.When trained to interpret natural language to identify which database to access, the transformer based model may be capable of generating appropriate database query (DB query) for, e.g., EDL production database. Said connector / API may allow to execute DB query and retrieve requested data from a one or more of database systems. DB query and data retrieval may be extended to different types of databases where a different query language is used. In this case, a pre-trained or trained transformer based model (trained in a first and / or second training cycles) may be trained on one or more query languages to analyze DB queries and retrieve data from one or more database systems comprising said one or more database query languages. In this case the training data for training a pretrained or trained transformer based model (trained in a first and / or second training cycles) may be a set of DB queries in one or more database query languages.Thus, pre-trained on text data (generic text) and additionally trained one or more DB query languages, the trained transformer-based model (or generative data-driven model) may be able to interpret natural language of an operator into a database query to access a relevant database eliminating the need for an operator to write the relevant DB query, setting up API connections to one or more databases. When the trained transformer based model pre-trained on generic text and trained on DB query languages is additionally trained on the first 10131 and / or second training data 10132, the trained transformer based model not only can retrieve one or more production-based data (10131 , 10132; 1005; 1011 ) requested by an operator, but also said model can analyze this data.When access / authorization is required to access one or more databases, the trained transformer-based model trained on the set of DB queries may prompt an operator / user to authorize / provide access rights to access the database wherein requested data by the operator / user may be stored. The trained transformer-based model may then retrieve the requested data automatically by a compute processor without the user having to write DB queries / setting up connectors to access the database. Automatic data retrieval carried out by a computer processor may accelerate retrieval of one or more production-based data (10131 , 10132; 1005; 1011) for training a pre-trained transformer-based model and / or using a trained transformer based model, e.g., for detecting an anomaly occurring in a production cycle, and optionally, assisting an operator in removing the source of the anomaly. For example, the trained transformer based model may retrieve data requested by a prompt “Get production data from April 07, 2023 between time 10:10 and 18:00 for plant area key Z”. The user / operator may then prompt the model to analyze this data and find an anomaly in this data. The user / operator may further prompt the model to identify the source of any detected anomaly. The user / operator may further prompt the model to provide instructions how to remove said source of the anomaly. Accelerated data retrieval (due to that a user / operator does not need to write DB queries and setting up connectors / APIs) may result in an accelerated anomaly detection and removal of the source of the anomaly in a production cycle, in turn, resulting in more efficient chemical production reducing waste and mitigating climate change.Example 2 (Data sorting, labeling, selecting for training, definition of quality of data for efficient training of a- pretrained transformer based model).An operator may define additional constraints or definitions for selecting training data (10132). For example, the operator may select training data withing a predefined standard range associated with standard operating condition of a production line (e.g., when one or more parameters, associated with one or more production operation of a production line, is withing a predefined range). The operator may label the data withing the predefined standard range as “good quality data”. The operator may select data outside the predefined standard range and label this data as “bad quality data”. The operator may select data associated with extreme operating conditions (e.g., under which a piece of equipment 1003 could explode, e.g., when one or more operating parameters exceeds a predefined extreme threshold) and label this data as “extreme data”. Sorting, labeling, categorizing and selecting a category of data fortraining may improve the efficiency of the training by, for example, reducing computing resources, increasing training speed, improving properties of the model to identify and remove the source of anomaly more efficiently and earlier in a production cycle.Training a pre-trained transformer-based model on “extreme data” may improve safety of controlling and / or monitoring of the distributed production. When the pre-trained transformer-based model is trained on extreme data, the trained transformer-based model may be prompted by an operator to generate one or more operating parameters associated with one or more production operation, wherein the one or more parameters do not reach the extreme threshold. Upon the prompt, the trained transformer-based model may generate machine readable instructions that do not go against safety range of one or more parameters associated with the one or more production operations of the distributed production (chemical production) thereby improving safety of the production environment and simultaneously providing a trained transformer based model that for safe deployment in the chemical production. Training a pre-trained transformer based model using the “good quality data” may improve the capabilities of the model to identify anomaly in the production data.A production cycle in a production plant may be a continuous process. For example, certain raw material flows through a reactor may have an increasing rate until it reaches its optimal rate (full capacity). An operator may sort, label and / or select “good quality data” by selecting the data associated with this optimal rate (e.g mass flow rate of raw material A = 1000 kg / h). An operator may sort, label and / or select “good quality data” by selecting datapoints associated with a one or more predefined thresholds (e.g., mass flow rate of raw material 1001 > A > 800 kg / h). At the end of a production cycle, the rate slowly decreases to zero or a residual flow. For choosing a production cycle running at its full capacity and optimal condition, the operator maz specify the data selection condition(s) (e.g. take the data after it reaches the desired flow rate of 1000 kg / h) thereby choosing the data labeled as “good quality data” for training.Example 3 (Uploading training or input data).An operator may upload data from one or more categories of data as described in the context of Embodiment 2 (technical documents, plant-based data, etc.) for training the model. An operator may also upload data (e.g., production based input data 101 1 , 1005, e.g., for a period of time) for analysis by the trained transformer based model. Uploading data by an operator rather than prompting a transformer based model to access the data in astorage may allow faster data analysis because authorization processes for accessing the data may not be required (e.g., the operator may have the data readily available for an upload).The uploading function of uploading data by an operator may be implemented in addition to or as an alternative to the function of providing data via a database query as, for example, described in the context of Embodiment 1 .In cases when data is not stored in a well-established storage system (e.g. where APIs / connectors are not fully developed); a pre-trained (or trained in the first and / or second training cycle) transformer based model may still be not trained on the relevant DB query language; a user may have problems with data access control (permission); or a quick analysis on small volume of data may be sufficient or necessary, the feature of uploading the data via a computer interface may be advantageous. In particular, in such cases the operators may upload the training data by themselves via the uploading feature of the computer interface eliminating said problems associated with accessing a storage or a database.Once a pre-trained transformer based model is trained in the first and the second training cycles, the trained transformer based model acquires an embedding layer that is based on the context and data patterns associated with the training datasets 10131 , 10132. Due to said embedding layer, the trained transformer based model acquires abilities to generate data distribution associated with, e.g., standard operating conditions, extreme operating conditions, not optimal operating conditions associated with “bad quality data”. As a result, the trained transformer based model may be able to compare said standard, extreme and not optimal data distribution with real-time production data for identifying any anomalies.The detected anomalies may be provided, via a user interface as a response (comprising, e.g., a text string and / or audio signal) to a plant operator in natural language. Outputting said response in the natural language to the operator may provide an improved man-machine interactions since the operator facing extreme situations related to anomalous behavior of production environment may be under time pressure and stress. Providing instructions for the operator in the natural language may enable the operator to act faster (as compared to the scenario when an output may be provided as an error message code or alike when the operator needs to interpret the error or look up the solutions, for example) in order to remove the source of the anomaly and thereby restoring the quality of the intermediate or final product in a production cycle.Additionally, or alternatively to providing the response as a text or audio response in the natural language, the response may comprise generating a technical report comprising a description of, e.g., one or more anomalies and, optionally, a one or more sources of the one or more anomalies, and optionally, operating instructions to the production environment for removing said one or more sources of anomalies. Said report may also be generated in a scheduled manner or upon request of an operator.At the end of each training cycle, a trainer of the model (operator that may be a human or a computer processor) may test the trained model as described in the context of FIG. 8, and based on the testing, the trainer may run additional training cycles or release the trained model for use in production as described, for example, in the context of FIGs. 3a, 3c; and 7, 8.The released trained transformer-based model may be stored in a database or a cloud. The released model may be operated by a computing unit / nod comprising a computer processor such as the analytics engine. A copy of the trained transformer-based model may be provided to a user for use and / or further training. Alternatively, only an access to operate the trained transformer-based model may be provided to a user.Instead of the pre-trained transformer-based model as described in the context of FIGs. 9 to 12, another foundation model may be used by the analytics engine 1004, such as, for example, ChatGPT (GPT-2, GPT-3, GPT-3-5-turbo, GPT-4 or higher / similar), Davinci, BERT (Bidirectional Encoder Representations from Transformers), DistilBERT, Transformer-XL, XLNet (extreme Language understanding Network), T5 (Text-to-Text Transfer Transformer), RoBERTa (Robustly Optimized BERT approach), ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacements Accurately), Reformer, Longformer, DeBERTa (Decoding-enhanced BERT with disentangled attention) and alike, or any other large language model pre-trained on big data such as generic text, image, video. Depending on availability of plant-based training data, availability of computing resources and required precision in analyzing data, a user may choose a foundation model with smaller or larger number of parameters. The transformer-based model shown in FIGs. 9 to 12 may be pre-trained to release a pre-trained model with the required number of parameters to suit the technical purpose of the user. The pre-trained model pre-trained in the context of FIGs. 2b, 9 to 12 may be further refined to release a model with even less parameters for improved computing speed and reduced computing resources.Re-training or fine-tuning described in the context of FIG. 7 and / or FIG. 3B may be continuous. Re-training or fine- tuning may be carried on the background of ongoing production without the need to stop / interrupt the production.FIG. 8 illustrates a muti-step tasks flow for using the model trained in a first and a second training cycles as described in the context of FIG.7 for chemical production.Using the model as trained in the context of FIG. 7 may comprise sequence of steps 1 to 11 described below. Stepl : A plant operator may intend to detect one or more anomalies in production data 1005 for Plant A. Step2: The operator may log into the system (e.g., analytics engine 1004).Step3: The operator may select the training data based on plant historic data 1010 using a prompt comprising natural language, e.g., “Select production data for plant area key ‘Plant A’ between 1st April, 2022 and 31st May, 2023 where the mass flow rate of F123X > 1000 kg / h”.Step4: The system (e.g., analytics engine 1004) may analyze the prompt and generate a database query for the production database, PIMS, e.g.,:ProductionData| where Areakey=="Plant A" and time between (datetime(2022-04-01 ).. datetime(2023-05-31))Step5: The system (e.g., analytics engine 1004) may execute this generated query and may retrieve the data (requested by the prompt of Step 3) from said PI MS.Step6: The operator may be additionally prompted to upload plant specific documents forming training data for additional training of the model.Step / : The system (e.g., analytics engine 1004) may generate an anomaly detection model using the training data and the uploaded documents, wherein the anomaly detection model means the trained transformer based model trained in the context of Steps 1 -6 above.Step8: Simultaneously, the operator may select new data (input data 101 1 based on production data 1005) for anomaly detection by providing a prompt to the trained transformer based model as described in the context of Step 7, wherein the prompt may comprise natural language, e.g., “Select production data for plant area key ‘Plant A’ from 1st of June 2023”.Step9: The system (e.g., analytics engine 1004) may retrieve this new data requested in Step 8 using the generated database query from PI MS of Step 4.StepI O: The system (e.g., analytics engine 1004) may compare data distribution under normal / standard operating condition (embedded in the trained transformer based model during, e.g., the second training cycle as described above) to this new data (requested in Step 8) for finding any anomaly in this new data.Stepl 1 : The system (e.g., analytics engine 1004) may generate a notification- (Response) to the operator, e.g.,: “There is a sudden drop of temperature in sensor T7235 by 10 degrees on 5th of June 2023 and remained similar for the rest of time frame. The value of F6324 remained normal with the exception that there was a sudden increase of 12% than expected on 17th of June 2023 between 16:04 and 17: 15 before being stabilized again. The average pressure at P7692 is 30% higher than expected over the period.”FIG. 9 illustrates an embodiment of training an embedding layer.The embedding layer may be obtained by training for example a continuous bag of words model (CBOW) or a skipgram model. The embedding layer may be suitable for generating embedded input data based on input data. Generating embedded input data may refer to embedding input data. Embedding input data may result in a representation associated with the input data. Thus, the embedded input 1 14 may be the representation associated with the input data. The input data may comprise one or more elements. The one or more elements may be represented by the input vector 106. In particular, the embedded input 114 and / or the input vector 106 may be machine- readable and / or processable by a processor. For this purpose, the embedded input 114 and / or the input vector 106 may be a tensor, in particular a first-rank tensor. Specifically, the input vector 106 may be a one hot vector or a summation of a plurality of one hot vectors. A one hot vector may be a vector with one entry unequal to zero. Examples for one hot vectors may be 108, 110 and 112. The entries unequal to zero in the one hot vector and / or in the input vector 106 may indicate the element. For example, a look up table may define the relation between the position of the entries unequal to zero and the element indicated by the one hot vector. The look up table may specify a plurality of different elements. The number of different elements may be equal to the number ofentries in the one hot vector. The number of different elements may be referred to as vocabulary size. In an example, the elements may be represented by tokens and a sequence of elements may refer to at least a part of a sentence. The at least a part of the sentence may be represented by a plurality of tokens. A token may represent at least a part of the element and / or word. For example, where one element would be associated with only one word, words such as “embeddings", “embedding” or “embed” would constitute different elements. A first token may represent the stem “embed” and the endings, typically appearing in a plurality of word, may be represented by a second token, a third token and a fourth token. The second token, the third token and the fourth token may be used for representing other words such as “look”, “looking” or the like, preferably together with a fifth token representing the stem “look”. Ultimately, this tokenization of elements associated with a plurality of stems and a plurality of endings results in less tokens to be used for representing a plurality of elements and thus, uses less computational resources.A look up table specifying a subset of the vocabulary size eg of the English language may comprise 10,000 words or more. The embedded input 114 may be a lower-dimensional representation than the input vector 106. For example, typical embedded inputs 114 may comprise some hundreds of different entries. Followingly, the embedded inputs 1 14 constitute a densified representation of one or more elements using less computational resources. More than that, the embedded input 114 may represent a relation between two or more elements. For example, the words “Italy” and “Germany” may be similar or may be more closely related since they both define european countries, whereas the the word “embodiment” may be very different from the two respective words. The smaller the dot product between two embedded inputs 114 may be the more similar the two elements associated with the embedded inputs 114 may be. Hence, the embedded inputs 1 14 may represent one or more elements accurately and lead to accurate results based on processing the embedded inputs 114.For transforming the input vector 106 into the embedded input 114, the embedding layer may comprise a number of neurons equal to the number of entries in the embedded input 114. Based on the embedded inputs 114, the output layer may generate the output vector 116. The output vector may be a vector and / or may indicate one or more elements. The output vector 116 may indicate one or more elements different from the input vector 106 and / or the one hot vectors associated with the input vector 106. For this purpose, the output layer may comprise a number of neurons equal to the number of entries of the input vector 106 and / or the output vector 116. The output layer may apply a softmax function to the embedded inputs 1 14. By doing so, the output vector may comprise the probabilities associated with the elements associated with the entries of the output vector 116 unequal to zero. Hence, from the output vector 116 one or more elements may be obtained with a corresponding probability. Where the input vector 106 may specify one or more sequence(s) of elements, the output vector 116 may specify one or more elements corresponding to the sequence(s) of elements specified by the input vector 106. In the example of FIG. 9, the element associated with vector 118 may correspond to the input vector with a probability of 71 %. Additional or alternative elements may correspond to the input vector as indicated by the output vector with lower probability. By defining a threshold to which the probability may be compared, the selection of the corresponding elements may be tailored to the needs of the user. The elements generated by the model comprising the embedding layer 102 andthe output layer 104 may refer to the most probable elements indicated by the output vector 116. Hence, the model depicted in FIG. 9 may generate the element associated with the vector 1 18 with a confidence score of 71 %.The model of FIG. 9 may be continuous bag of words (CBOW) model. The CBOW model may be trained based on a training data set comprising a plurality of input vectors and corresponding output vectors. As the training data set may not be labeled, the training of the CBOW model may be referred to as self-supervised. Before training of the CBOW model, the CBOW model may be initialized with random values assigned to the weights of the neurons. During the training of the CBOW model, the input vectors may be passed through the initialized embedding layer and the output layer and a loss may be determined by comparing the output vector obtained by passing the input vector 106 through the model to the output vector corresponding to the input vector 106 as specified by the training data set. Based on the determined loss, backpropagation may be applied to determine the gradients associated with the neurons of the embedding layer 102 and the output layer 104 to lower the loss. According to the determined gradients, the weights of the neurons may be updated by using a gradient descent algorithm. If a predetermined loss may be achieved by the CBOW model, the training may be terminated and a trained CBOW model may be obtained. From the trained CBOW model, the embedding layer 102 may be suitable for embedding input data comprising one or more elements. This embedding layer 102 may be used in other machine-learning architectures requiring an embedding layer 102 such as a transformer encoder, transformer decoder or transformer encoder decoder architecture as described within the context of FIG. 10A, FIG. 10B and FIG. 10C. For training these architectures, a trained embedding layer 102 may be required. Hence, a model such as a CBOW model may be trained prior to training the transformer encoder, transformer decoder or transformer encoder decoder architecture.FIG. 10A illustrates an embodiment of a transformer encoder architecture.The transformer encoder comprises an encoder input 278, one or more encoder blocks 274, 214 and an encoder output. The transformer encoder architecture may be derived from the transformer encoder-decoder architecture as known in the art and shown in FIG. 10C. In particular, the transformer encoder may be referred to as X-former. The transformer encoder architecture may correspond to the encoder architecture associated with the transformer encoder-decoder architecture with an additional encoder output instead of connecting the encoder block directly to the decoder of the transformer encoder-decoder architecture. A plurality of transformer encoder architectures are available in the art such as the bi-directional encoder representations from transformers (BERT).The input data may be received at the encoder input 278. The encoder input 278 may apply an input embedding 202. Applying the input embedding 202 may refer to passing the input data through an embedding layer eg as described within the context of FIG. 9. Further, the encoder input 278 may apply positional encoding 204. Applying positional encoding 204 may refer to adding a positional factor to the embedded input obtained via input embedding. Preferably, the input data may specify a sequence of elements. The positional factor Ppos may be indicative of the position of the elements within the sequence.For example, the positional factor Pp°smay be obtained based on the following equation:where pos may refer to the position of the element within the sequence, I may refer to the dimension associated with the input embedding and d may refer to the dimension of the model, eg transformer decoder, transformer encoder or transformer encoder-decoder. This may be referred to as absolute positional embeddings. Alternatively, the positional encoding may be based on rotary positional embeddings (RoPE). Positional encoding is beneficial since it enables the processing of sequential data without requiring further dimensions indicating the position of each element. Followingly, the positional encoding 204 reduces the computational resources needed for embedding the input data. By passing the input data through the encoder input, the input data may be transformed into a second-rank tensor representing the sequence of elements. This second-rank tensor may be referred to as embedded input data. The embedded input data may be processed by the encoder block. The embedded input data may be provided to the layer normalization 208 by a residual connection. Multi-head self attention 206 may be applied to the embedded input data. Multi-head self attention 206 may comprise the two components multi-head and self-attention. Self-attention may be understood as being a filter applied to the embedded input data. By applying the filter to the embedded input data, the elements associated with the embedded input data contributing to the to be generated output data may be identified for generating the output data. Hence, the filter may represent the degree of contributing to the to be generated output data by the elements associated with the embedded input data. Applying the filter may be referred to as weighting the elements associated with the embedded input data. This is advantageous specifically regarding long sequences of elements. The filter may be learned and improved during the training by learning to identify the contribution of elements associated with the embedded input data. For example, in the partial sentence “I went to the bakery to buy a” the last word may be generated by the data-driven model such as the transformer encoder. The self attention may focus the transformer encoder to attend to the word “bakery” and “buy” mostly to generate the word “bread”. Self attention may refer to attention generated based on the input data. Hence, the filter may be determined based on the input data, preferably the embedded input data. The embedded input data may serve as query Q, key K and value V with respect to the self attention operation. The self attention may refer to attention based on the received input data. Hence, the filter may be calculated based on the following formula by inserting the respective tensors based on the embedded input data:where dk corresponds to the dimension of the key.For improving the efficiency of the transformer encoder further, the multiple heads are used to apply the filter resulting in the multi-head self attention 206. Multi-head self attention 206 may comprise applying the filter to twoor more parts of the embedded input data. Hence, the tensor may be split into two or more parts and the filter may be applied to the two or more parts separately by two or more heads according to the following equation:with parameter matrices may refer to thenumber of heads, ^V, dx and dq mayreferto the dimensions of the value, key and query.The result of the two or more head may be concatenated according to the following equation: MultiHead(Q, K, V) = Concat {head 1, . . . , headh)WQ^hdvxd and h may refer to the number of heads.The embedded input data may be transformed via the multi-head self attention 206 into a context tensor. The context tensor may represent the sequence of elements and the relation between two or more elements of the input data. The context tensor may be a second rank tensor and / or may comprise one or more first rank tensor(s). After the multi-head self attention 206 layer normalization 208 may be applied based on the context tensor and / or the embedded input data from the residual connection. Applying layer normalization 208 may refer to normalizing the context tensor. Normalizing the context tensor may lower the values of the entries of the context tensor. This reduces the computational cost associated with processing the context tensor. Layer normalization 208 may be followed by passing the context tensor to a feed forward layer 210 again followed by layer normalization 212 based on the residual connection to the context tensor and / or the output of the feed forward layer 210. The feed forward layer 210 may be a feed-forward neural network. The feed-forward neural network may comprise of a plurality of fully connected neurons. Passing the context tensor through the feed-forward neural network may result in transforming the context tensor linearly. Additionally or alternatively, the neural network may comprise one or more activation functions such as a rectified linear unit (ReLU). Hence, the neural network may be configured for performing one or more non-linear operations to the context tensor and / or transforming the context tensor non-linearly. After the context tensor has been transformed and / or normalized by the feed forward layer 210 and the layer normalization 212, the context tensor may be provided to one or more further encoder blocks 214. Having passed the context tensor through the feed forward layer 210 may adapt the context tensor for the processing by a further attention layer of the one or more further encoder blocks 214 for applying a self attention filter, preferably multi-head self attention 206. The context vector after being transformed by the layer normalization 212 and the feed forward layer 210 may be referred to as hidden state.The encoder output 276 comprises of a linear layer 216 and a softmax layer 218. The linear layer 216 may transform the context vector into a logits vector. The linear layer may be fully-connected. The logits vector obtained by passing the context tensor through the linear layer 216 may be passed through the softmax layer 218. Passing the logits vector through the softmax layer 218 may refer to applying the softmax function to the logits vector. Applying the softmax function to the logits vector may result in a probability distribution of one or more elements corresponding to the sequence of elements in the input data. From the probability distribution based on predefined selection criteria,one or more elements may be chosen. The one or more chosen elements may be referred to as the one or more elements generated by the transformer encoder. The one or more generated elements may be provided to the encoder input for generating further one or more elements corresponding to the sequence of the input data and the one or more elements generated by the transformer encoder as described within the context of FIG. 11 .FIG. 10B illustrates an embodiment of a transformer decoder architecture.The transformer decoder comprises a decoder input 284, one or more decoder blocks 280, 232 and a decoder output 292. The transformer decoder architecture may be derived from the transformer encoder-decoder architecture as known in the art and shown in FIG. 10C. The transformer decoder may be referred to as X-former. The transformer decoder architecture may correspond to the decoder architecture associated with the transformer encoder-decoder architecture independent of receiving one or more hidden states from the encoder of the transformer encoder-decoder. A plurality of transformer decoder architectures are available in the art such as the generalized pretrained transformers (GPT).The decoder input 284 may apply input embedding 220 and positional encoding 222 analogous to analogous to the input embedding 202 and the positional encoding 204 as described within the context of FIG. 10A.The decoder block 280 may comprise the layer normalizations 226, the masked multi-head self attention 224, the feed forward layers 228 and / or the layer normalization 230. The embedded input data resulting from passing the input data through the decoder input 284 may be provided to the layer normalization 226 via a residual connection. Further, masked multi-head self attention 224 may be applied to the embedded input data. Masked multi-head self attention 224 corresponds to the multi-head self attention 206 as described within the context of FIG. 10A with additionally masking a part of the embedded input data associated with elements later in the sequence than the element to be generated. Additionally or alternatively, the part of the input data associated with elements later in the sequence than the element to be generated may not be received and / or transformed into the embedded input data. Thus, the transformer decoder may be suitable for generating a subsequent element to a sequence, whereas the transformer encoder may be suitable for generating a missing element in within one sequence and / or between two or more sequences. Therefore, the transformer encoder may be configured for classification tasks. The transformer decoder may be configured for text generation.Similar to the transformer encoder as described within the context of FIG. 10A, a context tensor may be generated by applying the masked multi-head self attention 224 and the layer normalization 226. The context tensor may be provided to the layer normalization 230 via a residual connection. Further, the feed forward layer 228 and the layer normalization 230 may be analogous to the feed forward layer 210 and the layer normalization 212 as described within the context of FIG. 10A. The context tensor may be provided to one or more further decoder blocks 232.The decoder output 292 may comprise of a linear layer 234 and a softmax layer 236. The linear layer 234 and the softmax layer 236 may be analogous to the linear layer 216 and the softmax layer 218 as described within the context of FIG. 10A.FIG. 10C illustrates an embodiment of a transformer encoder-decoder architecture. The transformer encoderdecoder may comprise the encoder input 288, the one or more encoder blocks 286, 264, the decoder input 294, the decoder block 290 and the decoder output 292. The encoder input 288 may correspond to the encoder input 278 of FIG. 10A. The one or more encoder block 286, 264 may correspond to the one or more encoder blocks 274, 214 of FIG. 10A. The decoder input 294 may correspond to the decoder input 284 of FIG. 10B.The decoder block 290 may comprise a masked multi-head self attention 270, a layer normalization 272, a feed forward layer 238 and a layer normalization 240 analogous to the masked multi-head self attention 224, the layer normalization 226, the feed forward layer 228 and the layer normalization 230 as described within the context of FIG. 2B. The decoder block 290 may further comprise a multi-head self attention 250 and a layer normalization 248. Analogous to the description of FIG. 10B, the context tensor may be obtained from the masked multi-head self attention 270 and the layer normalization 272. Multi-head self attention 250 analogous to the multi-head self attention 206 of FIG. 10A may be applied to the context vector obtained from the layer normalization 272 and the hidden states of the one or more encoder blocks 286, 264. Layer normalization 248 may be applied to the context vector obtained from the multi-head self attention 250 and the context vector obtained from the layer normalization 272 provided via a residual connection. The context vector resulting from the layer normalization 248 may be processed via the feed forward layer 238 and the layer normalization 240 analogous to the description of FIG. 10B. The context vector resulting from the layer normalization 240 may be provided to further decoder blocks 242 analogous to the decoder block 290. The context vector obtained from the one or more decoder blocks 290, 242 may be provided to the decoder output 292. The decoder output 292 may correspond to the decoder output 282 of FIG. 10B.With the above-described architecture, the transformer encoder-decoder may receive and process input data at the encoder input 288 and the one or more encoder blocks 286, 264 and the decoder block 290 and the decoder output 292. Based on the input data, the transformer encoder-decoder may generate output data part by part or sequentially. The sequentially generated output data may be provided to and / or may be processed by the decoder input 294, the one or more decoder blocks 290, 242 and the decoder output 292. Preferably, a sequence may be provided to the encoder input 288 and after having generated at least a part of the output data, the decoder input 294 may be provided with at least the part of the elements of the output data already generated. By doing so, the next elements of the output data may be generated with a higher accuracy by taking the input data and the generated output data into account since more data is received by the transformer encoder-decoder may be received over time.Because of the transformer encoder-decoder architecture, the transformer encoder-decoder may be configured for transforming a sequence into another representation of the sequence. An example for transforming one sequence into another representation may be translation of one sentence into another language. A plurality of transformer encoder-decoders are available in the art such as BART, T5 or the like.In an embodiment, the layer normalization 208, 212 may be applied prior to the masked multi-head self attention 224, multi-head self attention 206 and / or the feed forward layer 210 in the transformer decoder, the transformerencoder and / or the transformer encoder-decoder. By doing so, the computational resources for applying the multihead self attention 206 and / or the feed forward layer 210 to the embedded input data and / or the context tensor may be decreased as the entries of the respective tensors may be lower after normalization.In an embodiment, the decoder output 292 may comprise of a classification neural network, further feedforward layers, convolutional layers, fully connected layers or the like. For example, the transformer encoder-decoder may be configured for choosing between a plurality of options. For this purpose, the transformer encoder-decoder may be provided with three different input data sets and may classify the context vectors obtained from the one or more decoder blocks 290 via one or more linear layers. Followingly, the architecture may be extended depending on the use case to be solved. [1]FIG. 11 illustrates an embodiment of training and / or deploying the transformer encoder, the transformer decoder and / or the transformer encoder-decoder.The encoder / decoder / encoder-decoder architecture 302 may correspond to the transformer decoder, the transformer encoder and / or the transformer encoder-decoder as describe within the context of FIG. 10A- FIG. 10C. The output data generated by the encoder / decoder / encoder-decoder architecture 302 may comprise of one or more elements, in particular a sequence of elements. The previously generated elements of the output data may be provided as input for generating the next element in the sequence of the output data.In the example of FIG. 11 , the input data may comprise of N elements, in particular input tokens. An input token may be a token dedicated to be inputted into a data-driven model such as the transformer decoder, the transformer encoder or the transformer encoder-decoder. The output data to be generated may comprise of M elements. The encoder / decoder / encoder-decoder architecture 302 may generate one element of the output data based on receiving the input data and optionally previously generated elements of the output data at a timestep. Hence, for generating M elements M time steps are required. A time step comprises of providing input 310, 312, 314 to the encoder / decoder / encoder-decoder architecture 302 and receiving output data 304, 308, 306 from the encoder / decoder / encoder-decoder architecture 302. In a first timestep, the input 310 may comprise of N input tokens. The N input tokens may be associated eg with N words, stems or endings. Preferably, the N input tokens may specify a question. One or more input tokens may specify the beginning of the sequence of tokens and / or the end of the sequence of tokens. The input 310 may be processed by the encoder / decoder / encoder-decoder architecture 302. Based on the input 310 at least a part of the output data 304 may be generated. The at least a part of the output data may comprise a first output token. In the next timestep, the generated first output token may be provided together with the input 312. Specifically, where the input 312 may be received by a transformer encoderdecoder the input tokens may be received at the encoder input 288 and the first output token may be received at the decoder input 294. Where the input 312 may be received by the transformer encoder, the input 312 may be received by the encoder input 278 and analogously regarding the transformer decoder and the decoder input 284. Based on the input 312, the output data 308 comprising the first output token and a second output token may be generated. Generating the output data 308 based on the input 312 may refer to generating the second token basedon the first token and the N input tokens, wherein the first token may have been generated based on the N input tokens. This process may be repeated until the last token in the sequence of the output data 306 may be generated. Preferably, the last token may be an end token. The end token may terminate the generation of a further output token.Similarly, to the data processing during deployment of the encoder / decoder / encoder-decoder architecture 302, the encoder / decoder / encoder-decoder architecture 302 may be trained. The training data set may comprise a plurality of sequences comprising a plurality of elements. The sequences may be associated with the input data and / or the output data. Additionally or alternatively, the sequences may be independent of the input data and / or the output data. For example, where the input data and the output data may refer to chemical compositions represented via text, the training data set may comprise sequential text data independent of chemical compositions. In this example, the training data set may comprise sequences of words originating from a conversation. In an embodiment, the training data set may comprise at least partially input data sets and / or output data sets.The training may be initialized by initializing the encoder / decoder / encoder-decoder architecture 302. In an embodiment, the parameters associated with the encoder / decoder / encoder-decoder architecture 302 may be initialized randomly. Additionally or alternatively, the input embedding of the encoder / decoder / encoder-decoder architecture 302 may be obtained by training a CBOW model or a skip gram model as described within the context of FIG. 9. The trained embedding layer may be used during training. The parameters associated with the embedding layer may be kept constant and / or may be updated after a predefined number of training epochs. By doing so, the number of parameters to be updated is lower enabling a faster and less computational resources-consuming training. Further, the accuracy associated with the embedding layer may be constant and / or may be increased by avoiding error compensation in relation to the just initialized encoder / decoder / encoder-decoder architecture 302.During the training of the encoder / decoder / encoder-decoder architecture 302, at least a part of the sequences of the training data set may be provided to the encoder / decoder / encoder-decoder architecture 302 one by another and one or more elements may be generated based on the sequences of the training data set one by another. The elements generated based on the sequences may follow the elements of the parts of sequences the encoder / decoder / encoder-decoder architecture 302 may have been provided with. The generated one or more elements may be compared to the one or more elements following the at least a part of the sequences provided to the encoder / decoder / encoder-decoder architecture 302 as specified by the training data set. Hence, during the training the encoder / decoder / encoder-decoder architecture 302 may generate a guess on the next element and the guess on the next element in a sequence may be compared to the ground truth specifying the actual next element according to the training data set. Based on the guess on the next element and the ground truth a loss may be determined. The loss may define the similarity between the guess on the next element and the ground truth. The loss may be determined by forming a vector dot product between the token associated with the one or more elements and the token associated with the ground truth. A loss unequal to zero may result in updating the parameters associated with encoder / decoder / encoder-decoder architecture 302. Preferably the parameters associated with the encoder / decoder / encoder-decoder architecture 302 may be independent of the embedding layer. For example, theparameters associated with the encoder / decoder / encoder-decoder architecture 302 may be weights of the neurons of the encoder / decoder / encoder-decoder architecture 302.Based on the determined loss, backpropagation may be applied to determine the gradients associated with the parameters of the parameters associated with encoder / decoder / encoder-decoder architecture 302 to lower the loss. According to the determined gradients, the parameters associated with the encoder / decoder / encoder-decoder architecture 302, preferably the weights of the neurons associated with the encoder / decoder / encoder-decoder architecture 302, may be updated by using a gradient descent algorithm.The training data set may be unlabeled. The sequences of elements within the training data set may inherently comprise the ground truth for determining the loss with respect to the one or more elements generated during the training of the encoder / decoder / encoder-decoder architecture 302. Hence, the encoder / decoder / encoder-decoder architecture 302 may be trained self-supervised. This is advantageous since time and resources for creating a labeled training data set may be saved. Furthermore, this enables the usage of large training data sets associated with a size of several tera bytes. Consequently, the data-driven model may be accurate in generating elements of a sequence. In addition, the large training data set enables few shot predictions or even zero shot predictions. Hence, the data-driven models trained as described above are versatile contributing to saving resources needed for training and / or hosting a plurality of purpose-driven models such as CNNs. The training described above may be referred to as pretraining. The data-driven model may be configured for performing few shot or even zero shot predictions with respect to a plurality of use cases after pretraining. The performance of the data-driven model may be increased further by additional training referred to as fine-tuning.FIG. 12 illustrates an embodiment of input embedding.Where the sequence of elements associated with the input data, preferably comprised in the input data, may be of one type, the input embedding 202, 220, 252, 266 as described within the context of FIG. 10A - FIG. 10C may be used. For example, a type of input data may be text where the elements may be associated with at least a part of a word, a punctuation character, a start token specifying the beginning of one or more sequences associated with the input data and / or the end token. In another example, the input data may be at least partially numerical. Hence, the input data may comprise a plurality of numbers. Numerical input data may be for example tabular data. Tabular data may specify one or more rows and / or one or more columns. Hence, the tabular data may comprise one or more cells, wherein the cells may be associated with one or more numerical values.Numerical input data may require a different embedding than text input data. Input embeddings for numerical input data may comprise a token embedding, a positional embedding, a column embedding, a row embedding or a combination thereof.Applying a token embedding to one or more elements, in particular tokens associated with the input data may result in a machine-processable representation associated with the one or more elements, in particular tokens. Applying the token embedding to one or more elements may refer to passing the one or more elements through the embedding layer, eg as described within the context of FIG. 9. Hence, token embeddings may specify the one or more elements,in particular tokens in a machine-processable representation. For example, the token embedding may transform a numerical value into a vector. This is advantageous since this representation can be enriched by further information such as the position of the token within the sequence and / or within a table associated with the sequence of tokens. The positional embedding may be analogous to the positional embedding as described within the context of FIG. 9, FIG. 10A-FIG. 10C. Where the input data may be tabular data, column embedding may be applied. Applying a column embedding to one or more elements, in particular tokens associated with the input data may result in a machine-processable representation specifying the location of the one or more elements within a table 402, preferably within the columns of the table 402. Applying the column embedding may refer to adding a column factor to the input data embedded via token embeddings, in particular the embedded input data. The column factor may be the same for elements associated with the same column and / or may differ between two or more elements associated with different columns. Analogous, row embeddings may be applied where the input data may be tabular data. Applying a row embedding to one or more elements, in particular tokens associated with the input data may result in a machine-processable representation specifying the location of the one or more elements within a table 402, preferably within the rows of the table 402. Applying the row embedding may refer to adding a column factor to the input data embedded via token embeddings, in particular the embedded input data. The row factor may be the same for elements associated with the same row and / or may differ between two or more elements associated with different rows.In an embodiment, input data may be at least partially numerical and at least partially text. Hence, the input data may comprise two or more types of data. A type of data may refer to a modality. Followingly, different embeddings may be applied to the input data. To parts of the input data comprising text the input embedding referred to in FIG. 9, FIG. 10A- FIG. 10C may be applied. To parts of the input data being numerical token embeddings, positional embeddings, column embeddings and row embeddings may be applied. Further, segment embeddings may be applied to the input data independent of the type of input data. The segment embedding may specify the type of input data one or more elements may be associated to. For example, if the input data comprises of text and numbers, the input data may comprise of two types of input data. Applying the segment embedding to the input data may refer to adding a segment factor to the input data, preferably the embedded input data and / or the input data after having applied the token embedding. The segment factor may specify the type of data associated with the one or more elements. The segment factor may be the same for one or more elements associated with the same type of input data and / or may differ between two or more elements associated with different types of input data.Applying the token embedding, the positional embedding, the segment embedding, the column embedding, the row embedding or a combination thereof may result in embedded input data and / or may be the output of any one of the encoder input 278, 284, 288 or decoder input 284, 294. The data obtained by applying the token embedding, the positional embedding, the segment embedding, the column embedding, the row embedding or a combination thereof may be processed by the encoder block 274, 286, decoder block 280, 290, encoder output 276, decoder output 292, 282.The following embodiments describe yet further possible ways of implementing the aspects and embodiments of this disclosure for use in an industrial environment of chemical production for climate change mitigation through advanced manufacturing.The training plant-based data 1013, untrained / pre-trained transformer-based model as described in the context of FIGs. 9-12 and / or the trained transformer-based model as described in the context of FIGs. 2a, 2c, 6 and 8, may be stored in a database, on an electronic data carrier or in a cloud. An access to the training plant-based data 1013 suitable for training a pre-trained transformer-based model, an untrained / pre-trained transformer-based model as described in the context of FIGs. 9-12 and / or the trained transformer-based model as described in the context of FIGs. 2a, 2c, 6 and 8, may be granted to a user as a computer readable token. The computer readable token may be an authorization key generated by a computer processor upon a request of a requesting computing node associated with a user. The request may be sent to an authorization engine comprising at least one computer processor having rights to grant one or more authorization keys to access the training plant-based data and / or one or more of the untrained, the pre-trained, the trained, the re-trained or fine-tuned transformer based models as described withing the context of FIGs. 2a, 2c, 6, 8 and 9-12.A user may access the training data via the token for processing the data and receiving the processed result without receiving the actual training data. The user may also receive the actual training plant-based data or a part of the training plant-based data. Depending on the access rights, a user may use (upon receiving the token) one or more of the untrained, pre-trained and / or trained transformer-based models for user’s technical purpose. The user may train / re-train one or more models that the user is granted a permission using the training plant-based data as claimed or using their own training data. An access token for using training plant-based data as claimed may be the same or different as to an access token for using one or more models as described withing the context of FIGs. 2a, 2c, 6, 8 and 9-12. Having a separate access token for each of the data products (i.e., training plant-based data) or data services (i.e., using one or more models as described withing the context of FIGs. 2a, 2c, 6, 8 and 9-12) may increase security aspects of using data products and / or data services. An access token may comprise user credentials, one or more generation algorithms, user authentication, two-factor authentication and / or alike.Example 4 (Transferring of one or more embedding layers for various tasks associated with one or more production operations)Data embedding as described in the context of FIG. 9 may be transferred from one or more trained transformer based models trained on one or more training datasets to another transformer based model that is not yet trained using said one or more training datasets. For example, a first dataset comprising unspecific text based data may be used to train a transformer based model to generate a first pre-trained transformer based model comprising a first embedded data as described in the context of FIGs. 9-12. The first embedded data may be transferred from the first pre-trained transformer based model to another transformer based mode. Furthermore, a first pre-trained transformer based model may be trained on a second dataset comprising one or more database queries in one or more database query languages for generating a first trained transformer based model trained to generate database queries based on a prompt comprising instructions from an operator provided in a natural language. The first trainedtransformer based model may comprise embedded data based on the second training dataset. This embedded data may be transferred to one or more pre-trained or trained transformer based models to generate database queries. A pre-trained or first trained transformer based models may be trained on a first 10131 and / or second training data 10132 generating a second and / or third trained transformer based model trained to analyze data associated with production. The second and / or the third trained transformer based model may comprise embedded data based on a first 10131 and / or second training data 10132. Said embedded data embedded in the second and / or third trained transformer based model may be transferred to a pre-trained or trained transformer based model not yet trained using the a first 10131 and / or second training data 10132.Said transfer of embedded data between transformer based models may accelerate training and improve properties of the trained models to perform one or more tasks associated with one or more production operations of the distributed production environment. For example, an embedded data based on a first training dataset embedded in a first transformer based model may be transferred to a second transformer based model that has better properties than the first transformer based model but wherein the second transformer based model was not yet trained on the first training dataset. Thus, by embedding transfer, the second, better model may acquire embedded data of the first model in an accelerated manner. The use of the second transformer based model may be more advantageous because the initial properties of the second transformer based model (properties before the embedding transfer). In addition to that, the second transformer based model may acquire the embedded data of the first model thereby combining the properties of the first model (the embedded data) and the second model (the initial properties).It may be understood that the features of the above and below embodiments are combinable unless otherwise disclosed in the current application.The following example embodiments shall also be disclosed:Clause 1 :A computer implemented method for climate change mitigation by reducing waste of a product and waste of production resources by detecting one or more anomalies in a production cycle of a distributed chemical production, identifying one or more sources of the one or more anomalies and removing one or more sources of the one or more anomalies, the method comprising: receiving, by a computer processor, access to a trained transformer based model; receiving, via a computer interface, input plant-based data associated with one or more production operations of one or more production lines of the distributed chemical production; prompting the trained transformer-based model to analyze the input plant-based data for detecting one or more anomalies in the input plant based-data, and, when one or more anomalies is present, identifying one or more sources of anomalies and identifying one or more instructions for removing one or more sources of the one or more anomalies.Clause 2:The computer implemented method of clause 1 , further comprising: prompting the trained transformer based model to generate a technical report comprising a description indicating presence of the one or more anomalies or indicating absence of an anomaly, wherein, in case the one or more anomalies is detected, optionally, the technical report further comprising an indication of one or more sources of the one or more anomalies, and optionally, the technical report further comprising one or more operating instructions for removing the one or more sources of the one or more anomalies, and optionally, the operating instructions further comprising: instructions to an operator of the distributed production environment output via the computer interface for controlling and / or monitoring the one or more production operations, and / or machine readable instructions for automatically, by a computer processor, controlling equipment and / or sensors for removing the one or more sources of the one or more anomalies.Clause 3:A computer implemented method for providing the trained transformer-based model for use according to clause 1 or 2, the method comprising: providing, via the computer interface, one or more training data sets; providing, via the computer interface, access to a pre-trained transformer-based model comprising at least a transformer component; and prompting the pre-trained transformer-based model to fine-tune or re-train using the one or more training data sets for generating the trained transformer based model for use according to clause 1 or 2.Clause 4:The computer implemented method of clause 3, wherein the one or more training data sets comprises: a first training dataset comprising one or more plant technical documents, and a second training dataset comprising plant-based data associated with one or more production operations of one or more production lines of the distributed chemical production.Clause 5:The computer implemented method of clause 4, wherein the second training dataset comprises datapoints associated with one or more production operations and associated with standard operating conditions of a production line, wherein the standard operating conditions are defined by a predefined range of one or more parameters associated with the one or more production operation of the production line;Clause 6:The computer implemented method of clause 4 or 5, wherein the second training dataset comprises datapoints associated with the one or more production operations and extreme operating conditions of the production line, wherein the extreme operating conditions are defined by one or more parameters exceeding the predefined range of the one or more parameters associated with the one or more production operation of the production line.Clause 7:The computer implemented method of any one of clauses 1-6, further comprising: receiving, by the trained transformer based model, a one or more database query training datasets comprising one or more database query languages; prompting the trained transformer based model to train using the one or more database query training datasets for retrieving one or more data of any one of clauses 1-5 from a one or more databases upon a prompt of a user comprising instructions for retrieving said one or more data, wherein the instructions comprise natural language.Clause 8:The computer implemented method of clause 7 further comprising: prompting the trained transformer based model of clause 7 to retrieve one or more data according to anyone of clauses 1 -6, identifying, by the trained transformer based model of clause 7, whether access and / or authorization credentials to access one or more databases comprising the one or more data is required, if access and / or authorization credentials are required, generating, by the trained transformer based model of clause 6, a request for a user to provide an authorization and / or access credentials to access the one or more databases, retrieving, by the trained transformer based model of clause 7 the one or more data, and optionally, using the one or more data according to any one of clause 1-6.Clause 9:The computer implemented method of any one of clauses 1-8, wherein one or more data of any one of clauses 1-6 is provided by uploading the one or more data via a user interface.Clause 10:A computer implemented method for detecting a one or more anomalies in a production cycle of the distributed chemical production, the method comprising:when authorization is requested by an artificial intelligence engine, Al comprising at least one computer processor for operating one or more transformer based models, authorizing, by a user, via a computer interface, to access the Al engine, selecting, via a computer interface, a transformer based model from a list of the one or more transformer based models, when authorization and / or access credentials are requested by the Al engine to access the selected transformer based model, providing, via the computer interface, by the user, the authorization and / or access credentials, wherein the selected transformer based model is a pre-trained transformer based model comprising a transformer component and comprising embedded data based on unstructured data comprising at least one or more text data, accessing, via the computer interface, the pre-trained transformer based model, prompting, via the computer interface, the pre-trained transformer based model to access a first training dataset comprising one or more database queries in one or more database query languages, prompting, via the computer interface, the pre-trained transformer based model to train using the first training dataset to generate a first trained transformer based model trained to interpret instructions from an operator of the distributed chemical production comprising natural language into database queries; prompting, via the computer interface, the first trained transformer based model to retrieve a second dataset associated with one or more production operations, wherein the prompt comprises one or more parameters associated with one or more data characteristics characterizing the data of the second dataset, retrieving, by the first trained transformer based model the second dataset, prompting the first trained transformer based model to train using the second dataset to generate the second trained transformer based model comprising an embedded data distribution model suitable to interpret data associated with one or more production operations of the distributed chemical production, prompting, via the computer interface, the second trained transformer based model to retrieve a third dataset associated with one or more production operations, wherein the prompt comprises one or more parameters associated with one or more data characteristics characterizing the data of the third training dataset, and wherein the prompt further comprises instructions for analyzing the third dataset, retrieving, by the second trained transformer based model, the third dataset, analyzing, by the second trained transformer based model, the third dataset according to the instructions for analyzing the third dataset, wherein the analyzing comprises comparing the third dataset to the embedded data distribution model, the analyzing further comprises identifying a deviation of a datapoint of the third dataset from the embedded data distribution model, generating, by the second trained transformer based model, a response to the user, wherein the response comprises instructions having a format of a natural language, and / or the response comprises a technical report characterizing the deviation of the datapoint of the third dataset from the embedded data distributionmodel for removing the one or more sources of the one or more anomalies causing the deviation of the datapoint.Clause 11 :The computer implemented method of clause 10, wherein the first and the second trained transformer based models may be one transformer based model having one embedding layer embedding data from the first and the second training datasets, or wherein the first and the second trained transformer based model may be two transformer based models each having an embedding layer embedding the first or the second training dataset, further comprising transferring one or more embedded data comprising an embedded data based on unstructured data, an embedded data based on the first training dataset and / or an embedded data based on the second dataset to another transformer based model not yet comprising said one or more embedded data, wherein the another transformer based model is selected from the list of the one or more transformer based models, the first or the second transformer based models.Clause 12:Using the another transformer based model of clause 11 instead of any one of the pre-trained or trained transformer based models as specified in any one of clauses 1-11 .Clause 13:The computer implemented method of clause 12, wherein the one or more parameters associated with the one or more data characteristics characterizing the data of the second and / or the third dataset is any one of a time period, a plant identification, a production line identification, a production cycle identification, one or more operating parameters associated with one or more production operation, an identification of a product, a plant, one or more parameters associated with one or more properties of a product, data type such as temperature, pressure, flow, mass.Clause 14:The computer implemented method of any one of clauses 11 -13, further comprising providing, via a user interface, the first training dataset, the second training dataset, the third dataset and / or an additional dataset associated with one or more production operations of the distributed facility.Clause 15:A computer program product comprising computer readable instructions that when executed on a computer cause the computer to carry out steps as specified in any one of clauses 1-14.Clause 16:A method for detecting a one or more anomalies in a production cycle of the distributed chemical production, the method comprising: selecting, via a computer interface, a transformer based model from a list of the one or more transformer based models, wherein the selected transformer based model is a pre-trained transformer based model comprising a transformer component and comprising embedded data based on unstructured data comprising at least one or more text data, accessing, via the computer interface, the pre-trained transformer based model, prompting, via the computer interface, the pre-trained transformer based model to access a first training dataset comprising one or more database queries in one or more database query languages, prompting, via the computer interface, the pre-trained transformer based model to train using the first training dataset to generate a first trained transformer based model trained to interpret instructions from an operator of the distributed chemical production comprising natural language into database queries; prompting, via the computer interface, the first trained transformer based model to retrieve a second dataset associated with one or more production operations, wherein the prompt comprises one or more parameters associated with one or more data characteristics characterizing the data of the second dataset, retrieving, by the first trained transformer based model the second dataset, prompting the first trained transformer based model to train using the second dataset to generate the second trained transformer based model comprising an embedded data distribution model suitable to interpret data associated with one or more production operations of the distributed chemical production, prompting, via the computer interface, the second trained transformer based model to retrieve a third dataset associated with one or more production operations, wherein the prompt comprises one or more parameters associated with one or more data characteristics characterizing the data of the third training dataset, and wherein the prompt further comprises instructions for analyzing the third dataset, retrieving, by the second trained transformer based model, the third dataset, analyzing, by the second trained transformer based model, the third dataset according to the instructions for analyzing the third dataset, wherein the analyzing comprises comparing the third dataset to the embedded data distribution model, the analyzing further comprises identifying a deviation of a datapoint of the third dataset from the embedded data distribution model, generating, by the second trained transformer based model, a response to the user, wherein the response comprises instructions having a format of a natural language, and / or the response comprises a technical report characterizing the deviation of the datapoint of the third dataset from the embedded data distribution model for removing the one or more sources of the one or more anomalies causing the deviation of the datapoint.The publication Prior Art Disclosure; Issue 684; paragraphs

[1000] to

[8005] ; ISSN: 2198-4786; published: February 12, 2024 will be regarded as Reference RF1 , which is incorporated herein by reference in itsentirety. Preferably, the (chemical) product is a product as described in Reference RF1 ; paragraphs

[1000] to

[8005] , Preferably, the method / process described herein is further a method / process for the production of a product.The converting step to obtain the product preferably comprises one or more step(s) as described below and can be performed by conventional methods well known to a person skilled in the art. The converting step preferably comprises one or more step(s) selected from: recycling, preferably depolymerizing, gasifying, pyrolyzing, and / or steam cracking; and / or purifying, preferably crystallizing, (solvent) extracting, distilling, evaporating, hydrotreating, absorbing, adsorbing and / or subjecting to ion exchanger; and / or assembling, preferably foaming, synthesizing, chemical conversion, chemically transforming, polymerizing and / or compounding; and / or forming, preferably foaming, extruding and / or molding; and / or finishing, preferably coating and / or smoothing.In addition, the one or more step(s) are described in detail in Reference RF1 ; paragraphs

[1000] to

[8005] ,The present disclosure has been described in conjunction with preferred embodiments and examples as well. However, other variations can be understood and effected by those persons skilled in the art and practicing the claimed subject-matter, from the studies of the drawings, this disclosure and the claims. Notably, in particular, any steps presented can be performed in any order, i.e. the present disclosure is not limited to a specific order of these steps. Moreover, it is also not required that the different steps are performed at a certain place or at one node of a distributed system, i.e. each of the steps may be performed at different nodes using different equipment / data processing.The sequence of all method steps presented above is not mandatory, also alternative sequences may be possible. Nevertheless, the specific sequence of method steps shown as examples in the figures shall be considered as one possible sequence of method steps, e.g. for the respective embodiment described by the respective figure or an embodiment comprising at least some of the steps described by the respective figure.In the present specification, any presented connection in the described embodiments is to be understood in a way that the involved components are operationally coupled. Thus, the connections can be direct or indirect with any number or combination of intervening elements, and there may be merely a functional relationship between the components.The indefinite article “a” or “an” is not to be understood as “one”, i.e. use of the expression “an element” does not preclude that also further elements are present. A single element or other unit may fulfill the functions of severalentities or items recited in the claims. The mere fact that certain measures are recited in the mutual different dependent claims does not indicate that a combination of these measures cannot be used in an advantageous implementation or further elements may be included.The expressions “A and / or B” and “at least one of: A or B” are considered interchangeable and meant to comprise any one of the following three scenarios: (I) A, (ii) B, (ill) A and B. More generally, the expression “at least one of the following: ” and “at least one of ” and similar wording, where the list of two or more elements are joined by “and” or “or”, mean at least any one of the elements, or at least any two or more of the elements, or at least all the elements.Providing in the scope of this disclosure may include any interface configured to provide data. This may include an application programming interface, a human-machine interface such as a display and / or a software module interface. Providing may include communication of data or submission of data to the interface, in particular display to a user or use of the data by the receiving entity.Obtaining in the scope of this disclosure may include any interface configured to obtain or receive data. This may include an application programming interface, a human-machine interface such as a display and / or a software module interface. Obtaining may include communication of data or submission of data from the interface, in particular use of the data by the receiving entity. Any obtaining of data, data structures, data sets, or the like may comprise receiving the data, data structures, data sets, or the like from a server providing (e.g. hosting) a data base comprising the data, data structures, data sets, or the like.Various units, circuits, entities, nodes or other computing components may be described as “configured to” perform a task or tasks. Configured to shall recite structure meaning “having circuitry that” performs the task or tasks on operation. The units, circuits, entities, nodes or other computing components can be configured to perform the task even when the unit / circuit / component is not operating. The units, circuits, entities, nodes or other computing components that form the structure corresponding to “configured to” may include hardware circuits and / or memory storing program instructions executable to implement the operation. The units, circuits, entities, nodes or other computing components may be described as performing a task or tasks, for convenience in the description. Such descriptions shall be interpreted as including the phrase “configured to.” Any recitation of “configured to” is expressly intended not to invoke 35 U.S.C. § 112(f) interpretation.In general, the methods, apparatuses, systems, computer elements, nodes or other computing components described herein may include memory, software components and hardware components. The memory can include volatile memory such as static or dynamic random-access memory and / or nonvolatile memory such as optical or magnetic disk storage, flash memory, programmable read-only memories, etc. The hardware components mayinclude any combination of combinatorial logic circuitry, clocked storage devices such as flops, registers, latches, etc., finite state machines, memory such as static random-access memory or embedded dynamic random-access memory, custom designed circuitry, programmable logic arrays, etc.In the present specification, any presented connection in the described embodiments is to be understood in a way that the involved components are operationally coupled. Thus, the connections can be direct or indirect with any number or combination of intervening elements, and there may be merely a functional relationship between the components.Moreover, any of the methods, processes and actions described or illustrated herein may be implemented using executable instructions in a general-purpose or special-purpose processor and stored on a computer-readable storage medium (e.g., disk, memory, or the like) to be executed by such a processor. References to a ‘computer- readable storage medium’ should be understood to encompass specialized circuits such as signal processing devices, and other devices.Any disclosure and embodiments described herein relate to the methods, the systems, devices, the computer program element lined out above and vice versa. Advantageously, the benefits provided by any of the embodiments and examples equally apply to all other embodiments and examples and vice versa.All terms and definitions used herein are understood broadly and have their general meaning if not indicated otherwise.It will be understood that all presented embodiments are only examples, and that any feature presented for a particular example embodiment may be used with any aspect on its own or in combination with any feature presented for the same or another particular example embodiment and / or in combination with any other feature not mentioned. In particular, the example embodiments presented in this specification shall also be understood to be disclosed in all possible combinations with each other, as far as it is technically reasonable and the example embodiments are not alternatives with respect to each other. It will further be understood that any feature presented for an example embodiment in a particular category (method / apparatus / computer program / system) may also be used in a corresponding manner in an example embodiment of any other category. It should also be understood that presence of a feature in the presented example embodiments shall not necessarily mean that this feature forms an essential feature and cannot be omitted or substituted.REFERENCE LIST1001 Production line(s);1002 Equipment;1003 Sensors;1004 Analytics engine;1005 Production data / plant data, e.g., sensors data, analytics data;1006 Control and / or monitoring engine; 1007, 1008 Machine readable instructions;1010 Production / Plant historic data;1011 Input data;1012 Operator data;1013 Training data;

Claims

CLAIMS1 . A method for anomaly mitigation in a production cycle of a distributed chemical production comprising:Receiving an analyze instruction related to detecting one or more anomalies in plant based-data, wherein the analyze instruction comprises a reference indicator indicating at least one reference data set and an analysis indicator indicating at least one data set to be analyzed;Obtaining, based on the analyze instruction, the at least one reference data set and the at least one data set to be analyzed;Determining at least one anomaly mitigation instruction, the determining the at least one anomaly detection result comprising: providing a task instruction, based on the analyze instruction, to at least one generative data-driven model, the at least one generative data-driven model having been trained to generate the at least one anomaly mitigation instruction based on the at least one reference data set and the at least one data set to be analyzed in response to receiving the task instruction;Providing the at least one anomaly mitigation instruction.

2. The method of claim 1 , wherein the task instructing comprises at least a part of the at least one reference data set and / or at least a part of the at least one data set to be analyzed.

3. The method of claim 1 or 2, wherein the generative data-driven model is fine-tuned or re-trained using the at least one reference data set and wherein the fine-tuned or re-trained generative data-driven model is used as the generative data-driven model in the determining the at least one anomaly mitigation instruction.

4. The method of any one of claim 1 to 3, wherein obtaining the at least one reference data set and / or the at least one data set to be analyzed comprises determining whether the at least one reference data set and / or the at least one data set to be analyzed are available in a data base;Upon determining that the at least one reference data set and / or the at least one data set to be analyzed is available in the data base: retrieving the at least one reference data set and / or the at least one data set to be analyzed from the data base.

5. The method of claim 4, further comprising:Upon determining that the at least one reference data set is not available in the data base: requesting a user to provide the at least one reference data set;Receiving at least one reference data set via a computer interface.

6. The method of claim 4 or 5, further comprising:Upon determining that the at least one data set to be analyzed is not available in the data base: requesting a user to provide the at least one data set to be analyzed;Receiving at least one data set to be analyzed via a computer interface.

7. The method of any one of claims 1 to 6, further comprising:Operating and / or controlling the production cycle of the distributed chemical production based on the at least one anomaly mitigation instruction, in particular to produce a chemical product.

8. The method of any one of claims 1 to 7, wherein the task instructing comprises an instruction to generate a technical report comprising a description indicating presence of the one or more anomalies or indicating absence of an anomaly, wherein, in case the one or more anomalies is detected, optionally, the technical report further comprising an indication of one or more sources of the one or more anomalies.

9. The method of any one of claims 1 to 8, wherein the generative data-driven model is fine-tuned or re-trained using a technical training dataset comprising one or more plant technical documents. and wherein the fine-tuned or re-trained generative data-driven model is used as the generative data-driven in the determining the at least one anomaly mitigation instruction.

10. The method of any one of claim 9, wherein the technical training dataset comprises datapoints associated with one or more production operations and associated with standard operating conditions of a production line, wherein the standard operating conditions are defined by a predefined range of one or more parameters associated with the one or more production operation of the production line, wherein optionally the one or more parameters is any one of a time period, a plant identification, a production line identification, a production cycle identification, one or more operating parameters associated with one or more production operation, an identification of a product, a plant, one or more parameters associated with one or more properties of a product, data type such as temperature, pressure, flow, mass.11 . The method of any one of claims 1 to 10, further comprising:Determining whether access and / or authorization credentials to access one or more databases comprising the at least one reference data set, the at least one data set to be analyzed, and / or the technical data set of claims 9 or 10 is required,Upon determining that access and / or authorization credentials are required, generating a request for a user to provide an authorization and / or access credentials to access the one or more databases, retrieving the at least one reference data set, the at least one data set to be analyzed, and / or the technicaldata set of claims 9 or 10 and optionally, using the at least one reference data set and / or the technical data set of claims 9 or 10 to fine-tune or re-train the generative data-driven model; and wherein the fine-tuned or re-trained generative data-driven model is used as the generative data-driven model in the determining the at least one anomaly mitigation instruction.

12. The method of any one of claims 1 to 1 1 , further comprising:Providing, via the computer interface, access to the at least one generative data-driven model, wherein the determining the at least one anomaly mitigation instruction is carried out using the accessed at least one generative data-driven model.

13. The method of any one of claims 1 to 12, wherein obtaining the at least one reference data set, the at least one data set to be analyzed, and / or the technical data set of claims 9 or 10 comprises:Generating by the at least one generative data-driven model a retrieval query for obtaining the at least one reference data set, the at least one data set to be analyzed, and / or the technical data set;Providing the generated retrieval query to a database comprising the at least one reference data set, the at least one data set to be analyzed, and / or the technical data set for obtaining the at least one reference data set, the at least one data set to be analyzed, and / or the technical data set;Obtaining, from the database, the at least one reference data set, the at least one data set to be analyzed, and / or the technical data set based in the provided retrieval query.

14. An apparatus comprising respective means for carrying out or performing the steps of any one of claims 1 to 13 or comprising at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to carry out the steps of the method according to any one of claims 1 to 13.

15. Use of an anomaly mitigation instruction generated according to the methods of any one of claims 1 to 13, or by the apparatus of claim 14 for displaying the anomaly mitigation instruction to an operator of the chemical plant and / or for producing a chemical product.