System and computer-implemented method for a distributed production environment
A generative data-driven model, like a transformer-based model, improves the control and safety of chemical production by generating precise operating instructions and adapting to real-time data, addressing the challenges of complex chemical processes in distributed environments.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- BASF SE
- Filing Date
- 2024-07-22
- Publication Date
- 2026-05-07
AI Technical Summary
Distributed production environments, such as chemical plants, face challenges in efficiently and safely controlling complex processes involving hazardous chemicals at high temperatures and pressures, requiring precise timing and compliance with safety regulations to mitigate risks.
A computer-implemented method using a generative data-driven model, such as a transformer-based model, to analyze plant-based input data and generate operating instructions for production processes, with the ability to fine-tune and retrain the model using historical production data and operator inputs for improved efficiency and safety.
Enhances the efficiency and safety of chemical production by providing precise operating instructions, reducing risks associated with hazardous materials, and ensuring compliance with safety regulations through continuous model updates and operator interaction.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical field
[0001] The following disclosure concerns the field of computer-aided production, such as chemical production. The following disclosure may relate to trustworthy AI. State of the art
[0002] A distributed production environment, such as a chemical plant, is highly complex, involving various chemical processes controlled by a multitude of chemical devices. These processes can include the handling, storage, and transformation of various chemicals (including hazardous, flammable, and / or toxic chemicals). These procedures may be carried out at high temperatures and / or high pressures, and safe and efficient control requires the precise timing of multiple devices, for which strict operating procedures may be in place. Compliance with these regulations can help mitigate risks associated with hazardous materials, ensure the health and safety of personnel, and protect the environment. Brief description
[0003] According to a first aspect, a method for controlling and / or monitoring a distributed production environment is disclosed, wherein the distributed production environment comprises at least one device for producing a product. The method includes: - Receiving, via a computer interface, plant-based input data associated with one or more production processes; - Determining operating instructions for one or more production processes of the distributed production environment, wherein determining the operating instructions includes providing a prompt to at least one generative data-driven model trained to generate the operating instructions in response to receiving the prompt; - Providing the operating instructions.
[0004] Further aspects will be disclosed, including respective devices, systems and uses. Designs
[0005] In a distributed production environment, such as a chemical plant, various chemical processes can be controlled by a multitude of chemical devices. These processes may include the handling, storage, and transformation of various chemicals (including hazardous, flammable, and / or toxic chemicals). These processes may, for example, be carried out at high temperatures and / or high pressures, where safe and efficient control requires the precise timing of multiple devices. Compliance with these regulations can help mitigate risks associated with hazardous materials, ensure the health and safety of personnel, and protect the environment.
[0006] For example, WO2020165045 (A1), WO2021116123 (A1), WO2021156157 (A1) show that chemical production can be data-intensive and can provide multi-parameter data streams.
[0007] This disclosure may relate to how the control and / or monitoring of a distributed production environment can be improved.
[0008] The aspects, embodiments, and examples provided in this disclosure make it possible to use a generative, data-driven model (e.g., a transformer-based model) for this purpose. The aspects, embodiments, and examples provided in this disclosure can enable an improvement in the efficiency of chemical production.
[0009] According to a first aspect, a (particularly computer-implemented) method for controlling and / or monitoring a distributed production environment is disclosed, wherein the distributed production environment comprises one or more pieces of equipment (e.g., a chemical device such as a reactor) that produce a product, and wherein the method comprises the following: - Receiving, via a computer interface, plant-based input data associated with one or more production processes; - Determining operating instructions for one or more production processes of the distributed production environment, wherein determining the operating instructions includes providing a prompt to at least one generative data-driven model trained to generate the operating instructions in response to receiving the prompt; - Providing the operating instructions (e.g., to an operator of the plant / distributed production environment or, e.g., after verification by an operator, to a control system, e.g., if the operating instructions are provided at least partially in a machine-readable structure, e.g., machine-readable instructions).
[0010] Providing a prompt to at least one generative data-driven model trained to generate operating instructions in response to receiving the prompt may involve instructing the trained generative data-driven model (e.g., a transformer-based model) to analyze the plant-based input data and provide operating instructions for one or more production operations in the distributed production environment.
[0011] According to a second aspect, a (particularly computer-implemented) method for generating the at least one generative data-driven model for use in a method according to the first aspect is disclosed, the method comprising the following: - Providing, via a computer interface, plant-based training data associated with one or more production processes; - Provide, via the computer interface, a pre-trained generative data-driven model (e.g., a transformer-based model that includes at least one transformer component); - Fine-tuning the pre-trained, generative, data-driven model using the plant-based training data; - Releasing the trained generative data-driven model for one or more production processes of the distributed production environment according to the first aspect.
[0012] According to one embodiment of the first aspect, plant-based raw data is preprocessed by a preprocessing engine to provide the plant-based input data.
[0013] According to one embodiment of the second aspect, plant-based raw data is preprocessed by a preprocessing engine to provide the plant-based training data.
[0014] Raw data can be unprocessed, unfiltered, and unaltered data directly captured or generated by a source, such as a sensor, in a distributed production environment. Raw data may lack structure or context. Because it is captured directly from the source, it may contain errors, noise, or inconsistencies resulting from the data acquisition process itself. It may also contain irrelevant or redundant information.
[0015] According to one embodiment of any aspect, the plant-based input data includes data from an operator, provided to the operator via a computer interface. An operator, for example, of the distributed production environment, can provide specific data points that may be of particular relevance to the distributed production environment. This might occur, for instance, if the operator suspects an anomaly or wishes to restrict the operating instructions provided by the generative data-driven model, e.g., to modify parameter ranges of the production process or only parameters of specific chemical devices / plants, or to exclude the operating instructions from applying to specific parts of the production process. The operator's data may relate, for example, to anomaly data points, typical or optimal operating parameters of the distributed production environment.The operator's data can form part of the prompt for the trained generative data-driven model.
[0016] According to one embodiment of the method described in the first aspect, the prompt (e.g., from an operator or generated from a request by the operator) includes an instruction for the trained at least one generative data-driven model to provide machine-readable instructions. Such machine-readable instructions can be in a format used to control at least one part (e.g., a specific chemical device) of the distributed production environment. An operator can therefore directly use the provided machine-readable instructions to control the (e.g., chemical) distributed production environment, or the system can directly use the machine-readable instructions, e.g., after an operator has been prompted to verify the generated operating instructions.
[0017] According to one embodiment of the method as described in the first aspect, the prompt comprises one or more contexts relating to one or more production operations in the distributed production environment. Providing one or more production operations in the distributed production environment as a specific context can improve the quality and accuracy of the generated operating instructions and can therefore further enhance the efficient, safe, and reliable operation of the distributed production environment.
[0018] According to one embodiment of the method described in the first aspect, the operating instructions for production are available in one or more languages. This can, for example, enable operators with different native languages to quickly understand the operating instructions and act accordingly, and furthermore facilitate the transfer of operational knowledge in a distributed production environment between plants or production lines located in different countries.
[0019] According to one embodiment of the method described in the second aspect, the plant-based training data is plant-based data in one or more languages. The availability of training data in different languages can improve the ability of generative data-driven models to understand different input languages and generate operating instructions in different languages.
[0020] According to one embodiment of the method as described in the second aspect, the method further comprises updating the plant-based training data and fine-tuning or retraining the trained generative data-driven model based on the updated plant-based training data. This can enable continuous updating of the generative data-driven model with new production data and an improvement of the generated operating instructions.
[0021] According to one embodiment of the method according to the second aspect, retraining or fine-tuning the trained generative data-driven model is performed by a schedule-based retraining or fine-tuning, a continuous retraining or fine-tuning, or a trigger-based retraining or fine-tuning.
[0022] According to one embodiment of the method according to the second aspect, the trigger (e.g. for trigger-based fine-tuning) is based on a threshold and an evaluation relating to the operating instructions provided by the trained generative data-driven model, and / or optionally on a number of retraining or fine-tuning cycles.
[0023] According to one embodiment of any aspect, the operating instructions include machine-readable instructions for controlling and / or monitoring production.
[0024] According to another aspect, a computer program product is disclosed which includes computer-readable instructions which, when executed on a computer, cause the computer to perform the steps of any example of the procedure according to the first or second aspect.
[0025] According to a further aspect, a device is disclosed which includes respective means for carrying out or performing the steps of any example of the method according to the first or second aspect, or includes at least one processor and at least one memory which stores instructions which, when executed by the at least one processor, cause the device to perform at least the steps of the method according to any example of the method according to the first or second aspect.
[0026] According to a further aspect, a use of operating instructions generated according to the methods of the first aspect is disclosed for displaying the operating instructions for an operator of the distributed production environment and / or for producing a (e.g. chemical) product.
[0027] According to another exemplary aspect, a computer element is disclosed, wherein the computer element comprises instructions which, when executed by a processor or a computing device, perform or carry out the steps according to the methods or as defined by the devices disclosed herein.
[0028] According to another exemplary aspect, a computer program or computer program product is disclosed, wherein, when executed by a processor, the computer program or computer program product causes a device, for example a server, to perform and / or control the actions of the method according to any aspect.
[0029] According to another exemplary aspect, a (e.g. tangible and / or non-volatile) computer-readable storage medium is disclosed, wherein the computer-readable storage medium comprises a computer program, wherein the computer program, when executed by a processor, causes a device, e.g. a server, to perform and / or control the actions of the method according to any aspect.
[0030] According to another exemplary aspect, a computer product is disclosed that includes a computer-readable token for accessing plant-based training data and / or the trained or pre-trained or fine-tuned model of any aspect.
[0031] A generative, data-driven model (e.g., a transformer-based model) can be trained using "big data," i.e., large datasets (e.g., non-specific text and image data). Trained generative data-driven models, such as transformer-based models, can exhibit improved ability to predict data patterns, such as patterns in natural language. This improved ability can be attributed to the large number of parameters obtained through training. For example, transformer-based models, such as OpenAL GPT, can include 117 million parameters (GPT-1), 1.5 billion parameters, 175 billion parameters (GPT-3), or 170 trillion parameters (GPT-4). These parameters can enable GPT models to produce improved data output compared to other models, such as recurrent neural networks (RNNs) or long short-term memory (LSTM) networks, which do not include a transformer component.
[0032] “Attention Is All You Need” by Vaswani et al., 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA (6 Dec 2017, arXiv:1706.03762v5) can describe a mechanism in machine learning that includes a transformer component (a transformer-based model) which is hereby included by reference.
[0033] ISO / IEC 23053:2022(en), ISO / IEC TR 24372:2021(en), ISO / IEC 22989, and ISO / IEC 23053 can define standards in the field of artificial intelligence (AI) and machine learning (ML). Big Data can be specified in ISO / IEC 20546:2019(en) – Information technology – Big Data. Data quality can be specified, for example, in ISO / IEC 20546:2019(en), ISO 8000-66:2021(en) – Data quality, and ISO / IEC DIS 5259-1(en).
[0034] Generative data-driven models, such as transformer-based architectures, can enable the capture of long-term dependencies and the parallelization of computations. Furthermore, transformer-based architectures can be pre-trained using larger text-based datasets and then fine-tuned (or retrained) for specific tasks using smaller (annotated) datasets. Fine-tuning can be a process in which a pre-trained generative, data-driven model, for example, trained with a large dataset, is further trained on a smaller, specific dataset, thereby allowing the knowledge learned by the pre-trained model to be applied to the specific task. During fine-tuning, the model's weights can be updated based on the provided specific dataset, using the pre-trained weights as a starting point and, for example,Only a small number of additional training steps need to be performed.
[0035] Fine-tuning or retraining can enable the use of the excellent analytical capabilities of pre-trained transformer-based models to analyze data patterns other than patterns in natural languages.
[0036] The present disclosure may relate to the use of generative transformer-based models for analyzing data patterns using a pre-trained transformer-based model for distributed production, such as chemical production. Data suitable for retraining or fine-tuning a pre-trained transformer-based model for production can be obtained by providing historical production data accumulated over more than 150 years.
[0037] The amount of training data used to fine-tune a pre-trained transformer-based model can depend on the specific task, the complexity of the model, and the desired level of performance. In many cases, transformer models can be fine-tuned (retrained) with smaller amounts of purpose-specific training data compared to the size of the pre-training dataset. However, if the specific application is very different from the pre-training data, fine-tuning or retraining a pre-trained transformer-based model may require larger datasets compared to a scenario where the specific application is similar to the pre-training data. For example, text classification or sentiment analysis might require anywhere from a few hundred megabytes to several gigabytes of annotated data.Machine translation may require tens to hundreds of gigabytes of text data for retraining. Question answering can require several gigabytes of training data to retrain and fine-tune a pre-trained transformer-based model.
[0038] With improved training data quality, the amount of training data required can be reduced (the definition of "data quality" is given, for example, in ISO / IEC 20546:2019(en), ISO 8000-66:2021 (en) / Data quality, ISO / IEC DIS 5259-1 (en)). Preprocessing data to generate high-quality training data can include annotating data, removing noise, irrelevant information, and / or similar actions. Therefore, preprocessed data can be beneficial for generating high-quality training data, enabling more efficient retraining or fine-tuning.
[0039] Larger amounts of training data can enable fine-tuning or retraining of pre-trained transformer-based models with more parameters. For example, fine-tuning the DistillBERT model may require less training data compared to the amount needed to fine-tune the GPT-4 model. A larger set of these parameters can also provide improved data analysis capabilities (e.g., GPT-4 is more powerful than GPT-3 at analyzing data patterns).
[0040] A generative, data-driven model can be a transformer-based model, such as TinyBERT, DistilBERT, Llama 7B, Mistral 7B, GPT-Neo, or a larger GPT variant, or another model, such as a structured state-space model. Furthermore, pre-trained transformer-based models can include, for example, ChatGPT (GPT-2, GPT-3, GPT-3-5-turbo, GPT-4), Davinci, BERT (Bidirectional Encoder Representations from Transformers), DistilBERT, Transformer-XL, XLNet (eXtreme Language Understanding Network), T5 (Text-to-Text Transfer Transformer), RoBERTa (Robustly Optimized BERT Approach), ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacements Accurately), Reformer, Longformer, or DeBERTa (Decoding-Enhanced BERT with Disentangled Attention). The properties and therefore the output data of transformer-based models can vary for the same input data due to differences in the architectures and / or pre-training datasets.Thus, one or more of the transformer-based models (GPT-2, GPT-3, GPT-3-5-Turbo, GPT-4, Davinci, BERT, DistilBERT, Transformer-XL, XLNet, T5, RoBERTa, ELECTRA, Reformer, Longformer, DeBERTa) can be used as an alternative to a pre-trained transformer-based model for a technical purpose, or one or more of the models can be used in combination for a technical purpose to provide multiple data outputs for complementary data analysis.
[0041] In particular, providing training data for retraining or fine-tuning a pretrained generative data-driven model for production, such as chemical production, may involve providing historical plant-based data. For example, stored production data (plant-based data) spanning more than 150 years may be used. Providing plant-based training data required for retraining or fine-tuning a pretrained transformer-based model may involve providing historical production data (plant-based data) from one or more databases of one or more distributed production plants.
[0042] Furthermore, the procedure according to the first aspect can include retraining or fine-tuning a pre-trained transformer-based model to provide (release) a trained transformer-based model suitable for production. The trained transformer-based model can then be used to control and / or monitor a distributed production environment, such as chemical production.
[0043] Furthermore, the procedure according to the first aspect can also include providing a transformer-based model and pretraining the model with generic data, such as text-based data (any data types that can be converted into text data). The pretraining data can also include publicly available, production-associated data, such as data available on production facility websites.
[0044] Another large language model (e.g., a foundation model, such as a transformer-based model) that is pre-trained on a large dataset (e.g., a non-specific, publicly available dataset) and then trained (fine-tuned, retrained) on a smaller dataset (e.g., a specific dataset, such as production data) to make the trained model suitable for controlling and / or monitoring production can also be used.
[0045] A pre-trained transformer-based model, pre-trained on larger datasets (publicly available data, non-specific data) and fine-tuned or retrained on smaller datasets (i.e., plant-based datasets), is superior in analyzing plant-based data patterns, identifying anomalies in the data patterns, and generating operating instructions based on the analysis of the data patterns to improve the control and / or monitoring of distributed production, such as chemical, pharmaceutical, or biotechnological production.
[0046] The provision of plant-based training data enables the fine-tuning or retraining of a pre-trained transformer-based model for production.
[0047] Preprocessing raw plant data (e.g., sorting, filtering, annotating, structuring data) to generate plant-based training data improves the quality of training data and thus training efficiency by requiring less training data and computing power to retrain or fine-tune a pre-trained transformer-based model.
[0048] One or more alternative pre-trained, transformer-based models may be used instead of or in addition to the claimed pre-trained, transformer-based model, provided that the one or more models can be fine-tuned or retrained for production. The one or more alternative or additional models may exhibit different properties based on different sets of parameters obtained during pre-training. Thus, the one or more additional or alternative models may be a complementary or alternative component of the various aspects of the disclosure provided herein for analyzing data patterns in plant-based input data.
[0049] Trigger-based retraining or fine-tuning of a trained, transformer-based model, as well as schedule-based or continuous retraining or fine-tuning, can improve model performance because the model can be trained with the most up-to-date data. A feedback rating can be provided as a quality indicator for the generated output data.
[0050] Using the trained transformer-based model to control and / or monitor production, such as chemical production, can improve overall production efficiency. This is because the model can identify many factors contributing to inefficient production by analyzing data patterns in input data based on plant information. Based on this analysis, the model can generate a solution for improving production in the form of operating instructions (machine-readable instructions) for controlling and / or monitoring production. Validating these operating instructions can ensure the secure integration of the model into a distributed production environment like chemical manufacturing. Description of the drawings
[0051] The present disclosure is described in more detail below with reference to the accompanying figures. Identical reference numerals in the drawings and in this disclosure refer to identical or similar elements, components and / or parts. Fig. Figure 1 illustrates a distributed production environment, such as one or more chemical plants. Fig. 2A illustrates an operating system of the distributed production environment of Fig. 1. Fig. Figure 2B illustrates the training of a pre-trained transformer-based model using plant data generated by the production environment of Fig. 1 were produced. Fig. Figure 2C illustrates one embodiment of using a trained transformer-based model to control and / or monitor the in Fig. 1. Distributed production environment shown. Fig. Figure 3 illustrates a preprocessing engine of the in Fig. 2 operating systems shown. Fig. Figure 4 illustrates a plant data structure defined by the information in Fig. The production environment shown in section 1 was created. Fig. Figure 5 illustrates the data structure of an operator. Fig. Figure 6 illustrates further details of retraining or fine-tuning a pretrained transformer-based model for production, in addition to the embodiments described in Fig. 2a and Fig. 2b. Fig. Figure 7 illustrates the contextualization of prompts. Fig. Figure 8 illustrates trigger-based retraining or fine-tuning of a trained transformer-based model when the model is used for production. Fig. Figure 9 illustrates one embodiment for training an embedding layer. Fig. Figure 10A illustrates an embodiment of a transformer-encoder architecture. Fig. Figure 10B illustrates an embodiment of a transformer-decoder architecture. Fig. Figure 10C illustrates an embodiment of a transformer-encoder-decoder architecture. Fig. Figure 11 illustrates one embodiment of training and / or deploying the transformer encoder, the transformer decoder and / or the transformer encoder decoder. Fig. Figure 12 illustrates one embodiment of input embedding. Detailed description
[0052] The following embodiments are merely examples of implementing the method, system, device, or application device disclosed herein and are not intended to be limiting. The following description serves to deepen understanding and is to be understood as a supplement to the preceding summary and the exemplary embodiments of this patent specification, and should be read together with them. Some aspects may use different terminology than, for example, the description above. However, the person skilled in the art will understand that these terms refer to the same subject matter, for example, by being more specific.
[0053] A generative data-driven model (generative artificial intelligence, AI model) in the context of the current application can refer to a foundation model (a machine learning, ML, model) that includes, for example, a transformer component, as in Fig. Described in 9-12.
[0054] Generative Artificial Intelligence (AI) can refer to a computer program capable of producing output, as described, for example, in ISO / IEC 23053:2022(en), ISO / IEC TR 24372:2021(en), ISO / IEC 22989, ISO / IEC 23053, ISO / IEC DIS 5259-1(en), and ISO / IEC 24661:2023(en). A generative AI program may include a machine learning (ML) model, such as a transformer-based model (generative pretrained transformer, GPT model). The generative pretrained transformer model (or simply the transformer-based model) may also be referred to as the foundation model.
[0055] An “engine” in the context of Fig. 1-8 comprises at least one computer processor. A “computer interface” in the context of this disclosure may be, for example, a graphical user interface, an application programming interface, or a web-based interface.
[0056] For example, a pre-trained transformer-based model can be pre-trained for a first purpose (e.g., analyzing non-specific text-based data) and retrained / fine-tuned for a second purpose (e.g., production). A trained generative data-driven model (e.g., a transformer-based model) suitable for the second purpose can be further retrained / fine-tuned for that purpose to improve the output data provided by the model.
[0057] "Distributed production environment" or "facility(s)" can refer, without limitation, to any technical infrastructure used for the industrial purpose of manufacturing, producing, or processing one or more process products; that is, a manufacturing or production process or processing carried out by the distributed production environment. The distributed production environment can be a "facility" (infrastructure) with distributed units for production. The distributed production environment can be a technical infrastructure (facility) that encompasses distributed operations. The distributed production environment can consist of more than one facility distributed geographically and / or directed toward distributed operations.A distributed production environment can consist of one or more facilities, such as a chemical plant, a process plant, a pharmaceutical plant, a fossil fuel processing plant (e.g., an oil and / or natural gas well), a refinery, a petrochemical plant, a cracking plant, and the like. It can even include a distillery, a processing plant, or a recycling plant. A distributed production environment can also be a combination of any of the above examples or similar.
[0058] The “product” generated by the distributed production environment can be any physical product, such as a chemical, biological, or pharmaceutical product; a food product; a dietary supplement; a beverage; a textile; a metal, plastic, or semiconductor product; a cosmetic product; or any combination thereof. Additionally or alternatively, the product can be a service product, such as a recovery or waste treatment, like recycling, chemical treatment, or the extraction or dissolution of materials into one or more chemical products. Some non-restrictive examples of chemical products include organic or inorganic compounds, monomers, polymers, foams, pesticides, herbicides, fertilizers, animal feed, nutritional products, precursors, pharmaceuticals, or treatment products, as well as one or more components or active ingredients thereof.In some cases, the chemical product may be a product usable by an end user or consumer, for example, a cosmetic or pharmaceutical composition. The chemical product may also be a product that can be used to manufacture one or more other products; for example, the chemical product may be a synthetic foam that can be used to manufacture shoe soles or a coating that can be used for vehicle exteriors. The chemical product may be in any form, for example, solid, semi-solid, paste, liquid, emulsion, solution, pellets, granules, or powder.
[0059] The distributed production environment may include equipment (e.g., a chemical apparatus) or process units, such as one or more of the following: a heat exchanger, a column (e.g., a fractionating column), a furnace, a reaction chamber, a cracking unit, a storage tank, an extruder, a pelletizer, a precipitation apparatus, a blender, a mixer, a cutter, a hardening tube, an evaporator, a filter, a sieve, a pipeline, a chimney, a valve, an actuator, a mill, a transformer, a conveying system, a circuit breaker, a machine (e.g., a heavy-duty rotating machine such as a turbine), a generator, a pulverizer, a compressor, an industrial fan, a pump, a transport element such as a conveyor system, a motor, etc.
[0060] Furthermore, a distributed production environment typically includes multiple sensors and at least one control system for controlling at least one parameter related to the process or process parameters within the plant. Such control functions are usually performed by the control system or control device in response to at least one measurement signal from at least one of the sensors. The plant's control device or control system may be implemented as a distributed control system (DCS) and / or a programmable logic controller (PLC). The multiple sensors may be distributed throughout the distributed production environment for monitoring and / or control purposes. Such sensors can generate large amounts of data. The sensors may or may not be considered part of the equipment. Therefore, production, such as chemical and / or service production, can be a data-intensive environment.A distributed production environment can generate a large amount of process-related data.
[0061] These sensors can be used to measure one or more process parameters and / or to measure the operating states of the plant or parameters related to the plant or process units. For example, the sensors can be used to measure a process parameter such as a flow rate within a pipeline, a fill level within a tank, a furnace temperature, the chemical composition of a gas, etc. Some sensors can be used to measure the vibration of a pulverizer, the speed of a fan, the opening of a valve, corrosion of a pipeline, the voltage across a transformer, etc. The difference between these sensors lies not only in the parameter they detect, but also in the sensor principle that each sensor employs.Examples of sensors based on the parameter they detect include: temperature sensors, pressure sensors, radiation sensors (such as light sensors), flow sensors, vibration sensors, displacement sensors, and chemical sensors (such as those used to detect a specific substance, like a gas). Examples of sensors that differ in terms of the sensor principle used include: piezoelectric sensors, piezoresistive sensors, thermocouples, impedance sensors (such as capacitive and resistive sensors), and so on.
[0062] A distributed production environment can consist of multiple distributed production environments. These multiple distributed production environments can be interconnected in such a way that they can share one or more of their value chains, feedstocks, and / or products. The multiple distributed production environments can also be referred to as a compound, compound site, or compound site. Such compound sites or chemical parks can be or comprise one or more distributed production environments, with products manufactured in one or more distributed production environments serving as feedstock for another.
[0063] "Production" refers to any industrial process that, when used on or applied to an input component, provides an output product that is different from the input component. Manufacturing can thus be any production or treatment process, or a combination of several processes, used to obtain the product as defined above. The production process may also include the packaging and / or stacking of one or more of the products.
[0064] The production process can be continuous or run in campaigns; for example, it might be a batch chemical production process, such as one based on catalysts that require recovery. A key difference between these production types lies in the frequencies of the data generated during production. For instance, in a batch process, production data spans from the start of the process to the final batch, encompassing various batches produced in that run. In a continuous design, the data is more continuous, with potential shifts in production operation and / or maintenance-related downtime. Consequently, the required data analysis may differ depending on whether the data flow is batch or continuous.For example, schedule-based retraining or fine-tuning of a trained generative data-driven model (e.g., a transformer-based model) may be advantageous for batch data flows, whereas continuous retraining or fine-tuning of a trained generative data-driven model (e.g., a transformer-based model) may be advantageous for continuous data flows.
[0065] The terms "equipment data" or "equipment-based data" can be used interchangeably and can refer to production data (e.g., product characteristics), process data (process parameters), and operating conditions. Equipment-based data can refer to data that includes values, such as numerical or binary signal values, measured during the production process, for example, via one or more sensors. Process data can be time-series data of one or more of the process parameters and / or the operating conditions of the equipment. Typically, equipment-based data can include temporal information of the process parameters and / or the operating conditions of the equipment; for example, the data contains timestamps for at least some of the data points relating to the process parameters and / or the operating conditions of the equipment. Equipment-based data can include space-time data, i.e.,temporal data and the location or data relating to one or more equipment zones that are physically separated from each other, so that a time-space relationship can be derived from the data.
[0066] “Process parameters” can refer to any of the production process-related variables, for example, any one or more of temperature, pressure, time, value, etc., that are relevant to the manufacture of the product, as defined above.
[0067] The above definitions of a distributed production environment, products generated by the distributed production environment, production processes, data generated by the production environment, and production control are merely examples and should not be considered limiting. It is understood that the system and method of the invention disclosed herein are applicable to any type of production that generates a product and to the generation of multi-parameter data flows relating to production. Any type of plant-based data can be preprocessed by a preprocessing unit to generate the necessary plant-based training data (annotated data, (pre)structured data, filtered data, data in numerical and / or text format, etc.) or plant-based input data for retraining or fine-tuning a pretrained generative data-driven model (e.g.,a transformer-based model) or using a trained generative data-driven model for production are suitable. Therefore, the distributed production environment should generally be designed as a technical environment that produces a product (a physical product and / or a service associated with a product; a product can also be a data product) and generates production-related multi-parameter data flows (plant-based data) during the production of the product through the technical environment.
[0068] Fig. Figure 1 illustrates a distributed production environment, such as one or more chemical plants.
[0069] A distributed production environment can include equipment 1002 and sensors 1003 that generate one or more sensor-related data flows. The distributed production environment can produce one or more products as defined above, where properties of one or more products can be measured, extracted, or calculated, thereby generating one or more product-related data flows. Plant data 1005 (plant-based data) can include data obtained from each of the one or more data flows.
[0070] The equipment 1002 can be any equipment in a distributed production environment, such as pumps, heat exchangers, valves, reaction vessels, separation chambers and / or the like.
[0071] The 1003 sensors can be any type of sensor in a distributed production environment, such as temperature sensor, flow sensor, pressure sensor and / or the like.
[0072] One or more products manufactured through distributed production environments can be of any type, as described above. The properties of the products can be measured, for example, using gas chromatography.
[0073] The plant data 1005 can be stored in a database, for example, as historical data. Plant data 1005 can be provided to the preprocessing engine, which preprocesses the data and provides plant-based input data to an analysis engine 1004. This analysis engine can analyze the plant-based input data and generate machine-readable instructions 1007 for the control and / or monitoring engine 1006. The generation of the machine-readable instructions 1007 can be performed automatically (i.e., without operator intervention) by the analysis engine 1004 based on the analysis of the input plant data. For example, the analysis engine 1004 can continuously receive input plant data and analyze the data in a continuous mode. If an anomaly occurs, the analysis engine can generate machine-readable instructions 1007 for the control and / or monitoring engine to correct the anomaly.The analysis engine can identify solutions for improving production efficiency by analyzing plant-based input data against the backdrop of production processes and send machine-readable instructions 1007 to the control and / or monitoring engine to improve production. The control and / or monitoring engine 1004 can display a push notification to an operator to review the instructions 1007 and, based on this review, generate machine-readable instructions 1008. Based on these machine-readable instructions 1008, the control system can modify the operating parameters of one or more pieces of equipment 1002. Reviewing the machine-readable instructions enables the secure integration of the trained transformer-based model into distributed production environments, such as chemical production.
[0074] Alternatively or in addition to the automatic generation of machine-readable instructions 1007 by the analysis engine, an operator can request the analysis engine 1004 to provide machine-readable instructions 1007 based on a prompt, such as in connection with Fig. 7 described.
[0075] Machine-readable instructions 1007 can be used by an operator to control and / or monitor one or more production operations in the distributed production environment.
[0076] Machine-readable instructions 1007 can include operating instructions for production, such as machine-readable instructions for controlling the equipment 1002 and / or machine-readable instructions for monitoring the equipment 1002 and / or the sensors 1003. Machine-readable instructions 1007 can include operating instructions for an operator to control and / or monitor the distributed production environment, where the operator can be a human operator, a computer-based operating system, or a hybrid system comprising a human operator and a computer-based assistance system.
[0077] An operator can review the machine-readable instructions 1007 and, based on the review, generate machine-readable instructions 1008 to control the equipment 1002 and / or the sensors 1003. An operator can request the analysis engine 1004 to provide operating instructions (machine-readable instructions 1007) based on a prompt, query, context, and / or the like, as described in Fig. 7 illustrates.
[0078] An operator can be a human operator who checks the machine-readable instructions 1007. An operator can be a human operator who has a computer-based supported system for checking the machine-readable instructions 1007. An operator who checks the machine-readable instructions 1007 can be a computer-based system based on a computer program that includes a set of instructions for checking machine-readable instructions 1007.
[0079] The control system of a distributed production environment can comprise one or more computing units capable of modifying one or more process parameters related to the production process by controlling one or more of the actuators or switches and / or end-effector units, for example, by changing one or more of the equipment's operating conditions. This control typically occurs in response to one or more signals received from the equipment.
[0080] The control and monitoring engine can include one or more computer processors for checking machine-readable instructions 1007 and generating machine-readable instructions 1008 for the distributed production control system. Based on the machine-readable instructions 1008, the distributed production control system can adjust the equipment's operating conditions so that the modified process parameters and / or equipment operating conditions result in a controlled product (such as a chemical product) that exhibits one or more required or predetermined properties or performance parameters. This allows for production to be controlled during operation, ensuring that the equipment's operating conditions are adjusted to accommodate undesirable changes in process parameters.
[0081] It is understood that the control and monitoring of the distributed production environment generally involves controlling equipment and / or production lines to manufacture a product by sending machine-readable instructions to the production environment. A product should be broadly defined as described above. As another example, a product could even be a data product and / or a data service product provided by the trained transformer-based model, which was trained using plant-based training data.
[0082] The generative data-driven model or transformer-based model (first ML model), which is generated by the KL engine as in connection with Fig. The first ML model, as described in sections 2a, 2b, and 9-12, can be integrated with another computer program, such as a second ML model, via a computer interface (e.g., an API, a GUI, or a web application). The second ML model may employ a different algorithm (e.g., a classical ML model not based on a transformer architecture, a variation of the transformer-based architecture of the first model, or similar). Alternatively, the second ML model may have the same architecture as the first model, but it may be trained on a different dataset.
[0083] For example, a second machine learning (ML) model could be a data-driven model that integrates with the first model (the transformer-based model) via an API, GUI, or web-based interface. This second ML model can be used to preprocess plant-based raw data to generate plant-based training data. Raw data preprocessing might include noise removal, data filtering, data annotation, data sorting, raw data conversion to a different format, operator conversion to a format better suited for the first ML model (the transformer-based model), and / or similar actions. Alternatively, raw data preprocessing could be performed by a computer program based on a set of computer-implemented instructions (e.g., a set of programming instructions, filters, annotations, mathematical steps, and / or similar actions that do not constitute an ML model).Preprocessing plant-based raw data can improve the quality of training or input data based on the raw data and can reduce the computing power required to train / use the first ML model (e.g., transformer-based model).
[0084] Fig. 2A illustrates an operating system of the distributed production environment of Fig. 1. The operating system includes the (re)training of a pre-trained transformer-based model and the use of a trained transformer-based model to control and / or monitor production, such as chemical production. The training of a pre-trained transformer-based model is further related to… Fig. 2B described. The use of the purpose-trained transformer-based model is explained in more detail in connection with Fig. 2C described.
[0085] Fig. Figure 2B illustrates the training of a pre-trained transformer-based model using plant data generated by the production environment of Fig. 1 were produced.
[0086] Training (fine-tuning, retraining) a pre-trained transformer-based model, such as in the context of Fig. As described in 9-12, this involves accessing a pre-trained transformer-based model via a computer interface (GUI, API, web-based interface) and receiving, via the interface, historical plant data 1010 from the preprocessing engine 1009. An operator can retrain or fine-tune the pre-trained transformer-based model, request the pre-trained transformer-based model to access the historical plant data via the computer interface, or alternatively, the operator can upload the historical plant-based data via the computer interface from a database to the AI engine, which includes at least one processor used to operate the pre-trained transformer-based model.The operator who retrains or fine-tunes the pre-trained model can be a human operator, an automated operating system with a computer processor, or a hybrid operating system comprising a human operator and a computer-based operating system that prompts the pre-trained model to train, fine-tune, or retrain using plant-based data.
[0087] Historical plant data 1010 can be based on plant data 1005. Plant data 1005 and historical plant data can be stored in a database. Plant data 1005 can include several production data points, such as those related to... Fig. 4 described. During training, historical plant data can be embedded via an embedding layer, as described in connection with Fig. 9 described. Embedding the plant data can lead to embedded plant data.
[0088] The above examples of plant data are merely examples to illustrate possible implementations of the method disclosed herein. Plant data should be interpreted generally and understood as any type of data associated with the manufacture of a product by a technical infrastructure. Any type or format of production-associated data can be preprocessed by the preprocessing engine to make the data suitable for training or use with the transformer-based model in production.
[0089] At the end of a training cycle, the 1004 analysis engine can output (release) a trained transformer-based model suitable for use in a distributed production environment, such as chemical manufacturing. The released trained model can be stored in a database for purposes such as version control. The released trained model can also be a computer program product. Access to the released trained model can be provided to a user as a data service to support production.
[0090] Training plant-based data can involve plant-based data in one or more languages. Training pre-trained transformer-based models to train plant-based data in one or more languages allows for the expansion of the plant-based training dataset and additionally enables the provision of operating instructions for production in one or more languages. Providing instructions in one or more languages can improve user interaction with the trained transformer-based model.
[0091] Fig. Figure 2C illustrates one embodiment of using a trained transformer-based model to control the in Fig. 1. Distributed production environment shown.
[0092] The plant data 1005 can be provided to the preprocessing engine 1004 for preprocessing. Preprocessing the data 1005 can include steps related, for example, to... Fig. 3 are described. Preprocessing may also include other steps, such as annotating data, removing noise, structuring unstructured data, converting data into different formats, and / or the like, which are necessary to provide training / input data based on plant data that is suitable for training or using a generative data-driven model (e.g., a transformer-based model) for production.
[0093] The preprocessed data from the preprocessing engine 1004 can be used as plant-based input data for the trained generative data-driven model for production.
[0094] The trained generative data-driven model can receive plant-based input data and, for example, predict anomalies in that data. Based on this analysis, the trained generative data-driven model can, for example, predict how to resolve plant operation errors, improve production efficiency, and handle user queries related to production control and / or monitoring, among other things. Based on these predictions, the analysis engine (1004) can generate machine-readable instructions (1007) for the control and / or monitoring engine (1006).
[0095] The analysis engine 1004, which operates the pretrained transformer-based model, may include at least one computer interface (e.g., graphical user interface, GUI, web-based computer interface, and / or application programming interface, API) for operating the transformer-based model (uploading data, providing prompts, such as prompts that provide instructions for accessing plant-based training data and / or plant-based input data, providing prompts for retraining or fine-tuning, providing prompts for generating machine-readable instructions 1007, means for verifying machine-readable instructions 1007, means for receiving notifications when new machine-readable instructions are generated, and the like). The model (untrained, pretrained, trained) may be stored in a database or in the cloud. The model may be accessed by the AI engine via the at least one computer interface (e.g., graphical user interface, GUI, web-based computer interface, and / or application programming interface, API).B. graphical user interface, GUI, a web-based computer interface and / or application programming interface, API).
[0096] A generative data-driven model, such as a transformer-based model (e.g., any pre-trained, trained, and / or newly trained generative data-driven model), can be provided to the AI engine via a computer interface. In other words, one or more generative data-driven models (e.g., transformer-based models) can be accessible to the AI engine. For example, the AI engine (e.g., at least one computer processor) can access a transformer-based model via the computer interface (e.g., API, user interface, web interface). Alternatively, a transformer-based model can be part of the AI engine (e.g., stored in the AI engine's computer memory), or the transformer-based model can be integrated into the AI engine (e.g., stored in a database that the AI engine can access).
[0097] An operator of the trained transformer-based model can provide input (e.g., plant-based data input) to the trained generative data-driven model (e.g., trained transformer-based model) via a computer interface (e.g., a graphical user interface, an application programming interface, a web-based interface), such as a text and / or audio query, as in the context of Fig. As described in section 5, the operator of the trained generative data-driven model (e.g., a trained transformer-based model) can additionally provide extracted data points, which are extracted, for example, from the data provided by the preprocessing engine. These extracted data points can relate to anomaly data points, typical operating parameters, or optimal operating parameters of the distributed production environment. The extracted data points, together with the operator query, can form part of the input data for the trained generative data-driven model.
[0098] The trained generative data-driven model (e.g., trained transformer-based model) can process the input data and provide a solution for the user query, for example, in the form of machine-readable instructions 1007.
[0099] The control and / or monitoring engine can verify the machine-readable instructions 1007.
[0100] After verification, the control and / or monitoring engine 1006 can generate machine-readable instructions 1008, which may be identical to, partially based on, or different from machine-readable instructions 1007. An operator of the control and / or monitoring engine may use machine-readable instructions 1007 solely for monitoring production or forward them as machine-readable instructions 1008 to control production. The operator may be a human operator and / or an operating system comprising a processor and, optionally, a human operator.
[0101] If machine-readable instructions 1008 generated by the control and / or monitoring engine differ from machine-readable instructions 1007 generated by the trained transformer-based model, the control and / or monitoring engine can trigger retraining or fine-tuning of the trained generative data-driven model (e.g., the trained transformer-based model). The retraining or fine-tuning can be triggered automatically based on feedback from the control and / or monitoring engine, which may indicate that machine-readable instructions 1007 and 1008 differ. The feedback may include a feedback score indicating the degree of deviation.
[0102] A retraining or fine-tuning process can be initiated by an operator. Alternatively or additionally, retraining or fine-tuning can be scheduled or continuous.
[0103] The trained generative data-driven model can be trained using plant-based training data in one or more languages. The trained generative data-driven model can provide operating instructions for production in one or more languages. Providing operating instructions in one or more languages can be advantageous, for example, because training data may be more readily available in one language than in another. The trained generative data-driven model (e.g., a transformer-based model) trained on plant-based training data in one or more languages can be configured to provide operating instructions in the language in which the most training data is available.Additionally or alternatively, the trained generative data-driven model can be configured to provide production operating instructions in a user-selected language, making the model's operation more user-friendly. Additionally or alternatively, the trained generative data-driven model (e.g., a transformer-based model) can be required to provide production operating instructions in more than one language to cross-check the instructions and select the most suitable ones for improved production.
[0104] Fig. Figure 3 illustrates a preprocessing engine of the in Fig. 2 operating systems shown.
[0105] Plant raw data 1005, as in connection with Fig. As shown in Figure 4, the preprocessing engine 1009 can be provided via a computer interface.
[0106] The 1009 preprocessing engine can preprocess plant-based raw data to provide plant-based training data and / or plant-based input data to the generative data-driven model (e.g., a transformer-based model). The input data can be stored in a database. The input data can be provided to the generative data-driven model via a computer interface (e.g., by requesting the model to access the data). The preprocessing steps can include selecting required parameters, merging / aggregating, and computing plant-based training data, such as calculating derived parameters, removing outliers, and the like.Preprocessing can include filtering data, removing noise, annotating data, sorting data, converting data from formats unsuitable for training / using the model into formats suitable for training / using the model, and / or similar actions. The output data from the preprocessing engine can be stored in a database and used as plant-based training data to retrain or fine-tune a pretrained generative data-driven model (e.g., a transformer-based model), as described in the context of [reference to relevant example]. Fig. 2a, 2c, 6 and 9 to 12 are described. The output data of the preprocessing engine can also be used as input data for the trained generative data-driven model (e.g., transformer-based model), as described in the context of Fig. 1, Fig. 2a, Fig. 2c and Fig. 7 described.
[0107] Fig. Figure 4 illustrates a plant data structure from one or more plants, as shown in Fig. 1 shown.
[0108] Plant data can be received from the distributed production environment via a computer interface (e.g., a graphical user interface, an application programming interface, a web-based interface). This plant data can encompass different categories, such as sensor data, operational data, plant metadata, and analytical data.
[0109] Sensor data can refer to measured quantities that are available in production facilities by means of installed sensors, e.g. temperature sensors, pressure sensors, flow sensors, etc.
[0110] Analytical data can refer to quantities derived from analytical measurements of samples taken at any point in a production plant, such as the composition of a reactant, starting material, product and / or by-product, determined, for example, by gas chromatography on samples taken during production at different stages of the production process, e.g., before or after catalytic reactors.
[0111] Operational data can refer to raw data (basic, unprocessed analysis and / or sensor data) or processed or derived parameters (derived directly or indirectly from raw data).
[0112] Plant metadata can specify a physical plant layout and include plant-specific parameters that describe, for example, the characteristics of one or more reactors predefined by a physical plant layout and which may be relevant to plant or reactor performance.
[0113] The plant data 1005 can include text and / or numbers (structured data). Plant data can also be unstructured. The unstructured data (such as scans, datasheets with images, QR codes, and the like) can be preprocessed by the preprocessing engine and converted into text and / or numbers suitable for training a pretrained, transformer-based model, as described in the context of Fig. Described in sections 9 to 12.
[0114] Fig. Figure 5 illustrates the data structure of an operator.
[0115] A user (e.g., an operator in a production environment) can submit a query to the analytics engine, for example, as a text or audio query, regarding the monitoring and control of a distributed production environment. The query might be a request to provide operating instructions for production, such as instructions related to controlling and / or monitoring one or more production processes within the distributed production environment. For example, the operator could request the trained transformer-based model to provide operating instructions for resolving an anomaly in plant data, fixing a plant operation error, troubleshooting plant operations issues, or providing steps for performing a task related to one or more processes within the distributed production environment.The analysis engine can predict a query solution based on operator input and input of plant-based data. The solution can include machine-readable instructions 1007 (operating instructions), for example, to change operating parameters of the equipment 1002, to replace the equipment 1002 and / or the sensors 1003, to perform maintenance, to carry out one or more steps related to one or more production processes, and / or the like.
[0116] Fig. Figure 6 illustrates further details of retraining or fine-tuning a pre-trained transformer-based model for production, in addition to the embodiments described in Figure 6. Fig. 2a and Fig. 2b.
[0117] A pre-trained transformer-based model that is pre-trained using text and / or numbers (as in, for example, Fig. (illustrated in Figures 9 to 12), can be trained for production based on historical plant data (plant-based training data). In particular, the transformer-based model, as described in the context of Fig. 2b and Fig. 11 are described, trained, and / or parameterized, whereby the transformer encoder, the transformer decoder, and / or the transformer encoder decoder are trained and / or used. For fine-tuning or retraining the model for production using plant-based data instead of generic text / numbers, the following can be used in connection with Fig. The training steps described in 9-12 must be followed, as shown in Fig. Figures 9-12 illustrate this. A trainer of a pre-trained transformer-based model can generate a prompt to trigger fine-tuning or retraining of the pre-trained transformer-based model and specify or upload plant-based training data for training, based on Plant Data 1005, for example, via a computer interface (e.g., web-based interface, API, GUI). At the end of the training cycle, the model trainer can test the trained model and, based on the testing, perform additional training cycles or release the trained model for use in production, as described in the context of Fig. 2a and Fig. 2c described.
[0118] The shared, trained transformer-based model can be stored in a database or the cloud. The shared model can be run by a compute unit / node, which includes a computer processor, such as the analysis engine. A copy of the trained transformer-based model can be made available to a user for use and / or further training. Alternatively, a user can be granted access only to run the trained transformer-based model.
[0119] Instead of the one related to Fig. In addition to the pre-trained transformer-based model described in sections 9 to 12, the analysis engine 1004 can also use a different Foundation model, such as ChatGPT (GPT-2, GPT-3, GPT-3-5-turbo, GPT-4 or higher / similar), Davinci, BERT (Bidirectional Encoder Representations from Transformers), DistilBERT, Transformer-XL, XLNet (eXtreme Language Understanding Network), T5 (Text-to-Text Transfer Transformer), RoBERTa (Robustly Optimized BERT Approach), ELECTRA (Efficiently Learning an Encoder that Classifies Token Replacements Accurately), Reformer, Longformer, DeBERTa (Decoding-Enhanced BERT with Disentangled Attention) or similar, or any other large language model that has been pre-trained on large datasets, such as generic text, image or video data.
[0120] Depending on the availability of plant-based training data, the availability of computing resources, and the required accuracy in data analysis, a user can choose a Foundation model with a smaller or larger number of parameters. The in Fig. The transformer-based model shown in Figures 9 to 12 can be pre-trained to generate a pre-trained model with the required number of parameters to meet the user's technical purpose. The pre-trained model, which is related to Fig. The model pre-trained with parameters 2b, 9 to 12 can be further refined to release a model with even fewer parameters for improved computing speed and reduced computing resources.
[0121] Fig. Figure 7 illustrates the contextualization of prompts.
[0122] A user (an operator in a production environment) can request a trained transformer-based model via a computer interface (e.g., a graphical user interface, an application programming interface, a web-based interface), as in the context of Fig. 2a, Fig. 2b and Fig. Section 6 describes how to resolve a user query regarding the control / monitoring of production. The user can further provide a query context, such as sensor data, operational data, plant metadata, or analytical data. The user can select the context via the computer interface or input the context via text and / or audio channels in one or more languages. The user can also provide a second context. This second context can include, for example, one or more keywords (e.g., anomaly, fault associated with equipment X, fault in software operating equipment X), one or more data points from plant data, such as one or more anomaly data points, one or more standard operating parameters, and the like.
[0123] Based on one or more contexts and optionally plant-based input data, the user can request the trained transformer-based model to predict a solution to the user query. The predicted solution can include machine-readable instructions 1007, which are sent to the control and / or monitoring engine 1006 to operate the production environment 1001.
[0124] An operator can provide one or more contexts in one or more languages and request the trained transformer-based model to provide operating instructions for production in one or more languages.
[0125] Fig. Figure 8 illustrates trigger-based fine-tuning or retraining of a trained transformer-based model when the model is used for production.
[0126] Once trained, as in connection with Fig. 2b, Fig. 6 and Fig. As shown in Figure 11, the trained transformer-based model can be further retrained or fine-tuned based on feedback regarding the operating instructions provided by the trained transformer-based model. For example, an operator can monitor machine-readable instructions 1007 provided by the trained transformer-based model to the control and / or monitoring engine 1006. The operator can also assign a feedback rating and a threshold value regarding the operating instructions provided by the trained transformer-based model. If the rating falls below a threshold value, retraining or fine-tuning can be triggered automatically by a computer processor (e.g., the control / monitoring engine or the analysis engine), or the training can be triggered by a human operator.For example, the rating may fall below the threshold if the machine-readable instructions 1007 for the operator to control the equipment 1002 are not acceptable (e.g., if the machine-readable instructions 1007 are contradictory regarding optimal operating conditions of the equipment 1002).
[0127] The retraining or fine-tuning cycle can include steps such as those related to Fig. 2b, Fig. 6 and Fig. Figures 9-11 illustrate this, with the training data being plant-based data. After a fine-tuning or retraining cycle, the operator can provide a new feedback score for new machine-readable instructions 1007. Training can continue until the feedback score reaches the threshold or a maximum number of training cycles. At the end of the fine-tuning or retraining, the model can be saved and / or released for use in production.
[0128] The retraining can be carried out in the background of ongoing production without having to stop / interrupt production.
[0129] Fig. Figure 9 illustrates an embodiment for training an embedding layer. The embedding layer can be obtained, for example, by training a continuous bag-of-words (CBOW) model or a skip-gram model. The embedding layer can be suitable for generating embedded input data based on input data. Generating embedded input data can refer to embedding input data. Embedding input data can result in a representation associated with the input data. Thus, the embedded input 114 can be the representation associated with the input data. The input data can comprise one or more elements. The one or more elements can be represented by the input vector 106. In particular, the embedded input 114 and / or the input vector 106 can be machine-readable and / or processable by a processor.For this purpose, the embedded input 114 and / or the input vector 106 can be a tensor, in particular a first-rank tensor. Specifically, the input vector 106 can be a one-hot vector or a summation of a plurality of one-hot vectors. A one-hot vector can be a vector with a single non-zero entry. Examples of a one-hot vector might be 108, 110, and 112. The non-zero entries in the one-hot vector and / or in the input vector 106 can specify the element. For example, a lookup table can define the relationship between the position of the non-zero entries and the element specified by the one-hot vector. The lookup table can specify several distinct elements. The number of distinct elements can be equal to the number of entries in the one-hot vector. The number of distinct elements can be referred to as the vocabulary size.In one example, the elements can be represented by tokens, and a sequence of elements can refer to at least one part of a sentence. That at least one part of the sentence can be represented by multiple tokens. A token can represent at least one part of the element and / or word. For example, if an element were associated with only one word, words such as "embeddings," "embedding," or "embed" would represent different elements. A first token can represent the stem "embed," and the endings, which typically appear in a variety of words, can be represented by a second, a third, and a fourth token. The second, third, and fourth tokens can be used to represent other words, such as "see," "seeing," or the like, preferably together with a fifth token representing the stem "see."Ultimately, this tokenization of elements associated with multiple word stems and multiple endings results in fewer tokens to be used to represent multiple elements, and thus uses fewer computing resources.
[0130] A lookup table that specifies a subset of the vocabulary size of, for example, the English language, can contain 10,000 words or more. The embedded input 114 can be a lower-dimensional representation than the input vector 106. For example, typical embedded inputs 114 can contain several hundred different entries. Consequently, the embedded inputs 114 represent a condensed representation of one or more elements using fewer computational resources. Furthermore, the embedded input 114 can represent a relationship between two or more elements. For example, the words "Italy" and "Germany" may be similar or closely related, as they both define European countries, whereas the word "execution form" may be very different from either of these words.The smaller the dot product between two embedded inputs 114, the more similar the two elements associated with the embedded inputs 114 can be. Therefore, the embedded inputs 114 can accurately represent one or more elements and lead to accurate results based on the processing of the embedded inputs 114.
[0131] To transform the input vector 106 into the embedded input 114, the embedding layer can include a number of neurons equal to the number of entries in the embedded input 114. Based on the embedded input 114, the output layer can generate the output vector 116. The output vector can be a vector and / or specify one or more elements. The output vector 116 can specify one or more elements that differ from the input vector 106 and / or the one-hot vector associated with the input vector 106. For this purpose, the output layer can include a number of neurons equal to the number of entries in the input vector 106 and / or the output vector 116. The output layer can apply a softmax function to the embedded input 114.This allows the output vector to include the probabilities associated with the elements that are non-zero and associated with the entries of output vector 116. Therefore, one or more elements with a corresponding probability can be obtained from output vector 116. If input vector 106 specifies one or more sequences of elements, output vector 116 can specify one or more elements corresponding to the sequences of elements specified by input vector 106. In the example of... Fig. 9. The element associated with vector 118 can correspond to the input vector with a probability of 71%. Additional or alternative elements can correspond to the input vector with a lower probability, as indicated by the output vector. By defining a threshold against which the probability can be compared, the selection of the corresponding elements can be tailored to the user's needs. The elements generated by the model, which comprises embedding layer 102 and output layer 104, can be related to the most probable elements specified by output vector 116. Therefore, the in Fig. Model 9 shows that the element associated with vector 118 can be generated with a confidence value of 71%.
[0132] The model of Fig. 9 can be a CBOW (Continuous Bag of Words) model. The CBOW model can be trained based on a training dataset that includes multiple input vectors and corresponding output vectors. Since the training dataset may not be annotated, the training of the CBOW model can be described as self-supervised. Before training the CBOW model, it can be initialized with random values assigned to the neuron weights. During training, the input vectors may be passed through the initialized embedding layer and the output layer, and any loss can be determined by comparing the output vector obtained by the model through passing input vector 106 with the output vector corresponding to input vector 106, as specified by the training dataset.Based on the determined loss, backpropagation can be applied to determine the gradients associated with the neurons of embedding layer 102 and output layer 104 in order to reduce the loss. According to these gradients, the neuron weights can be updated using a gradient descent algorithm. If a predetermined loss can be achieved by the CBOW model, training can be terminated, and a trained CBOW model can be obtained. From the trained CBOW model, embedding layer 102 can be used to embed input data comprising one or more elements. This embedding layer 102 can be used in other machine learning architectures that require an embedding layer 102, such as a transformer-encoder, transformer-decoder, or transformer-encoder-decoder architecture, as described in [reference to relevant worksheet]. Fig. 10A, Fig. 10B and Fig. 10C described. A trained embedding layer 102 may be required to train these architectures. Therefore, a model, such as a CBOW model, can be trained before training the transformer-encoder, transformer-decoder, or transformer-encoder-decoder architecture.
[0133] Fig. Figure 10A illustrates an embodiment of a transformer-encoder architecture. The transformer-encoder comprises an encoder input 278, one or more encoder blocks 274, 214, and an encoder output. The transformer-encoder architecture can be derived from the transformer-encoder-decoder architecture known in the art and described in Fig. Figure 10C shows that the transformer-encoder can be referred to as an X-former. The transformer-encoder architecture can correspond to the encoder architecture associated with the transformer-encoder-decoder architecture, with an additional encoder output, instead of directly connecting the encoder block to the decoder of the transformer-encoder-decoder architecture. Several transformer-encoder architectures are available in the prior art, such as the Bi-directional Encoder Representations from Transformers (BERT).
[0134] The input data can be received at encoder input 278. Encoder input 278 can apply an input embedding 202. Applying the input embedding 202 can refer to routing the input data through an embedding layer, e.g., as in the context of Fig. 9 described. Furthermore, the encoder input 278 can apply a position encoding 204. Applying the position encoding 204 can refer to adding a position factor to the embedded input obtained via an input embedding. Preferably, the input data can specify a sequence of elements. The position factor p pos can specify the position of the elements within the sequence. For example, the position factor p can pos can be obtained based on the following equation: ppos(2i)=sin(pos100002id) ppos(2i+1)=cos(pos100002id) where pos can refer to the position of the element within the sequence, i can refer to the dimension associated with the input embedding, and d can refer to the dimension of the model, e.g., transformer-decoder, transformer-encoder, or transformer-encoder-decoder. This can be referred to as absolute positional embeddings. Alternatively, positional encoding can be based on rotary positional embeddings (RoPE). Positional encoding is advantageous because it allows the processing of sequential data without requiring additional dimensions to specify the position of each element. Consequently, positional encoding reduces the computational resources required to embed the input data. By passing the input data through the encoder input, the input data can be transformed into a second-rank tensor representing the sequence of elements.This second-rank tensor can be referred to as embedded input data. The embedded input data can be processed by the encoder block. The embedded input data can be provided to the layer normalization 208 via a residual connection. A multi-head self-attention 206 can be applied to the embedded input data. Multi-head self-attention 206 can comprise the two components multi-head and self-attention. Self-attention can be understood as a filter applied to the embedded input data. By applying the filter to the embedded input data, the elements associated with the embedded input data that contribute to the output data to be generated can be identified. Therefore, the filter can represent the degree of contribution to the output data to be generated by the elements associated with the embedded input data.Applying the filter can be described as weighting the elements associated with the embedded input data. This is particularly advantageous for long sequences of elements. The filter can be learned and improved during training to identify the contribution of elements associated with the embedded input data. For example, in the phrase "I went to the bakery to buy a loaf of bread," the last word might be generated by the data-driven model, such as the Transformer Encoder. Self-attention can focus the Transformer Encoder on the words "bakery" and "bread" to primarily generate the word "buy." Self-attention can refer to attention generated based on the input data. Therefore, the filter can be determined based on the input data, preferably the embedded input data.The embedded input data can serve as query Q, key K, and value V with respect to the self-attention process. Self-attention can refer to attention based on the received input data. Therefore, the filter can be calculated based on the following formula by inserting the respective tensors based on the embedded input data: Attention(Q,K,V) = Softmax(QKTdk)V where d k corresponds to the dimension of the key.
[0135] To further improve the efficiency of the transformer encoder, multiple heads are used to apply the filter, resulting in multi-head self-attention 206. Multi-head self-attention 206 can involve applying the filter to two or more parts of the embedded input data. Therefore, the tensor can be split into two or more parts, and the filter can be applied separately to the two or more parts by two or more heads according to the following equation: Head i=Attention(QWiQ,KWiK,VWiV) with the parameter matrices W i Q ∈ ℝ d×dQ , W i K ∈ ℝ d×dv , W i V ∈ ℝ d×dV , where i is the number of heads and d V , d K and d Q The dimensions are value, key, and query.
[0136] The result of two or more heads can be combined according to the following equation: MoreHeads(Q, K, V) = Concat(Head 1, ..., Head h )W 0 where W0 ∈ ℝ hdV×d and can specify the number of heads.
[0137] The embedded input data can be transformed into a context tensor using multi-head self-attention 206. The context tensor can represent the sequence of elements and the relationship between two or more elements of the input data. The context tensor can be a second-rank tensor and / or can include one or more first-rank tensors. After multi-head self-attention 206, layer normalization 208 can be applied based on the context tensor and / or the embedded input data from the residual connection. Applying layer normalization 208 can refer to normalizing the context tensor. Normalizing the context tensor can decrease the values of the context tensor entries. This reduces the computational cost associated with processing the context tensor.Layer normalization 208 can be followed by passing the context tensor to a feed-forward layer 210, which in turn is followed by layer normalization 212 based on a residual connection to the context tensor and / or to the output of the feed-forward layer 210. The feed-forward layer 210 can be a neural feed-forward network. The neural feed-forward network can comprise several fully connected neurons. Passing the context tensor through the neural feed-forward network can result in a linear transformation of the context tensor. Additionally or alternatively, the neural network can include one or more activation functions, such as a rectified linear unit (ReLU). Therefore, the neural network can be designed to perform one or more nonlinear operations on the context tensor and / or to nonlinearly transform the context tensor.After the context tensor has been transformed and / or normalized by the feed-forward layer 210 and the layer normalization layer 212, it can be provided to one or more additional encoder blocks 214. After passing through the feed-forward layer 210, the context tensor can be adapted for processing by another attention layer of the one or more additional encoder blocks 214 to apply a self-attention filter, preferably a multi-head self-attention filter 206. After being transformed through the layer normalization layer 212 and the feed-forward layer 210, the context vector can be referred to as the hidden state.
[0138] The encoder output 276 comprises a linear layer 216 and a softmax layer 218. The linear layer 216 can transform the context vector into a logits vector. The linear layer can be fully connected. The logits vector obtained by passing the context tensor through the linear layer 216 can be passed through the softmax layer 218. Passing the logits vector through the softmax layer 218 can refer to applying the softmax function to the logits vector. Applying the softmax function to the logits vector can result in a probability distribution of one or more elements corresponding to the sequence of elements in the input data. One or more elements can be selected from this probability distribution based on predefined selection criteria.The one or more selected elements can be referred to as the one or more elements generated by the transformer encoder. The one or more generated elements can be provided to the encoder input to generate further one or more elements corresponding to the sequence of input data and the one or more elements generated by the transformer encoder, as described in the context of [reference to relevant example]. Fig. 11 described.
[0139] Fig. Figure 10B illustrates an embodiment of a transformer-decoder architecture.
[0140] The transformer-decoder has a decoder input 284, one or more decoder blocks 280, 232, and a decoder output 292. The transformer-decoder architecture can be derived from the transformer-encoder-decoder architecture, as is known in engineering and in Fig. Figure 10C shows that the transformer-decoder can be called an X-former. The transformer-decoder architecture can correspond to the decoder architecture associated with the transformer-encoder-decoder architecture, regardless of whether it receives one or more hidden states from the encoder of the transformer-encoder-decoder. Several transformer-decoder architectures are available in the prior art, such as the generalized pretrained transformers (GPTs).
[0141] The decoder input 284 can apply an input embedding 220 and a position encoding 222 analogously to the input embedding 202 and the position encoding 204, as described in connection with Fig. 10A described.
[0142] Decoder block 280 can include layer normalizations 226, masked multi-head self-attention 224, feed-forward layers 228, and / or layer normalization 230. The embedded input data resulting from routing the input data through decoder input 284 can be provided via a residual connection of layer normalization 226. Furthermore, masked multi-head self-attention 224 can be applied to the embedded input data. Masked multi-head self-attention 224 is equivalent to multi-head self-attention 206, as described in the context of Fig. As described in section 10A, a transformer decoder can be used to additionally mask a portion of the embedded input data associated with elements that appear later in the sequence than the element to be generated. Alternatively, or in addition, the portion of the input data associated with elements that appear later in the sequence than the element to be generated can be either not received and / or not transformed into the embedded input data. Thus, the transformer decoder can be suitable for generating a subsequent element to a sequence, while the transformer encoder can be suitable for generating a missing element within a sequence and / or between two or more sequences. Therefore, the transformer encoder can be designed for classification tasks, while the transformer decoder can be designed for text generation.
[0143] Similar to the transformer encoder, as in connection with Fig. As described in 10A, a context tensor can be generated by applying the masked multi-head self-attention 224 and the layer normalization 226. The context tensor can be provided to the layer normalization 230 via a residual connection. Furthermore, the feed-forward layer 228 and the layer normalization 230 can be analogous to the feed-forward layer 210 and layer normalization 212, as described in the context of Fig. 10A described. The context tensor can be provided to one or more additional decoder blocks 232.
[0144] The decoder output 292 can include a linear layer 234 and a softmax layer 236. The linear layer 234 and the softmax layer 236 can be analogous to the linear layer 216 and the softmax layer 218, as described in the context of Fig. 10A described.
[0145] Fig. Figure 10C illustrates an embodiment of a transformer-encoder-decoder architecture. The transformer-encoder-decoder can include the encoder input 288, one or more encoder blocks 286, 264, the decoder input 294, the decoder block 290, and the decoder output 292. The encoder input 288 can be connected to the encoder input 278 of Fig. 10A. The one or more encoder blocks 286, 264 can correspond to the one or more encoder blocks 274, 214 of Fig. 10A. Decoder input 294 can be used to correspond to decoder input 284 of Fig. 10B corresponds.
[0146] Decoder block 290 can include a masked multi-head self-attention 270, a layer normalization 272, a feed-forward layer 238, and a layer normalization 240 analogous to the masked multi-head self-attention 224, layer normalization 226, feed-forward layer 228, and layer normalization 230, as described in connection with Fig. 2B described. The decoder block 290 can further include a multi-head self-awareness unit 250 and a layer normalization unit 248. Analogous to the description of Fig. 10B, the context tensor can be obtained from the masked multi-head self-awareness 270 and the layer normalization 272. Multi-head self-awareness 250 analogous to multi-head self-awareness 206 from Fig. 10A can be applied to the context vector obtained from layer normalization 272 and the hidden states of one or more encoder blocks 286, 264. Layer normalization 248 can be applied to the context vector obtained from multi-head self-attention 250 and to the context vector obtained from layer normalization 272, which is provided via a residual connection. The context vector resulting from layer normalization 248 can be applied via feed-forward layer 238 and layer normalization 240 analogously to the description of Fig. 10B is processed. The context vector resulting from layer normalization 240 can be provided to further decoder blocks 242, analogous to decoder block 290. The context vector obtained from one or more decoder blocks 290 and 242 can be provided to decoder output 292. Decoder output 292 can be passed to decoder output 282. Fig. 10B corresponds.
[0147] With the architecture described above, the transformer-encoder-decoder can receive and process input data at encoder input 288 and at one or more encoder blocks 286 and 264, as well as decoder block 290 and decoder output 292. Based on the input data, the transformer-encoder-decoder can generate output data partially or sequentially. The sequentially generated output data can be provided to and / or processed by decoder input 294, one or more decoder blocks 290 and 242, and decoder output 292. Preferably, a sequence can be provided to encoder input 288, and after it has generated at least part of the output data, decoder input 294 can be supplied with at least that part of the elements of the already generated output data.This allows the next elements of the output data to be generated with higher accuracy, taking into account the input data and the generated output data, since more data can be received from the transformer-encoder-decoder.
[0148] Due to its transformer-encoder-decoder architecture, the transformer-encoder-decoder can be designed to transform a sequence into a different representation of the sequence. An example of converting a sequence into a different representation could be translating a sentence into another language. Several transformer-encoder-decoders, such as BART, T5, or similar, are available in the prior art.
[0149] In one embodiment, the layer normalization 208, 212 can be applied before the masked multi-head self-attention 224, the multi-head self-attention 206, and / or the feed-forward layer 210 in the transformer-decoder, transformer-encoder, and / or transformer-encoder-decoder. This can reduce the computational resources required to apply the multi-head self-attention 206 and / or the feed-forward layer 210 to the embedded input data and / or the context tensor, since the entries of the respective tensors may be lower after normalization.
[0150] In one embodiment, the decoder output 292 can comprise a neural classification network, further feed-forward layers, convolutional layers, fully connected layers, or the like. For example, the transformer-encoder-decoder can be designed to select between multiple options. To this end, the transformer-encoder-decoder can be provided with three different input data sets and can classify the context vectors obtained from the one or more decoder blocks 290 via one or more linear layers. Consequently, the architecture can be extended depending on the use case to be solved.[1]
[0151] Fig. Figure 11 illustrates one embodiment of training and / or deploying the transformer encoder, the transformer decoder and / or the transformer encoder decoder.
[0152] The encoder / decoder / encoder-decoder architecture 302 can correspond to the transformer-decoder, the transformer-encoder and / or the transformer-encoder-decoder, as described in connection with Fig. 10A- Fig. 10C described.
[0153] The output data generated by the encoder / decoder / encoder-decoder architecture 302 can consist of one or more elements, in particular a sequence of elements. Previously generated elements of the output data can serve as input for generating the next element in the output data sequence.
[0154] In the example of Fig. The input data can comprise N elements, specifically input tokens. An input token can be a token designed to be fed into a data-driven model, such as the transformer-decoder, the transformer-encoder, or the transformer-encoder-decoder. The output data to be generated can comprise M elements. The encoder / decoder / encoder-decoder architecture 302 can generate an element of the output data at a given time step based on receiving the input data and, optionally, previously generated elements of the output data. Therefore, generating M elements requires M time steps. A time step comprises providing an input 310, 312, 314 to the encoder / decoder / encoder-decoder architecture 302 and receiving output data 304, 308, 306 from the encoder / decoder / encoder-decoder architecture 302. In a first time step, the input 310 can comprise N input tokens. The N input tokens can, for example,The input tokens may be associated with N words, stems, or endings. Preferably, the N input tokens can specify a question. One or more input tokens can specify the beginning and / or end of the sequence of tokens. The input 310 can be processed by the encoder / decoder / encoder-decoder architecture 302. Based on the input 310, at least one part of the output data 304 can be generated. This at least one part of the output data can include a first output token. In the next time step, the generated first output token can be provided along with the input 312. In particular, while the input 312 can be received by a transformer-encoder-decoder, the input tokens can be received at the encoder input 288, and the first output token can be received at the decoder input 294.While input 312 can be received by the transformer encoder, input 312 can also be received by encoder input 278, and analogously by the transformer decoder and decoder input 284. Based on input 312, output data 308, comprising the first output token and a second output token, can be generated. Generating output data 308 based on input 312 can refer to generating the second token based on the first token and the N input tokens, where the first token may have been generated based on the N input tokens. This process can be repeated until the last token in the output data sequence 306 is generated. Preferably, the last token can be an end token. The end token can terminate the generation of a further output token.
[0155] Similar to data processing during the use of the Encoder / Decoder / Encoder-Decoder Architecture 302, the Encoder / Decoder / Encoder-Decoder Architecture 302 can be trained. The training dataset can comprise multiple sequences, each containing multiple elements. The sequences can be associated with the input data and / or the output data. Additionally or alternatively, the sequences can be independent of the input data and / or the output data. For example, if the input data and the output data can relate to chemical compositions represented by text, the training dataset can comprise sequential text data that is independent of chemical compositions. In this example, the training dataset can comprise sequences of words taken from a conversation. In one embodiment, the training dataset can comprise at least some input datasets and / or output datasets.
[0156] The training can be initialized by initializing the encoder / decoder / encoder-decoder architecture 302. In one embodiment, the parameters associated with the encoder / decoder / encoder-decoder architecture 302 can be initialized randomly. Additionally or alternatively, the input embedding of the encoder / decoder / encoder-decoder architecture 302 can be obtained by training a CBOW model or a skip-gram model, as described in the context of Fig. 9. The trained embedding layer can be used during training. The parameters associated with the embedding layer can be kept constant and / or updated after a predefined number of training epochs. This reduces the number of parameters to be updated, enabling faster and less computationally intensive training. Furthermore, the accuracy associated with the embedding layer can be kept constant and / or increased by avoiding error compensation with respect to the newly initialized encoder / decoder / encoder-decoder architecture 302.
[0157] During the training of the Encoder / Decoder / Encoder-Decoder Architecture 302, at least a portion of the sequences from the training dataset can be provided sequentially, and one or more elements can be generated sequentially based on these training dataset sequences. The elements generated based on these sequences can follow the elements of the portions of sequences with which the Encoder / Decoder / Encoder-Decoder Architecture 302 can be provided. The generated element(s) can be compared to the element(s) that follow the at least portion of the sequences provided to the Encoder / Decoder / Encoder-Decoder Architecture 302, as specified by the training dataset.Therefore, during training, the Encoder / Decoder / Encoder-Decoder Architecture 302 can generate an estimate of the next element, and this estimate can be compared to the ground truth, which indicates the actual next element according to the training dataset. Based on the next element estimate and the ground truth, a loss can be determined. The loss can define the similarity between the next element estimate and the ground truth. The loss can be determined by forming a vector dot product between the token associated with the one or more elements and the token associated with the ground truth. A non-zero loss can cause the parameters associated with the Encoder / Decoder / Encoder-Decoder Architecture 302 to be updated.Preferably, the parameters associated with the encoder / decoder / encoder-decoder architecture 302 can be independent of the embedding layer. For example, the parameters associated with the encoder / decoder / encoder-decoder architecture 302 can be weights of the neurons of the encoder / decoder / encoder-decoder architecture 302.
[0158] Based on the determined loss, backpropagation can be applied to determine the gradients associated with the parameters of the encoder / decoder / encoder-decoder architecture 302 in order to reduce the loss. According to the determined gradients, the parameters associated with the encoder / decoder / encoder-decoder architecture 302, preferably the weights of the neurons associated with the encoder / decoder / encoder-decoder architecture 302, can be updated using a gradient descent algorithm.
[0159] The training dataset can be unannotated. The sequences of elements within the training dataset can inherently include the ground truth for determining the loss with respect to the one or more elements generated during the training of the Encoder / Decoder / Encoder-Decoder Architecture 302. Therefore, the Encoder / Decoder / Encoder-Decoder Architecture 302 can be self-supervised. This is advantageous because it saves the time and resources required to create an annotated training dataset. Furthermore, it allows the use of large training datasets, which can be several terabytes in size. Consequently, the data-driven model can be accurate in generating elements of a sequence. Moreover, the large training dataset enables predictions with few examples or even predictions without any examples.Therefore, data-driven models trained as described above contribute significantly to saving resources needed to train and / or host multiple purpose-driven models, such as CNNs. The training described above can be referred to as pre-training. After pre-training, the data-driven model can be designed to make predictions with few examples, or even without any examples, across a wide range of use cases. The performance of the data-driven model can be further enhanced through additional training, known as fine-tuning.
[0160] Fig. Figure 12 illustrates an embodiment of the input embedding. If the sequence of elements associated with, and preferably contained in, the input data is of a certain type, the input embedding can be 202, 220, 252, 266, as described in connection with Fig. 10A - Fig. The input data described in Section 10C can be used. For example, one type of input data can be text, where the elements can be associated with at least part of a word, a punctuation mark, a start token that specifies the beginning of one or more sequences associated with the input data, and / or an end token. In another example, the input data can be at least partially numeric. Therefore, the input data can include multiple numbers. Numeric input data can, for example, be tabular data. Tabular data can specify one or more rows and / or one or more columns. Therefore, the tabular data can include one or more cells, where the cells can be associated with one or more numeric values.
[0161] Numeric input data may require a different embedding than text input data. Input embeddings for numeric input data can include token embedding, positional embedding, column embedding, row embedding, or a combination thereof.
[0162] Applying token embedding to one or more elements, particularly tokens associated with the input data, can result in a machine-processable representation associated with those one or more elements, particularly tokens. Applying token embedding to one or more elements can refer to passing those elements through the embedding layer, for example, as in the context of... Fig. 9 described. Thus, token embeddings can specify one or more elements, especially tokens, in a machine-processable representation. For example, token embedding can transform a numerical value into a vector. This is advantageous because this representation can be enriched with further information, such as the token's position within the sequence and / or within a table associated with the sequence of tokens. Positional embedding can be analogous to positional embedding, as described in the context of Fig. 9, Fig. 10A- Fig. 10C described. If the input data can be tabular, column embedding can be applied. Applying column embedding to one or more elements, in particular tokens, associated with the input data can result in a machine-processable representation that specifies the location of the one or more elements within a Table 402, preferably within the columns of Table 402. Applying column embedding can involve adding a column factor to the input data embedded via token embedding, in particular the embedded input data. The column factor can be the same for elements associated with the same column and / or can be different between two or more elements associated with different columns. Similarly, row embedding can be applied, where the input data can be tabular.Applying row embedding to one or more elements, particularly tokens, associated with the input data can result in a machine-processable representation that specifies the location of the one or more elements within a Table 402, preferably within the rows of Table 402. Applying row embedding can involve adding a column factor to the input data embedded via token embedding, particularly the embedded input data. The row factor can be the same for elements associated with the same row and / or can be different between two or more elements associated with different rows.
[0163] In one embodiment, input data can be at least partially numeric and at least partially textual. Therefore, the input data can comprise two or more data types. A data type can refer to a modality. Consequently, different embeddings can be applied to the input data. The portions of the input data that comprise text can be subjected to the embedding described in Fig. 9, Fig. 10A- Fig. The input embedding specified in 10C can be applied. Positional, columnal, and row embedding can be applied to portions of the input data that are numeric token embeddings. Furthermore, segment embedding can be applied to the input data regardless of the input data type. Segment embedding can specify the type of input data with which one or more elements can be associated. For example, if the input data includes text and numbers, the input data can comprise two types of input data. Applying segment embedding to the input data can involve adding a segment factor to the input data, preferably to the embedded input data and / or to the input data after token embedding has been applied. The segment factor can specify the data type associated with the one or more elements.The segment factor can be the same for one or more elements associated with the same type of input data and / or can be different between two or more elements associated with different types of input data.
[0164] Applying token embedding, position embedding, segment embedding, column embedding, row embedding, or a combination thereof, can result in embedded input data and / or the output of any encoder input 278, 284, 288 or decoder input 284, 294. The data obtained by applying token embedding, position embedding, segment embedding, column embedding, row embedding, or a combination thereof, can be processed by encoder block 274, 286, decoder block 280, 290, encoder output 276, or decoder output 292, 282.
[0165] The plant-based training data 1013, an untrained / pretrained transformer-based model, as in the context of Fig. 9-12 described, and / or the trained transformer-based model, as described in connection with Fig. 2a, Fig. 2c, Fig. 6 and Fig. The data described in section 8 can be stored in a database, on an electronic data carrier, or in a cloud. Access to the plant-based training data 1013, which is used to train a pre-trained transformer-based model, an untrained / pre-trained transformer-based model, as described in connection with Fig. 9-12 described, and / or the trained transformer-based model, as described in connection with Fig. 2a, Fig. 2c, Fig. 6 and Fig. As described in section 8, suitable tokens can be granted to a user as a computer-readable token. The computer-readable token can be an authorization key generated by a computer processor in response to a request from a requesting compute node associated with a user. The request can be sent to an authorization engine comprising at least one computer processor with the rights to grant one or more authorization keys to access the plant-based training data and / or one or more of the untrained, pretrained, trained, retrained, or fine-tuned transformer-based models, as described in the context of Fig. 2a, 2c, 6, 8 and 9-12 are described.
[0166] A user can access the training data via the token to process the data and receive the processed result without receiving the actual training data. The user can also receive the actual plant-based training data or a portion thereof. Depending on their access rights, a user (upon receiving the token) can use one or more of the untrained, pretrained, and / or trained transformer-based models for their technical purposes. The user can train or retrain one or more models for which they have been granted permission, using the claimed plant-based training data or their own training data. An access token for using the claimed plant-based training data can be combined with an access token for using one or more models, as described in the context of... Fig. 2a, 2c, 6, 8 and 9-12, may be identical or different. A separate access token may be used for each of the data products (i.e., training plant-based data) or data services (i.e., using one or more models, as described in connection with Fig. As described in sections 2a, 2c, 6, 8 and 9-12, access tokens can enhance the security aspects of using data products and / or data services. An access token may include user credentials, one or more generation algorithms, user authentication, two-factor authentication and / or the like.
[0167] One possible implementation of the disclosure of the current application is described in the list below.
[0168] Points: Item 1. Computer-implemented method for releasing a trained transformer-based model for controlling and / or monitoring a distributed production environment, wherein the distributed production environment comprises one or more pieces of equipment that manufacture a product, wherein the method comprises the following: Providing, via a computer interface, plant-based training data associated with one or more production processes; Provide, via the computer interface, a pre-trained transformer-based model that includes at least one transformer component; Requesting the pre-trained transformer-based model to fine-tune or retrain using the plant-based training data; Releasing the trained transformer-based model for one or more production processes in the distributed production environment. Item 2. Computer-implemented method for using a trained transformer-based model according to Item 1 to control and / or monitor a distributed production environment, wherein the distributed production environment comprises one or more pieces of equipment that manufacture a product, wherein the method comprises the following: Received, by a computer processor, from access to a trained transformer-based model; Receiving, via a computer interface, plant-based input data associated with one or more production processes; Instructing the trained transformer-based model to analyze the plant-based input data and provide operating instructions for production. Point 3. Computer-implemented method according to point 1, wherein plant-based raw data are pre-processed by a pre-processing engine to provide the plant-based training data. Point 4. Computer-implemented method according to point 2, wherein the plant-based raw data are pre-processed by a pre-processing engine to provide the plant-based input data. Point 5. Computer-implemented method according to point 2 or 4, wherein the plant-based input data includes data from an operator provided to the operator via a computer interface. Point 6. Computer-implemented method according to any of points 2, 4 or 5, further comprising a prompt from an operator requesting the trained transformer-based model to provide machine-readable instructions based on a request provided to the model by the operator via a computer interface. Point 7. Computer-implemented procedure according to one of points 2 and 4-6, wherein the prompt includes one or more contexts relating to one or more production operations in the distributed production environment. Item 8. Computer-implemented method according to one of items 1-7, wherein the plant-based training data are plant-based data in one or more languages and / or the operating instructions for production are in one or more languages. Item 9. Computer-implemented method according to one of items 1-8, wherein the method further includes updating the plant-based training data and requesting the trained transformer-based model to fine-tune or retrain based on the updated plant-based training data. Item 10. Computer-implemented method according to Item 9, wherein fine-tuning or retraining of the trained transformer-based model is schedule-based fine-tuning or retraining, continuous fine-tuning or retraining, or trigger-based fine-tuning or retraining. Item 11. Computer-implemented procedure according to Item 10, wherein the trigger is based on a threshold and an evaluation in relation to the operating instructions provided by the trained transformer-based model and / or optionally a number of fine-tuning or retraining cycles. Item 12. Computer-implemented procedure according to one of items 2-11, wherein the operating instructions include machine-readable instructions for controlling and / or monitoring production. Item 13. Computer program product comprising computer-readable instructions which, when executed on a computer, cause the computer to perform the steps of any of items 1-12. Item 14. Computer-readable storage medium that stores computer-readable instructions which, when executed on a computer, cause the computer to perform the steps of any of items 1-12. Item 15. Computer product comprising a computer-readable token for accessing plant-based training data and / or the trained or pre-trained model according to any of items 1-12.
[0169] The following exemplary embodiments shall also be disclosed: Clause 1: A computer-implemented method for using a trained transformer-based model to control and / or monitor a distributed production environment, wherein the distributed production environment comprises one or more pieces of equipment that manufacture a product, and wherein the method comprises: - Received, by a computer processor, access to the trained transformer-based model; - Receiving, via a computer interface, plant-based input data associated with one or more production processes; - Instructing the trained transformer-based model to analyze the plant-based input data and provide operating instructions for one or more production operations in the distributed production environment. Clause 2: A computer-implemented method for generating the trained transformer-based model for use in clause 1, wherein the method comprises: - Providing, via a computer interface, plant-based training data associated with one or more production processes; - Provide, via the computer interface, a pre-trained transformer-based model that includes at least one transformer component; - Requesting the pre-trained transformer-based model to perform a retraining or fine-tuning using the plant-based training data; - Releasing the trained transformer-based model for one or more production processes of the distributed production environment according to clause 1. Clause 3: Computer-implemented method according to clause 2, wherein plant-based raw data are preprocessed by a preprocessing engine to provide the plant-based training data. Clause 4: Computer-implemented method according to clause 1, wherein plant-based raw data are pre-processed by a pre-processing engine to provide the plant-based input data. Clause 5: Computer-implemented method according to clause 1 or 4, wherein the plant-based input data includes data from an operator provided to the operator via a computer interface. Clause 6: A computer-implemented method according to one of clauses 1, 4 or 5, further comprising a prompt from an operator requesting the trained transformer-based model to provide machine-readable instructions based on a query provided to the model by the operator via a computer interface. Clause 7: Computer-implemented procedure according to one of clauses 1 and 4-6, wherein the prompt includes one or more contexts relating to one or more production operations in the distributed production environment. Clause 8: Computer-implemented method according to one of clauses 2-7, wherein the plant-based training data are plant-based data in one or more languages and / or the operating instructions for production are in one or more languages. Clause 9: Computer-implemented method according to one of clauses 2-8, wherein the method further includes updating the plant-based training data and requesting the trained transformer-based model to retrain or fine-tune based on the updated plant-based training data. Clause 10: Computer-implemented procedure according to clause 9, wherein the retraining of the trained transformer-based model is one of a schedule-based retraining, continuous retraining or fine-tuning, or trigger-based retraining or fine-tuning. Clause 11: Computer-implemented procedure according to clause 10, wherein the trigger is based on a threshold and an evaluation in relation to the operating instructions provided by the trained transformer-based model and / or optionally a number of retraining or fine-tuning cycles. Clause 12: Computer-implemented procedure according to one of clauses 1-11, wherein the operating instructions include machine-readable instructions for controlling and / or monitoring production. Clause 13: Computer program product comprising computer-readable instructions which, when executed on a computer, cause the computer to perform the steps of any one of clauses 1-12. Clause 14: A computer-readable storage medium that stores computer-readable instructions which, when executed on a computer, cause the computer to perform the steps of any one of clauses 1-12. Clause 15: Computer product comprising a computer-readable token for accessing plant-based training data and / or the trained, pre-trained, or fine-tuned model according to any of clauses 1-12.
[0170] It is understood that the features of the above embodiments can be combined, unless otherwise disclosed in the present application.
[0171] The publication Prior Art Disclosure; Issue 684; paragraphs
[1000] to
[8005] ; ISSN: 2198-4786; published: February 12, 2024, is considered Reference RF1, which is incorporated herein by reference in its entirety. Preferably, the (chemical) product is a product as described in Reference RF1, paragraphs
[1000] to
[8005] . Preferably, the method / process described herein is furthermore a method / process for producing a product. The transformation step to obtain the product preferably comprises one or more steps as described below and may be carried out by conventional methods well known to those skilled in the art. The transformation step preferably comprises one or more steps selected from: Recycling, preferably depolymerization, gasification, pyrolysis and / or steam cracking; and / or purification, preferably crystallization, (solvent-)based extraction, distillation, evaporation, hydrotreating, absorption, adsorption and / or ion exchange treatment; and / or assembly, preferably foaming, synthesis, chemical conversion, chemical transformation, polymerization and / or compounding; and / or Shaping, preferably foaming, extruding and / or casting; and / or finishing, preferably coating and / or smoothing. In addition, one or more steps are described in detail in Reference RF1; paragraphs
[1000] to
[8005] .
[0172] The present disclosure has been described in conjunction with preferred embodiments and examples. However, further modifications can be recognized and implemented by those skilled in the art using the claimed subject matter with reference to the drawings, the present disclosure, and the claims. In particular, all the steps shown can be carried out in any order; that is, the present disclosure is not limited to a specific sequence of these steps. Furthermore, it is also not necessary that the various steps be carried out at a specific location or at a node of a distributed system; that is, each of the steps can be carried out at different nodes using different equipment / data processing.
[0173] The sequence of all process steps shown above is not mandatory; alternative sequences are also possible. Nevertheless, the specific sequence of process steps shown in the figures should be considered as one possible sequence, e.g., for the respective embodiment described by the respective figure or an embodiment that includes at least some of the steps described by the respective figure.
[0174] In the present description, each connection depicted in the described embodiments is to be understood as one in which the components involved are functionally coupled. The connections can be direct or indirect via any number or combination of interposed elements, whereby only a functional relationship between the components is required.
[0175] The indefinite article "a" is not to be understood as "one," meaning that the use of the expression "a single element" does not preclude the presence of other elements. A single element or other unit can fulfill the functions of several instances or elements mentioned in the claims. The mere fact that certain measures are specified in different dependent claims does not mean that a combination of these measures cannot be used in an advantageous embodiment or that further elements cannot be included.
[0176] The expressions “A and / or B” and “at least one of: A or B” are considered interchangeable and are intended to encompass one of the following three scenarios: (i) A, (ii) B, (iii) A and B. More generally, the expression “at least one of the following” means:<Liste von zwei oder mehr Elementen> “and “at least one of<Liste von zwei oder mehr Elementen> “and similar formulations where the list of two or more elements is joined by “and” or “or”, at least one of the elements or at least two or more of the elements or at least all of the elements.
[0177] Providing data within the scope of protection of this disclosure can include any interface designed to provide data. This can include an application programming interface, a human-machine interface such as a display, and / or a software module interface. Providing data can include communicating data or transmitting data to the interface, in particular displaying it to a user or using the data by the receiving instance. Receiving data within the scope of protection of this disclosure can include any interface designed to receive or obtain data. This can include an application programming interface, a human-machine interface such as a display, and / or a software module interface.Receiving data can involve communication or transmission of data from the interface, in particular the use of the data by the receiving instance. Any receiving of data, data structures, records, or the like can include receiving the data, data structures, records, or the like from a server that provides (e.g., hosts) a database containing the data, data structures, records, or the like.
[0178] Various units, circuits, instances, nodes, or other computing components may be described as "designed" to perform one or more tasks. "Designed" indicates a structure that means a circuit exists which, when operational, performs the task or tasks. The units, circuits, instances, nodes, or other computing components may be designed to perform the task even when the unit / circuit / component is not operating. The units, circuits, instances, nodes, or other computing components that form the "designed" structure may include hardware circuits and / or memory that stores program instructions executable to implement the operations. For simplicity, the units, circuits, instances, nodes, or other computing components may be described as performing one or more tasks.Such descriptions are to be interpreted as containing the phrase "interpreted for". Any reference to "interpreted" is expressly not intended to trigger an interpretation under 35 U.S.C. § 112(f).
[0179] In general, the methods, devices, systems, computer elements, nodes described herein, or other computing components can include memory, software components, and hardware components. Memory can include volatile memory such as static or dynamic random-access memory and / or non-volatile memory such as optical or magnetic disk storage, flash memory, programmable read-only memory, etc. Hardware components can include any combination of combinational logic circuitry, clocked memory devices such as FLOPS, registers, latches, etc., finite state machines, memory such as static random-access memory or embedded dynamic random-access memory, custom-designed circuitry, programmable logic matrices, etc.
[0180] In the present description, each connection depicted in the described embodiments is to be understood as one in which the components involved are functionally coupled. The connections can be direct or indirect via any number or combination of interposed elements, whereby only a functional relationship between the components is required.
[0181] Furthermore, any of the procedures, processes, and actions described or illustrated herein can be implemented in a general-purpose or specialized processor using executable instructions and stored on a computer-readable storage medium (e.g., hard disk, memory, or the like) for execution by such a processor. References to a "computer-readable storage medium" should be understood to include specialized circuitry, such as signal processing devices and other equipment.
[0182] Any disclosure and embodiments described herein relate to the methods, systems, devices, and computer program element set forth above, and vice versa. Advantageously, the advantages provided by one embodiment and example apply equally to all other embodiments and examples, and vice versa.
[0183] All terms and definitions used herein are to be interpreted broadly and have their general meaning unless otherwise stated.
[0184] It is understood that all embodiments shown are merely examples and that each feature shown for a particular embodiment can be used individually or in combination with any feature of the same or any other particular embodiment and / or in combination with any other feature not mentioned. In particular, the exemplary embodiments shown in this description are also to be understood as disclosed in all technically meaningful combinations with one another, insofar as the exemplary embodiments do not represent alternatives to one another. It is further understood that each feature shown for an embodiment in a particular category (method / device / computer program / system) can also be used accordingly in an exemplary embodiment of any other category.It is also understood that the presence of a feature in the illustrated embodiments does not necessarily mean that this feature is an essential feature and cannot be omitted or replaced. REFERENCE LIST 1001 plant(s); 1002 Equipment; 1003 sensors; 1004 Analysis Engine; 1005 Plant data, e.g. sensor data, analysis data; 1006 Control and / or monitoring engine; 1007, 1008 machine-readable instructions; 1010 historical plant data; 1011 input data; 1012 Operator data; 1013 training data points; 7001 Inviting the AI engine by providing a prompt, such as a text prompt (e.g., fix an anomaly in plant operation); 7002 Adding context to the prompt (e.g., temperature anomaly); 7003 Optional: Add another context (e.g., a data point from the preprocessing engine, such as an identification of a plant in metadata, a typical operating temperature of the plant, and / or the like); 7004 Generating instructions by AI to monitor and / or control production based on the prompt (e.g., instructions to correct the temperature anomaly). QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature
[0000] WO 2020165045 (A1
[0006] WO 2021116123 (A1
[0006] WO 2021156157 (A1
[0006] Cited non-patent literature
[0000] Attention Is All You Need” by Vaswani et al., 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA (6 Dec 2017, arXiv:1706.03762v5
[0032] ISO / IEC 23053:2022(en), ISO / IEC TR 24372:2021(en), ISO / IEC 22989, ISO / IEC 23053
[0033] ISO / IEC 20546:2019
[0033] ISO / IEC 20546:2019(en), ISO 8000-66:2021
[0033] ISO / IEC DIS 5259-1
[0033] ISO / IEC 20546:2019(en), ISO 8000-66:2021 (en) / Data quality, ISO / IEC DIS 5259-1
[0038] Prior Art Disclosure; Issue 684; Paragraphs
[1000] to
[8005] ; ISSN: 2198-4786; published: February 12, 2024
[0171]
Claims
[1] Method for controlling and / or monitoring a distributed production environment, wherein the distributed production environment comprises at least one piece of equipment for producing a product, wherein the method comprises: - Receiving, via a computer interface, plant-based input data associated with one or more production processes; - Determining operating instructions for one or more production processes of the distributed production environment, wherein determining the operating instructions includes providing a prompt to at least one generative data-driven model trained to generate the operating instructions in response to receiving the prompt; - Providing the operating instructions. [2] Method for generating the at least one generative data-driven model for use according to claim 1, wherein the method comprises: - Providing, via a computer interface, plant-based training data associated with one or more production processes; - Providing a pre-trained generative data-driven model via the computer interface; - Fine-tuning the pre-trained, generative, data-driven model using the plant-based training data; - Releasing the trained generative data-driven model for one or more production processes of the distributed production environment according to claim 1. [3] Method according to claim 2, wherein plant-based raw data are preprocessed by a preprocessing engine to provide the plant-based training data. [4] Method according to claim 1, wherein plant-based raw data are preprocessed by a preprocessing engine to provide the plant-based input data. [5] Method according to claim 1 or 4, wherein the plant-based input data comprises data from an operator provided via a computer interface. [6] Method according to one of claims 1, 4 or 5, wherein the prompt comprises an instruction for the trained at least one generative data-driven model to provide machine-readable instructions. [7] Method according to one of claims 1 and 4-6, wherein the prompt comprises one or more contexts relating to one or more production processes in the distributed production environment. [8] Method according to any one of claims 2-7, wherein the plant-based training data are plant-based data in one or more languages and / or the operating instructions for production are in one or more languages. [9] Method according to any one of claims 2-8, wherein the method further comprises updating the plant-based training data and fine-tuning or retraining the trained generative data-based model based on the updated plant-based training data. [10] Method according to claim 9, wherein the retraining or fine-tuning of the trained generative data-driven model is a schedule-based retraining or fine-tuning, a continuous retraining or fine-tuning, or a trigger-based retraining or fine-tuning. [11] Method according to claim 10, wherein the trigger is based on a threshold and an evaluation in relation to the operating instructions provided by the trained generative data-driven model and / or optionally a number of retraining or fine-tuning cycles. [12] Method according to any one of claims 1-11, wherein the operating instructions include machine-readable instructions for controlling and / or monitoring production. [13] Computer program product comprising computer-readable instructions which, when executed on a computer, cause the computer to perform the steps of any one of claims 1-12. [14] Device comprising means for performing or carrying out the steps of any one of claims 1 to 12 or comprising at least one processor and at least one memory which stores instructions which, when executed by the at least one processor, cause the device to carry out at least the steps of the method according to any one of claims 1 to 12. [15] Computer product comprising a computer-readable token for accessing plant-based training data and / or the trained or pre-trained or fine-tuned model according to any one of claims 1-12.
Citation Information
Patent Citations
Determining operating conditions in chemical production plants
WO2020165045A1
Manufacturing system for monitoring and / or controlling one or more chemical plant(s)
WO2021116123A1
Industrial plant monitoring
WO2021156157A1