System and computer-implemented method for distributed production environment

By using a generative data-driven model in distributed production environments such as chemical plants, which receives plant data and generates machine-readable instructions, the challenges of safely controlling chemical processes are addressed, and efficient and safe production management is achieved.

CN121729646APending Publication Date: 2026-03-24BASF SE
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-22
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

In distributed production environments, especially in places like chemical plants, there are challenges in safely and effectively controlling and monitoring chemical processes, including the handling, storage, and conversion of hazardous chemicals. Existing technologies struggle to manage these processes efficiently and safely.

Method used

Generative data-driven models (such as converter-based models) are employed to receive factory data, provide operational instructions, and fine-tune the model by combining operator prompts and training data to generate machine-readable instructions for controlling and monitoring the production process.

Benefits of technology

It improves the efficiency and safety of chemical production, ensures efficient, safe and reliable operation, adapts to different languages ​​and factory environments, and supports cross-plant knowledge transfer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present disclosure may relate to controlling and / or monitoring a distributed production environment, such as chemical production, in a secure manner using generative data driven models, such as transformer-based models. The disclosed method may involve determining and providing operational instructions for one or more production operations of the distributed production environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The following disclosure relates to the field of computer-aided production, such as production of chemicals. The following disclosure can relate to trustworthy AI. BACKGROUND

[0002] Distributed production environments, such as chemical plants, are very complex, in which various chemical processes are conducted under the control of various chemical equipment. These processes can involve handling, storage, and transformation of different chemicals, which can include hazardous, flammable, and / or toxic chemicals. These processes can be conducted, for example, at high temperatures and / or high pressures, and safe and efficient control of these processes can involve operating several pieces of equipment in a precise time sequence, for which strict work flows can be defined. Adhering to these work flows can help reduce risks associated with hazardous materials, protect the health and safety of personnel, and protect the environment. SUMMARY

[0003] According to a first aspect, a method for controlling and / or monitoring a distributed production environment comprising one or more pieces of equipment producing a product is disclosed. The method comprises:

[0004] - receiving, via a computer interface, factory-based data of input associated with one or more production operations;

[0005] - determining operational instructions for the one or more production operations of the distributed production environment, the determining of the operational instructions comprising providing a prompt to at least one generative data-driven model in response to receiving the prompt, the at least one generative data-driven model having been trained to generate operational instructions;

[0006] - providing the operational instructions.

[0007] According to further aspects, corresponding devices, systems, and uses are disclosed.

[0008] EMBODIMENTS

[0009] In a distributed production environment, such as a chemical plant, various chemical processes can take place, which are controlled by various chemical equipment. These processes can involve handling, storage, and transformation of different chemicals, which can include hazardous, flammable, and / or toxic chemicals. These processes can be conducted, for example, at high temperatures and / or high pressures, and safe and efficient control of these processes can involve operating several pieces of equipment in a precise time sequence. Adhering to these work flows can help reduce risks associated with hazardous materials, protect the health and safety of personnel, and protect the environment.

[0010] For example, WO2020165045 (A1), WO2021116123 (A1), WO2021156157 (A1) can illustrate that chemical production can involve large amounts of data, and that multi-parameter data streams can be provided.

[0011] The present disclosure can relate to how to improve control and / or monitoring of a distributed production environment.

[0012] The various aspects, embodiments and examples provided in the present disclosure can allow for this using a generative data-driven model (e.g., a transformer-based model). The aspects, embodiments and examples provided in the present disclosure can allow for improving efficiency of chemical production.

[0013] According to a first aspect, a (in particular computer-implemented) method for controlling and / or monitoring a distributed production environment comprising one or more pieces of equipment (e.g., chemical equipment such as reactors) producing a product, the method comprising:

[0014] - receiving, via a computer interface, input factory-based data associated with one or more production operations;

[0015] - determining operational instructions for the one or more production operations of the distributed production environment, the determining of the operational instructions comprising providing a prompt to at least one generative data-driven model in response to receiving the prompt, the at least one generative data-driven model having been trained to generate operational instructions;

[0016] - providing the operational instructions (e.g., to an operator of the factory / distributed production environment, or to a control system, e.g., after being validated by the operator, e.g., when the operational instructions are at least partially provided in a machine-readable structure (e.g., machine-readable instructions)).

[0017] The providing of the prompt to the at least one generative data-driven model having been trained to generate operational instructions in response to receiving the prompt can comprise prompting the trained generative data-driven model (e.g., a transformer-based model) to analyze the input factory-based data and to provide operational instructions for the one or more production operations of the distributed production environment.

[0018] According to a second aspect, a (in particular computer-implemented) method for generating at least one generative data-driven model for use in the method according to the first aspect, the method comprising:

[0019] - providing, via a computer interface, factory-based training data associated with one or more production operations;

[0020] - Provide a pre-trained generative data-driven model (e.g., a transformer-based model that includes at least a transformer component) via a computer interface.

[0021] - Use factory-based training data to fine-tune pre-trained generative data-driven models;

[0022] - Publish trained generative data-driven models for use in one or more production operations within a distributed production environment based on the first aspect.

[0023] According to the example implementation of the first aspect, the raw plant-based data is preprocessed by a preprocessing engine to provide input plant-based data.

[0024] According to the example implementation of the second aspect, the raw factory-based data is preprocessed by a preprocessing engine to provide factory-based training data.

[0025] Raw data can be unprocessed, unfiltered, and unmodified data collected directly from sources such as sensors in a distributed production environment. Raw data may lack structure or context. Because raw data is collected directly from the source, it may contain errors, noise, or inconsistencies introduced by the data collection process itself. Raw data may also include irrelevant or redundant information.

[0026] According to any example implementation, the plant-based input data includes data provided to the operator via a computer interface. For example, an operator in a distributed production environment can provide specific data points that may be particularly relevant to the distributed production environment, such as when the operator suspects or anomalies in or wishes to constrain operational instructions provided by a generative data-driven model, for example, constraining parameters to a range of production process parameters or changing only the parameters of specific chemical equipment / equipment or excluding operational instructions related to certain parts of the production process. Data from the operator may, for example, involve anomalous data points, typical or optimal operating parameters in the distributed production environment. Data from the operator can form part of the prompts used by a trained generative data-driven model.

[0027] According to an example implementation of the method according to the first aspect, the prompt (e.g., from an operator or generated by the operator from a query) includes instructions that provide machine-readable instructions to at least one trained generative data-driven model. Such machine-readable instructions may be in a format for controlling at least a portion of a distributed production environment (e.g., a specific chemical device). Therefore, an operator can directly use the provided machine-readable instructions to control the (e.g., chemical) distributed production environment, or the system can directly use the machine-readable instructions, for example, after the operator is prompted to verify the generated operational instructions.

[0028] According to an example implementation of the method according to the first aspect, the prompt includes one or more contexts related to one or more production operations in a distributed production environment. Providing one or more production operations as specific contexts in a distributed production environment can enhance the quality and accuracy of the generated operational instructions, and thus can further enhance the efficient, secure, and reliable operation of the distributed production environment.

[0029] According to an example implementation of the method according to the first aspect, the operating instructions for manufacturing exist in one or more languages. This can, for example, allow operators speaking different local languages ​​to quickly understand the operating instructions and act accordingly, and can allow the transfer of operational knowledge in a distributed production environment between factories or production lines located in different countries.

[0030] According to the example implementation of the method in the second aspect, the factory-based training data is factory-based data in one or more languages. Training data in different languages ​​can improve the generative data-driven model's ability to understand different input languages ​​and generate operation instructions in different languages.

[0031] According to an example implementation of the method in accordance with the second aspect, the method further includes updating the factory-based training data and fine-tuning or retraining the trained generative data-driven model based on the updated factory-based training data. This allows the generative data-driven model to be continuously updated with new production data and the generated operational instructions to be improved.

[0032] According to the example implementation of the method according to the second aspect, the retraining or fine-tuning of the trained generative data-driven model is any one of scheduled retraining or fine-tuning, continuous retraining or fine-tuning, or trigger-based retraining or fine-tuning.

[0033] According to an example implementation of the method according to the second aspect, the trigger (e.g., for trigger-based fine-tuning) is based on thresholds and scores related to operational instructions provided by a trained generative data-driven model, and / or optionally, based on the number of retraining or fine-tuning cycles.

[0034] According to any example implementation, the operating instructions include machine-readable instructions for controlling and / or monitoring production.

[0035] According to another aspect, a computer program product is disclosed, which includes computer-readable instructions that, when executed on a computer, cause the computer to perform steps according to any example of the method according to the first or second aspect.

[0036] According to another aspect, the present invention discloses an apparatus comprising corresponding components for performing or carrying out steps of any method example according to the first or second aspect of the invention, or comprising at least one processor and at least one memory for storing instructions which, when executed by at least one processor, cause the apparatus to perform at least the steps of any method example according to the first or second aspect of the invention.

[0037] According to another aspect, the use of the operating instructions generated according to the method of the first aspect is disclosed, for displaying operating instructions to operators in a distributed production environment and / or for producing (e.g., chemical) products.

[0038] According to another example aspect, a computer element is disclosed that includes instructions that, when executed by a processor or computing device, perform or carry out steps according to the methods disclosed herein or defined by the devices disclosed herein.

[0039] According to another example aspect, a computer program or computer program product is disclosed that, when executed by a processor, causes a device, such as a server, to perform and / or control actions according to any aspect of the method.

[0040] According to another example aspect, a (e.g., tangible and / or non-transitory) computer-readable storage medium is disclosed, which includes a computer program that, when executed by a processor, causes a device, such as a server, to perform and / or control actions according to any aspect of the method.

[0041] According to another example aspect, a computer product is disclosed that includes a computer-readable token for accessing factory-based training data and / or any aspect of a trained or pre-trained or fine-tuned model.

[0042] Generative data-driven models (e.g., transformer-based models) can be trained on large datasets (e.g., non-specific text and image data). Trained generative data-driven models (such as transformer-based models) can have improved capabilities for predicting data patterns (such as patterns presented in natural language). This improvement in capability can be attributed to the large number of parameters acquired through the training. For example, transformer-based models (such as OpenAI GPT) can include 117 million parameters (GPT-1), 1.5 billion parameters, 175 billion parameters (GPT-3), and 170 trillion parameters (GPT-4). These parameters allow the GPT model to generate improved data outputs compared to other models that do not include transformer components (such as recurrent neural networks (RNNs) or long short-term memory (LSTM) networks).

[0043] The paper “Attention Is All You Need”, published by Vaswani et al. at the 31st Neural Information Processing Systems Conference (NIPS2017) in Long Beach, California (December 6, 2017, arXiv:1706.03762v5), describes a mechanism in machine learning that includes a transformer component (a transformer-based model), which is incorporated herein by reference.

[0044] ISO / IEC 23053:2022(en), ISO / IEC TR 24372:2021(en), ISO / IEC 22989, and ISO / IEC 23053 define standards in the fields of artificial intelligence (AI) and machine learning (ML). Big data is specified in ISO / IEC 20546:2019(en) Information technology—Big data. Data quality can be specified, for example, in ISO / IEC 20546:2019(en), ISO 8000-66:2021(en) / Data quality, and ISO / IEC DIS 5259-1(en).

[0045] Generative data-driven models, such as transformer-based architectures, can allow for capturing long-range dependencies and parallelization of computations. Furthermore, transformer-based architectures can be pre-trained on large text-based datasets and fine-tuned (or retrained) for task-specific datasets with smaller (labeled) datasets. Fine-tuning can be a process of taking a pre-trained generative data-driven model (e.g., trained on a large dataset) and further training it on a smaller, specific dataset, which can allow the knowledge learned by the pre-trained model to be transferred to a specific task. During fine-tuning, the model's weights can be updated based on the provided specific dataset, where the pre-trained weights serve as a starting point, and, for example, only a small number of additional training steps are performed.

[0046] The fine-tuning or retraining allows the superior analytical capabilities of pre-trained transformer-based models to analyze data patterns beyond those found in natural language.

[0047] This disclosure may relate to using generative converter-based models to analyze data patterns, and to using pre-trained converter-based models for distributed production, such as chemical production. Data suitable for retraining or fine-tuning a pre-trained converter-based model for said production can be obtained by providing historical production data accumulated over more than 150 years.

[0048] The amount of training data used to fine-tune a pre-trained transformer-based model can depend on the specific task, the model's complexity, and the desired performance level. In many cases, a smaller amount of specialized training data can be used to fine-tune (retrain) a transformer model compared to the size of the pre-trained dataset. However, if the application is significantly different from the pre-trained data, fine-tuning or retraining a pre-trained transformer-based model may require a much larger dataset compared to scenarios where the application is similar to the pre-trained data. For example, text classification or sentiment analysis may require hundreds of megabytes to gigabytes of labeled data. Machine translation may require tens to hundreds of gigabytes of text data for retraining. Question answering may require several gigabytes of training data to retrain or fine-tune a pre-trained transformer-based model.

[0049] Utilizing higher-quality training data can reduce the amount of training data required (for example, the term "data quality" is defined in ISO / IEC 20546:2019(en), ISO 8000-66:2021(en) / Data quality, and ISO / IEC DIS 5259-1(en)). Data preprocessing for generating high-quality training data can involve labeling data, removing noise, and removing irrelevant information. Therefore, preprocessing data to generate high-quality training data can facilitate more effective retraining or fine-tuning.

[0050] Larger amounts of training data allow for fine-tuning or retraining of pre-trained transformer-based models with more parameters. For example, fine-tuning the DistillBERT model may require less training data compared to fine-tuning the GPT-4 model. The larger number of parameters can also provide improved data analysis capabilities (e.g., GPT-4 is more powerful than GPT-3 in analyzing data patterns).

[0051] Generative data-driven models can be transformer-based models, such as TinyBERT, DistilBERT, Llama 7B, Mistral 7B, GPT-Neo, or larger GPT variants, or other models, such as structured spatial state models. Furthermore, pre-trained transformer-based models can be, for example, ChatGPT (GPT-2, GPT-3, GPT-3-5-turbo, GPT-4), Davinci, BERT (bidirectional encoder representation from a transformer), DistilBERT, Transformer-XL, XLNet (Extreme Language Understanding Network), T5 (text-to-text transfer transformer), RoBERTa (robustly optimized BERT method), ELECTRA (efficient encoder for learning accurate token substitutions), Reformer, Longformer, and DeBERTa (decoding-enhanced BERT with disentangled attention). Due to differences in architecture and / or pre-trained datasets, the properties of these transformer-based models, and therefore the output data, may differ for the same input data. Therefore, one or more of the transformer-based models (GPT-2, GPT-3, GPT-3-5-turbo, GPT-4, Davinci, BERT, DistilBERT, Transformer-XL, XLNet, T5, RoBERTa, ELECTRA, Reformer, Longformer, DeBERTa) can be used as an alternative to a pre-trained transformer-based model for technical purposes, or one or more of the models can be combined for technical purposes to provide multiple data outputs for supplementary data analysis.

[0052] Specifically, training data provided for retraining or fine-tuning a pre-trained generative data-driven model used for production (such as chemical production) may include historical, plant-based data. For example, production data stored for over 150 years (plant-based data) could be used. Plant-based training data required for retraining or fine-tuning a pre-trained transformer-based model may include historical production data (plant-based data) from one or more databases from one or more distributed production facilities.

[0053] Furthermore, the method according to the first aspect may include retraining or fine-tuning a pre-trained converter-based model to provide (release) a trained converter-based model suitable for production. The trained converter-based model can then be used to control and / or monitor distributed production environments, such as chemical production.

[0054] Furthermore, the method according to the first aspect may also include providing a transformer-based model, and pre-training the model on general data such as text-based data (any type of data that can be converted into text data). The pre-trained data may also include publicly available data associated with production, such as data available on the website of the production facility.

[0055] Another large language model (e.g., a base model such as a transformer-based model) can also be used, which can be pre-trained on a large dataset (e.g., a non-specific, publicly available dataset) and trained (fine-tuned, retrained) on a smaller dataset (e.g., a specific dataset, such as production data) to make the trained model suitable for controlling and / or monitoring production.

[0056] Pre-trained transformer-based models, which are pre-trained on larger datasets (publicly available data, non-specific data) and fine-tuned or retrained on smaller datasets (i.e., factory-based datasets), are superior in analyzing factory-based data patterns, identifying anomalies in data patterns, and generating operational instructions based on the analysis of data patterns to improve the control and / or monitoring of distributed production (such as chemical, pharmaceutical, and biotechnology production).

[0057] Providing factory-based training data enables fine-tuning or retraining of pre-trained transformer-based models for production.

[0058] Preprocessing raw factory data (e.g., classifying, filtering, labeling, or structuring data) to generate factory-based training data improves the quality of the training data and thus increases training efficiency, since retraining or fine-tuning a pre-trained transformer-based model may require less training data and computational power.

[0059] Instead of or in addition to the claimed pre-trained transformer-based model, one or more alternative pre-trained transformer-based models may be used, provided that the one or more models are fine-tunable or retrainable for production use. The one or more alternative or additional models may have different properties based on different sets of parameters obtained during pre-training. Therefore, the one or more additional or alternative models may be supplementary or alternative components for analyzing data patterns in factory-based input data, according to different aspects of the disclosure provided herein.

[0060] Triggered, scheduled, or continuous retraining or fine-tuning of a trained transformer-based model can improve its performance because it can be trained on up-to-date data. Feedback scores can be provided as a quality metric for the generated output data.

[0061] Using trained transformer-based models to control and / or monitor production, such as in chemical production, can improve overall production efficiency because the models can identify many aspects leading to inefficient production by analyzing data patterns in the input data based on plant data analysis. Based on the analysis, the model can generate solutions on how to improve production in the format of operational instructions (machine-readable instructions) for controlling and / or monitoring production. Reviewing these operational instructions ensures that the model can be securely integrated into distributed production environments such as chemical production. Attached Figure Description

[0062] The present disclosure is further described below with reference to the accompanying drawings. The same reference numerals in the drawings and the present disclosure are intended to refer to the same or similar elements, components, and / or parts.

[0063] Figure 1 This illustrates a distributed production environment, such as one or more chemical plants.

[0064] Figure 2A shows Figure 1 The operating system for a distributed production environment.

[0065] Figure 2B shows the situation caused by... Figure 1 The pre-trained transformer-based model is trained on factory data generated in the production environment.

[0066] Figure 2C illustrates the use of a trained transformer-based model for control and / or monitoring. Figure 1 The implementation scheme of the distributed production environment shown is illustrated.

[0067] Figure 3 The preprocessing engine of the operating system shown in Figure 2 is illustrated.

[0068] Figure 4 It shows the result of Figure 1 The factory data structure generated in the production environment.

[0069] Figure 5 The data structure derived from the operator is shown.

[0070] Figure 6 It shows that, in addition to Figure 2a and Figure 2b Additional details beyond the implementation plan include retraining or fine-tuning the pre-trained transformer-based model used for production.

[0071] Figure 7 The contextualization of the prompt is shown.

[0072] Figure 8 This demonstrates trigger-based retraining or fine-tuning of a trained transformer-based model when it is used in production.

[0073] Figure 9 An implementation scheme for training the embedding layer is shown.

[0074] Figure 10A An implementation scheme of the converter encoder architecture is shown.

[0075] Figure 10B An implementation scheme of the converter decoder architecture is shown.

[0076] Figure 10C An implementation scheme of the converter encoder-decoder architecture is shown.

[0077] Figure 11 An implementation scheme for training and / or deploying a transformer encoder, transformer decoder, and / or transformer encoder-decoder is shown.

[0078] Figure 12 An implementation scheme for input embedding is shown. Detailed Implementation

[0079] The following embodiments are merely examples for implementing the methods, systems, devices, or application apparatuses disclosed herein and should not be considered limiting. The following description is intended to enhance understanding and should be understood as supplementing and reading in conjunction with the descriptions provided in the foregoing Summary and Embodiments section of this specification. Some aspects may have different terminology than, for example, that is provided in the above description. However, those skilled in the art will understand that these terms refer to the same subject matter, for example, in a more specific manner.

[0080] In the context of this application, a generative data-driven model (generative artificial intelligence AI model) can refer to, for example, a model including transformer components such as... Figures 9 to 12 The basic model described in the text (machine learning ML model).

[0081] Generative artificial intelligence (AI) can refer to computer programs that can generate outputs, as described in ISO / IEC 23053:2022(en), ISO / IEC 23053:2022(en), ISO / IEC TR 24372:2021(en), ISO / IEC 22989, ISO / IEC 23053, ISO / IEC DIS 5259-1(en), and ISO / IEC 24661:2023(en). Generative AI programs can include ML models, such as transformer-based models (generative pre-trained transformer GPT models). Generative pre-trained transformer models (or simply transformer-based models) can also be referred to as base models.

[0082] Figures 1 to 8In the context of “engine”, it includes at least one computer processor.

[0083] In the context of this disclosure, "computer interface" can be, for example, a graphical user interface, an application programming interface, or a web-based interface.

[0084] For example, a pre-trained transformer-based model can be pre-trained for a first purpose (e.g., analyzing non-specific text-based data) and retrained / fine-tuned for a second purpose (e.g., production). A pre-trained generative data-driven model (e.g., a transformer-based model) can be further retrained / fine-tuned for the second purpose to better suit the second purpose, thereby improving the output data provided by the model.

[0085] A “distributed production environment” or “plant” can refer to, but is not limited to, any technological infrastructure used for industrial purposes of manufacturing, producing, or processing one or more products (i.e., manufacturing or production processes or processing performed by a distributed production environment). A distributed production environment can be a “plant” (infrastructure) having distributed units for production. A distributed production environment can be a technological infrastructure (plant) that includes distributed operations. A distributed production environment can be more than one plant distributed in space and / or for distributed operations. A distributed production environment can be one or more of the following: chemical plant, processing plant, pharmaceutical plant, fossil fuel processing facility (such as oil wells and / or natural gas wells), refinery, petrochemical plant, cracking plant, etc. A distributed production environment can even be any of the following: distillery, processing plant, or recycling plant. A distributed production environment can be any of the examples given above or a combination of their analogues.

[0086] "Products" produced in a distributed production environment can be, for example, any physical product, such as chemicals, biological products, pharmaceuticals, food, nutritional products, beverages, textiles, metals, plastics, semiconductors, cosmetics, or even any combination thereof. Additionally, or alternatively, products can be service products, such as those resulting from recycling or waste disposal (e.g., recycling), or chemical processing (e.g., decomposition or dissolution into one or more chemical products). Some non-limiting examples of chemical products are organic or inorganic compositions, monomers, polymers, foams, pesticides, herbicides, fertilizers, feed, nutritional products, precursors, pharmaceutical or therapeutic products, or any or more of their components or active ingredients. In some cases, chemical products can be products that can be used by end-users or consumers, such as cosmetic or pharmaceutical compositions. Chemical products can be products that can be used to further manufacture one or more products; for example, a chemical product can be a synthetic foam that can be used to manufacture shoe soles or a coating that can be used on the exterior of automobiles. Chemical products can be in any form, such as solid, semi-solid, paste, liquid, emulsion, solution, granules, particles, or powder.

[0087] Distributed production environments may include equipment (e.g., chemical equipment) or process units such as any one or more of the following: heat exchangers, towers (such as fractionation towers), furnaces, reaction chambers, cracking units, storage tanks, extruders, granulators, settlers, agitators, mixers, cutters, curing tubes, evaporators, filters, sieves, pipes, chimneys, valves, actuators, mills, transformers, conveying systems, circuit breakers, machinery (e.g., heavy rotating equipment such as turbines, generators, crushers, compressors, industrial fans, pumps), conveying elements (such as conveyor systems), motors, etc.

[0088] Furthermore, a distributed production environment typically includes multiple sensors and at least one control system for controlling at least one parameter related to a process or process parameter within the plant. Such control functions are typically performed by the control system or controller in response to at least one measurement signal from at least one of the sensors. The plant's controller or control system can be implemented as a distributed control system (DCS) and / or a programmable logic controller (PLC). Multiple sensors can be distributed throughout the distributed production environment for monitoring and / or control purposes. Such sensors can generate large amounts of data. Sensors may or may not be considered part of the equipment. Therefore, production such as chemical and / or service production can be a data-intensive environment. Distributed production environments can generate large amounts of process-related data.

[0089] The sensors can be used to measure one or more process parameters and / or to measure the operating conditions of the equipment or parameters related to the equipment or process unit. For example, sensors can be used to measure process parameters such as flow rate in a pipe, liquid level in a tank, furnace temperature, chemical composition of a gas, etc., and some sensors can be used to measure vibration of a crusher, fan speed, valve opening, pipe corrosion, voltage across a transducer, etc. The differences between these sensors are not only based on the parameters they sense, but can even be based on the sensing principles used by the respective sensors. Some examples of sensors based on the parameters sensed by the sensor can include: temperature sensors, pressure sensors, radiation sensors such as light sensors, flow sensors, vibration sensors, displacement sensors, and chemical sensors such as those for detecting specific substances such as gases. Examples of sensors that differ in the sensing principles employed can be, for example: piezoelectric sensors, piezoresistive sensors, thermocouples, impedance sensors such as capacitive sensors and resistive sensors, etc.

[0090] A distributed production environment can consist of multiple distributed production environments. These environments can be coupled, allowing them to share one or more of their value chains, derivatives, and / or products. Multiple distributed production environments can also be referred to as a complex, a composite site, an "integrated" or "integrated site." Such an integrated site or chemical industrial park can be or may include one or more distributed production environments, where products manufactured in at least one distributed production environment can serve as raw materials for another.

[0091] "Production" refers to any industrial process that provides an output product different from the input product when used on or applied to an input component. Therefore, production can be any manufacturing or processing process or a combination of processes to obtain the aforementioned product. A production process can even include the packaging and / or stacking of one or more products.

[0092] Production processes can be continuous or periodic; for example, a batch chemical production process may be used when a catalyst that requires recovery is employed. A key difference between these production types lies in the frequency of occurrence of data generated during production. For instance, in a batch process, production data extends from the beginning of the process to the last batch of different batches produced during that run. In a continuous setup, data is more continuous due to potential changes in production operations and / or maintenance-driven downtime. Consequently, the required data analysis can vary depending on whether the data stream is batch or continuous. For example, periodic retraining or fine-tuning of a trained generative data-driven model (e.g., a transformer-based model) may be advantageous for batch data streams, while continuous retraining or fine-tuning of a trained generative data-driven model (e.g., a transformer-based model) may be advantageous for continuous data streams.

[0093] The terms "plant data" or "plant-based data" are used interchangeably and can refer to production data (e.g., product attributes), process data (process parameters), and operating conditions. Plant-based data can refer to data including values ​​(e.g., numerical or binary signal values) measured during a production process, such as via one or more sensors. Process data can be time-series data of one or more process parameters and / or equipment operating conditions. Typically, plant-based data can include temporal information about process parameters and / or equipment operating conditions; for example, the data may contain timestamps for at least some data points related to process parameters and / or equipment operating conditions. Plant-based data can include temporal-spatial data, i.e., time data and location or data associated with one or more physically separated equipment areas, allowing temporal-spatial relationships to be derived from the data.

[0094] "Process parameters" can refer to any variable related to the production process, such as any or more of the following: temperature, pressure, time, level, etc., which are related to the production of the product as defined above.

[0095] The above definitions of distributed production environment, products produced by distributed production environment, production process, data generated by production environment, and production control are merely illustrative and should not be construed as restrictive. It is understood that the systems and methods disclosed herein can be applied to any kind of production that produces products and generates production-related multi-parameter data streams. Any type of factory-based data can be preprocessed by a preprocessing unit to generate the required factory-based training data (labeled data, (pre-)structured data, filtered data, data in numerical and / or text formats, etc.) or factory-based input data, respectively suitable for retraining or fine-tuning a pre-trained generative data-driven model (e.g., a transformer-based model) or for production using a trained generative data-driven model. Therefore, a distributed production environment should be broadly interpreted as a technological environment for producing products (physical products and / or services associated with the products; products can even be data products), and, during the production of said products, generating multi-parameter production-related data streams (factory-based data) through said technological environment.

[0096] Figure 1 This illustrates a distributed production environment, such as one or more chemical plants.

[0097] The distributed production environment may include equipment 1002 and sensors 1003 that generate one or more sensor-related data streams. The distributed production environment may produce one or more products as defined above, wherein attributes of the one or more products can be measured, extracted, or calculated to generate one or more product-related data streams. Plant data 1005 (plant-based data) may include data obtained from each of the one or more data streams.

[0098] Equipment 1002 can be any equipment in a distributed production environment, such as pumps, heat exchangers, valves, reaction vessels, separation chambers, etc.

[0099] Sensor 1003 can be any type of sensor in a distributed production environment, such as temperature sensors, flow sensors, pressure sensors, etc.

[0100] One or more products produced in a distributed production environment can be any type of product as described above. Product properties can be measured, for example, by gas chromatography.

[0101] Plant data 1005 can be stored in a database, for example, as historical data. Plant data 1005 can be provided to a preprocessing engine, which preprocesses the data and provides plant-based input data to an analysis engine 1004. This analysis engine can analyze the plant-based input data and generate machine-readable instructions 1007 for controlling and / or monitoring the engine 1006. The generation of machine-readable instructions 1007 can be performed automatically by the analysis engine 1004 based on the analysis of the input plant data (i.e., without operator intervention). For example, the analysis engine 1004 can continuously receive input plant data and analyze the data in a continuous mode. When an anomaly occurs, the analysis engine can generate machine-readable instructions 1007 for the control and / or monitoring engine to remove the anomaly. The analysis engine can determine solutions to improve production efficiency by analyzing plant input data in the context of the production process and can send machine-readable instructions 1007 to the control and / or monitoring engine to improve production. The control and / or monitoring engine 1004 can display push notifications to the operator to review instructions 1007, and based on this review, can generate machine-readable instructions 1008. Based on the machine-readable instructions 1008, the control system can change the operating parameters of one or more pieces of equipment 1002. Reviewing the machine-readable instructions can allow trained transducer-based models to be safely integrated into distributed production environments such as chemical production.

[0102] Alternatively, in addition to the machine-readable instructions 1007 automatically generated by the analysis engine, the operator may also prompt the analysis engine 1004 based on prompts (e.g., such as...). Figure 7 (as shown) to provide machine-readable instruction 1007.

[0103] Machine-readable instruction 1007 can be used by an operator to control and / or monitor one or more production operations in a distributed production environment.

[0104] Machine-readable instructions 1007 may include operational instructions for production, such as machine-readable instructions for controlling equipment 1002 and / or machine-readable instructions for monitoring equipment 1002 and / or sensors 1003. Machine-readable instructions 1007 may also include operational instructions for operators controlling and / or monitoring a distributed production environment, wherein the operator may be a human-based operator, a computer-based operating system, or a hybrid system including human operators and computer-assisted systems.

[0105] The operator can review the machine-readable instructions 1007, and based on that review, can generate machine-readable instructions 1008 for controlling equipment 1002 and / or sensors 1003. The operator can then, based on... Figure 7The prompts, queries, context, etc. shown are used to prompt the analysis engine 1004 to provide operation instructions (machine-readable instructions 1007).

[0106] The operator may be a human operator reviewing the machine-readable instructions 1007. The operator may be a human operator with a computer-based assisted system for reviewing the machine-readable instructions 1007. The operator reviewing the machine-readable instructions 1007 may be a computer-based system based on a computer program that includes a set of instructions for reviewing the machine-readable instructions 1007.

[0107] A control system for a distributed production environment may include one or more computing units that can manipulate one or more parameters related to the production process by controlling one or more actuator or switch and / or end effector units, such as by manipulating one or more equipment operating conditions. Control is typically performed in response to one or more signals retrieved from the equipment.

[0108] The control and monitoring engine may include one or more computer processors for modifying machine-readable instructions 1007 and generating machine-readable instructions 1008 for a control system used in distributed production. The control system for distributed production can adjust equipment operating conditions based on the machine-readable instructions 1008, such that the adjusted process parameters and / or equipment operating conditions produce a controlled product (such as a chemical product) with one or more desired or predetermined properties or performance parameters. Therefore, production can be controlled in real time while ensuring that equipment operating conditions adapt to undesirable changes in process parameters.

[0109] It is understood that the control and monitoring of a distributed production environment generally involves controlling the equipment and / or production lines used to produce products by sending machine-readable instructions to the production environment. "Product" should be interpreted broadly as described above. As another example, a product could even be a data product and / or data service product provided by a trained transformer-based model trained on factory-based training data.

[0110] As in Figure 2a , Figure 2b , Figures 9 to 12 The generative data-driven model or transformer-based model (first ML model) described in the context of the AI ​​engine can be integrated with another computer program (such as second ML model) via a computer interface (e.g., API, GUI, web application), wherein the second ML can include different algorithms (e.g., classic ML that is not based on a transformer architecture, a transformer-based variant of the first model, etc.) and / or the second ML can include the same architecture as the first model, but the second model can be trained on a different dataset.

[0111] For example, the second ML model can be a data-driven model that can be integrated with the first model (transformer-based model) via an API, GUI, or web-based interface. The second ML model can be used to preprocess the raw factory-based data to generate factory-based training data. Preprocessing of the raw data may involve noise removal, data filtering, data labeling, data sorting, converting the raw data to different formats, converting operator data to different formats more suitable for the first ML model (transformer-based model), and so on. This preprocessing of the raw data can also be performed by a computer program based on a set of computer-implemented instructions (e.g., a set of programming instructions, filters, labels, mathematical steps, etc., excluding the ML model). Preprocessing of the raw factory-based data can improve the quality of training or input data based on the raw data and can reduce the computational power required to train / use the first ML model (e.g., the transformer-based model).

[0112] Figure 2A shows Figure 1 The operating system is for a distributed production environment. The operating system includes (re)training a pre-trained converter-based model and using the trained converter-based model to control and / or monitor production, such as chemical production. The training of the pre-trained converter-based model is further described in the context of Figure 2B. The purpose of using the trained converter-based model is further described in the context of Figure 2C.

[0113] Figure 2B shows the situation caused by... Figure 1 The pre-trained transformer-based model is trained on factory data generated in the production environment.

[0114] Such as in Figures 9 to 12 As described in the context, training (fine-tuning, retraining) a pre-trained converter-based model includes accessing the pre-trained converter-based model via a computer interface (GUI, API, web-based interface) and receiving factory history data 1010 from a preprocessing engine 1009 via the interface. An operator retraining or fine-tuning the pre-trained converter-based model may prompt the pre-trained converter-based model to access the factory history data via the computer interface, or alternatively, the operator may upload historical factory data from a database to an AI engine via the computer interface, the AI ​​engine including at least one processor for operating the pre-trained converter-based model. The operator retraining or fine-tuning the pre-trained model may be a human operator, an automated operating system including a computer processor, or a hybrid operating system including both human operators and computer-aided operating systems, prompting the pre-trained model to train, fine-tune, or retrain on factory-based data.

[0115] Factory historical data 1010 can be based on factory data 1005. Factory data 1005 and factory historical data can be stored in a database. Factory data 1005 may include, for example, data from... Figure 4 This describes multiple production data points within the context of [the data]. During training, historical factory data can be embedded via an embedding layer, such as [example data]. Figure 9 The context described above. Embedded factory data can generate embedded factory data.

[0116] The examples of factory data described above are merely illustrative of possible ways to implement the methods disclosed herein. Factory data should be interpreted broadly and should be understood as any kind of data associated with the production of products by technological infrastructure. Data of any type or format associated with production can be preprocessed by a preprocessing engine to be suitable for training or production using transformer-based models.

[0117] At the end of the training cycle, the analytics engine 1004 can output (publish) a trained transformer-based model suitable for use in distributed production environments such as chemical production. The published trained model can be stored in a database for purposes such as version control. The published trained model can be a computer program product. Access to the published trained model can be provided to users as a data service to aid production.

[0118] The training data for factory-based training can be factory-based data in one or more languages. Training a pre-trained transformer-based model on factory-based training data in one or more languages ​​allows for the expansion of the factory-based training dataset and also allows for the provision of operational instructions for production in one or more languages. Providing instructions in one or more languages ​​can improve user interaction with the trained transformer-based model.

[0119] Figure 2C illustrates the use of a trained transformer-based model to control... Figure 1 The implementation scheme for the distributed production environment is shown.

[0120] Factory data 1005 can be provided to preprocessing engine 1004 for preprocessing. Preprocessed data 1005 may include, for example, data from... Figure 3 The steps described in the context. Preprocessing may also involve other steps required to provide training / input data based on factory data suitable for training or using generative data-driven models (such as transformer-based models), such as labeling data, removing noise, constructing unstructured data, converting data into different formats, etc., for production.

[0121] Preprocessed data from preprocessing engine 1004 can be used as factory-based input data for a trained generative data-driven model for production.

[0122] A trained generative data-driven model can receive factory-based input data and, for example, predict anomalies in the factory-based input data. Based on the analysis of the factory-based input data, the trained generative data-driven model can, for example, predict how to resolve errors in factory operations, how to improve production efficiency, and how to resolve user queries related to controlling and / or monitoring production. Based on the predictions, the analysis engine 1004 can generate machine-readable instructions 1007 for use in controlling and / or monitoring the engine 1006.

[0123] The analysis engine 1004, which operates a pre-trained transformer-based model, may have at least one computer interface (e.g., a graphical user interface (GUI), a web-based computer interface, and / or an application programming interface (API)) for operating the transformer-based model (uploading data, providing prompts (such as prompts including instructions for accessing factory-based training data and / or input factory-based data), providing prompts for retraining or fine-tuning, providing prompts for generating machine-readable instructions 1007, a component for reviewing machine-readable instructions 1007, a component for receiving notifications when new machine-readable instructions are generated, etc.). The (untrained, pre-trained, trained) model may be stored in a database or in the cloud. The model may be operated by the AI ​​engine via at least one computer interface (e.g., a graphical user interface (GUI), a web-based computer interface, and / or an application programming interface (API)).

[0124] Generative data-driven models, such as transformer-based models (e.g., any of pre-trained, trained, and / or retrained generative data-driven models), can be provided to the AI ​​engine via a computer interface. In other words, one or more generative data-driven models (e.g., transformer-based models) can be accessed by the AI ​​engine. For example, the AI ​​engine (e.g., at least one computer processor) can access the transformer-based model via the computer interface (e.g., API, user interface, web interface). Alternatively, the transformer-based model can be part of the AI ​​engine (e.g., stored in the AI ​​engine's computer memory), or the transformer-based model can be integrated with the AI ​​engine (e.g., stored in a database accessible to the AI ​​engine).

[0125] An operator of a trained transformer-based model can provide input (e.g., input factory-based data) to a trained generative data-driven model (e.g., a trained transformer-based model) via a computer interface (e.g., a graphical user interface, an application programming interface, a web-based interface), such as text and / or audio queries, as shown in [example code]. Figure 5 As described in the context, an operator of a trained generative data-driven model (e.g., a trained transformer-based model) can additionally provide extracted data points, for example, from data provided by a preprocessing engine. These extracted data points may, for example, relate to outlier data points in a distributed production environment, typical or optimal operating parameters. The extracted data points, together with the operator's query, can form part of the input data to the trained generative data-driven model.

[0126] A trained generative data-driven model (e.g., a trained transformer-based model) can process input data and provide solutions to user queries, for example, in the form of machine-readable instructions 1007.

[0127] The control and / or monitoring engine can review machine-readable instructions 1007.

[0128] Following review, the control and / or monitoring engine 1006 can generate machine-readable instructions 1008, which may be identical to, partially based on, or different from machine-readable instructions 1007. The operator of the control and / or monitoring engine can use the machine-readable instructions 1007 solely for monitoring production, or can forward the machine-readable instructions 1007 into machine-readable instructions 1008 for controlling production. The operator can be a human operator and / or an operating system including the processor, and optionally a human operator.

[0129] If the machine-readable instruction 1008 generated by the control and / or monitoring engine differs from the machine-readable instruction 1007 generated by the trained transformer-based model, the control and / or monitoring engine may trigger retraining or fine-tuning of the trained generative data-driven model (e.g., the trained transformer-based model). Retraining or fine-tuning may be automatically triggered based on feedback from the control and / or monitoring engine indicating deviations in the machine-readable instructions 1007 and 1008. The feedback may include a feedback score indicating the degree of deviation.

[0130] Retraining or fine-tuning can be triggered by the operator. Alternatively, retraining or fine-tuning can be scheduled or continuous.

[0131] A trained generative data-driven model can be trained on factory-based training data in one or more languages. The trained generative data-driven model can provide operation instructions for production in one or more languages. Providing operation instructions in one or more languages ​​may be advantageous, for example, because the training data may be more available in one language than in another. A trained generative data-driven model trained on factory-based training data in one or more languages ​​(e.g., a transformer-based model) can be configured to provide operation instructions in the language where most of the training data is available. Alternatively, the trained generative data-driven model can be configured to provide operation instructions for production in a user-selected language, making the model more user-friendly. Alternatively, the trained generative data-driven model (e.g., a transformer-based model) can be requested to provide operation instructions for production in more than one language to cross-check the operation instructions and select the most suitable instruction for improved production.

[0132] Figure 3 The preprocessing engine of the operating system is shown in Figure 2.

[0133] As in Figure 4 The raw factory data 1005 described in the context can be provided to the preprocessing engine 1009 via a computer interface.

[0134] The preprocessing engine 1009 can preprocess raw factory-based data to provide factory-based training data and / or factory-based input data to a generative data-driven model (e.g., a transformer-based model). Input data can be stored in a database. Input data can be provided to the generative data-driven model via a computer interface (e.g., by prompting the model to access the data). Preprocessing steps may include selecting desired parameters, merging / aggregating, computing factory-based training data (e.g., computing derived parameters), removing outliers, etc. Preprocessing may include filtering data, removing noise, labeling data, sorting data, and converting data from a format unsuitable for training / using the model to a format suitable for training / using the model, etc. The output data of the preprocessing engine can be stored in a database and used as factory-based training data for retraining or fine-tuning a pre-trained generative data-driven model (e.g., a transformer-based model), such as... Figure 2a , Figure 2c , Figure 6 and Figures 9 to 12 As described in the context. The output data of the preprocessing engine can also be used as input data for trained generative data-driven models (e.g., transformer-based models), such as... Figure 1 , Figure 2a , Figure 2c and Figure 7As shown.

[0135] Figure 4 It shows the source such as Figure 1 The factory data structure shown is for the factory.

[0136] Plant data can be received from a distributed production environment via computer interfaces (e.g., graphical user interface, application programming interface, web-based interface). Plant data can include different categories of plant data, such as sensor data, operational data, plant metadata, and analytical data.

[0137] Sensor data can refer to quantities that are available for measurement in a production plant using installed sensors (such as temperature sensors, pressure sensors, flow rate sensors, etc.).

[0138] The analytical data may involve quantities provided by analytical measurements of samples taken from any point in the production plant, such as the composition of reactants, starting materials, products and / or byproducts determined, for example, by gas chromatography from samples taken from different stages of the production process (e.g., before or after the catalytic reactor).

[0139] Operational data may involve raw data (basic data, unprocessed analytical data, and / or sensor data) or processed or derived parameters (derived directly or indirectly from raw data).

[0140] Plant metadata can indicate the physical plant layout and can include plant-specific quantities that describe attributes such as reactors. These plant-specific quantities are predefined by the physical plant layout and can be related to plant or reactor performance.

[0141] Factory data 1005 can include text and / or numbers (structured data). Factory data can also be unstructured. Unstructured data (such as scans, data tables including images, QR codes, etc.) can be preprocessed by a preprocessing engine and transformed into a format suitable for training, such as... Figures 9 to 12 The text and / or numbers described in the context of a pre-trained transformer-based model.

[0142] Figure 5 The data structure derived from the operator is shown.

[0143] Users (e.g., operators in a production environment) can input queries into the analytics engine, such as text or audio queries regarding the monitoring and control of distributed production. Queries can be, for example, requests for operational instructions for production, such as operational instructions related to controlling and / or monitoring one or more production operations in the distributed production environment. For instance, an operator might request a trained transducer-based model to provide operational instructions for resolving anomalies in plant data, correcting errors in plant operations, resolving equipment operation problems, or providing steps for performing tasks related to one or more operations in the distributed production environment. The analytics engine can predict solutions to queries based on input from the operator and input plant-based data. The solution may include machine-readable instructions 1007 (operational instructions), such as changing operating parameters of equipment 1002, replacing equipment 1002 and / or sensor 1003, performing maintenance, performing one or more steps related to one or more production operations, etc.

[0144] Figure 6 It shows that, in addition to Figure 2a and Figure 2b Additional details beyond the implementation plan include retraining or fine-tuning the pre-trained transformer-based model used for production.

[0145] Pre-trained transformer-based models (such as those pre-trained on text and / or numbers) can be trained based on historical factory data (factory-based training data). Figures 9 to 12 (As shown in the diagram) for use in production. Specifically, the converter-based model can be as shown in... Figure 2b and Figure 11 The transformer encoder, transformer decoder, and / or transformer encoder-decoder are trained and / or deployed as described in the context of the training and / or deployment described in the context of the transformer encoder-decoder. This can be followed as... Figures 9 to 12 The training steps described in the context of [the previous sentence] are to use factory-based data instead of [the previous sentence]. Figures 9 to 12 The generic text / numbers shown are used to fine-tune or retrain the model for production. The trainer of the pre-trained transformer-based model can create prompts to trigger fine-tuning or retraining of the pre-trained transformer-based model, and can instruct or upload factory-based training data based on factory data 1005 used for training, for example, via a computer interface (e.g., a web-based interface, API, GUI). At the end of the training cycle, the model trainer can, as... Figure 2a and Figure 2c The trained model is tested as described in the context, and based on the test, the trainer can run additional training cycles or release the trained model for production.

[0146] The published, trained transformer-based model can be stored in a database or in the cloud. The published model can be operated by computing units / nodes, including computer processors (such as analytics engines). A copy of the trained transformer-based model can be made available to users for use and / or further training. Alternatively, users can be given only access to operate the trained transformer-based model.

[0147] Replacement Figures 9 to 12 In the context described, the pre-trained transformer-based model can be used by the analytics engine 1004 with another base model, such as ChatGPT (GPT-2, GPT-3, GPT-3-5-turbo, GPT-4 or higher / similar), Davinci, BERT (bidirectional encoder representation from transformers), DistilBERT, Transformer-XL, XLNet (Extreme Language Understanding Network), T5 (text-to-text transfer transformer), RoBERTa (robustly optimized BERT method), ELECTRA (efficient encoder for learning accurate classification of token replacements), Reformer, Longformer, DeBERTa (decoding-enhanced BERT with disentangled attention), or any other large-scale language model pre-trained on big data such as general text, images, and videos.

[0148] Depending on the availability of factory-based training data, the availability of computing resources, and the required accuracy for analyzing the data, users can choose a base model with fewer or more parameters. Figures 9 to 12 The transformer-based model shown can be pre-trained to publish a pre-trained model with the required number of parameters to meet the user's technical objectives. Figure 2b , Figures 9 to 12 In the context of [the context], pre-trained models can be further improved to release models with even fewer parameters, thereby increasing computational speed and reducing computational resources.

[0149] Figure 7 The contextualization of the prompt is shown.

[0150] Users (operators in the production environment) can receive prompts via computer interfaces (e.g., graphical user interfaces, application programming interfaces, web-based interfaces) such as... Figure 2a , Figure 2b and Figure 6The context described is a trained transducer-based model used to parse user queries related to how to control / monitor production. Users can also provide context for queries, such as sensor data, operational data, plant metadata, and analytics data. Users can select the context via a computer interface or input the context via text and / or audio channels in one or more languages. Users can provide a second context. The second context may include, for example, one or more keywords (e.g., anomaly, error associated with a piece of equipment X, error in the software operating a piece of equipment X), one or more data points from plant data (e.g., one or more anomalous data points), one or more standard operating parameters, etc.

[0151] Based on one or more contexts, and optionally, based on data from the input plant, a user can prompt a trained transformer-based model to predict a solution to a user query. The predicted solution may include machine-readable instructions 1007 sent to the control and / or monitoring engine 1006 for operating the production environment 1001.

[0152] The operator can provide one or more contexts in one or more languages, and prompt the trained transformer-based model to provide operational instructions for production in one or more languages.

[0153] Figure 8 This demonstrates trigger-based fine-tuning or retraining of a transformer-based model trained for production use.

[0154] Once such Figure 2b , Figure 6 and Figure 11 Trained as described in the context, the trained transducer-based model can be further retrained or fine-tuned based on feedback related to the operational instructions provided by the trained transducer-based model. For example, an operator can monitor machine-readable instructions 1007 provided by the trained transducer-based model to the control and / or monitoring engine 1006. The operator can also assign feedback scores and thresholds related to the operational instructions provided by the trained transducer-based model. Retraining or fine-tuning can be automatically triggered by a computer processor (e.g., a control / monitoring engine or analysis engine) if the score falls below the threshold, or training can be triggered by a human operator. For example, the score may fall below the threshold when machine-readable instructions 1007 are unacceptable to the operator for the control equipment 1002 (e.g., when machine-readable instructions 1007 conflict with the optimal operating conditions of said equipment 1002).

[0155] Retraining or fine-tuning cycles may include, for example Figure 2b , Figure 6 and Figures 9 to 11The steps are illustrated in the context of the training data, which is based on factory data. After a fine-tuning or retraining cycle, the operator can provide new feedback scores for new machine-readable instructions 1007. Training can continue until the feedback scores reach a threshold or the maximum number of training cycles. At the end of the fine-tuning or retraining, the model can be stored and / or published for production use.

[0156] Retraining can be performed in the context of ongoing production without stopping / interrupting production.

[0157] Figure 9 An implementation scheme for training the embedding layer is shown. The embedding layer can be obtained by training, for example, a Continuous Bag-of-Words (CBOW) model or a skip-gram model. The embedding layer can be adapted to generate embedded input data based on input data. Generating embedded input data can refer to embedding input data. Embedding input data can produce a representation associated with the input data. Therefore, embedded input 114 can be a representation associated with the input data. The input data can include one or more elements. One or more elements can be represented by input vector 106. In particular, embedded input 114 and / or input vector 106 can be machine-readable and / or processable by a processor. For this purpose, embedded input 114 and / or input vector 106 can be tensors, particularly first-order tensors. Specifically, input vector 106 can be a one-hot vector or the sum of multiple one-hot vectors. A one-hot vector can be a vector with one entry that is not equal to zero. Examples of one-hot vectors can be 108, 110, and 112. The non-zero entries in the one-hot vector and / or input vector 106 can indicate elements. For example, a lookup table can define the relationship between the positions of non-zero entries and elements indicated by one-hot vectors. A lookup table can specify multiple distinct elements. The number of distinct elements can be equal to the number of entries in the one-hot vector. The number of distinct elements can be referred to as the vocabulary. In one example, elements can be represented by tokens, and a sequence of elements can refer to at least a portion of a sentence. At least a portion of a sentence can be represented by multiple tokens. Tokens can represent at least a portion of an element and / or a word. For example, in the case where an element will be associated with only one word, words such as “embeddings,” “embedding,” or “embed” will constitute distinct elements. A first token can represent the stem “embed,” and suffixes that typically appear in multiple words can be represented by a second, third, and fourth token. The second, third, and fourth tokens can be used to represent other words, such as “look,” “looking,” etc., preferably together with a fifth token representing the stem “look.” Ultimately, this tokenization of elements associated with multiple stems and multiple suffixes results in using fewer tokens to represent multiple elements, and therefore using fewer computational resources.

[0158] A lookup table specifying, for example, a subset of the English vocabulary can include 10,000 or more words. The embedded input 114 can be a lower-dimensional representation than the input vector 106. For example, a typical embedded input 114 can include hundreds of different entries. Next, the embedded input 114 uses fewer computational resources to construct a dense representation of one or more elements. Furthermore, the embedded input 114 can represent relationships between two or more elements. For example, the words “Italy” and “Germany” can be similar or more closely related because they both define European countries, while the word “embodiment” can be quite different from both. The smaller the dot product between two embedded inputs 114, the more similar the two elements associated with the embedded input 114 can be. Therefore, the embedded input 114 can accurately represent one or more elements and produce accurate results based on processing the embedded input 114.

[0159] To transform input vector 106 into embedded input 114, the embedding layer may include a number of neurons equal to the number of entries in embedded input 114. Based on embedded input 114, the output layer may generate output vector 116. The output vector may be a vector and / or may indicate one or more elements. Output vector 116 may indicate one or more elements different from input vector 106 and / or the one-hot vector associated with input vector 106. For this purpose, the output layer may include a number of neurons equal to the number of entries in input vector 106 and / or output vector 116. The output layer may apply a softmax function to embedded input 114. By doing so, the output vector may include non-zero probabilities associated with elements of entries associated with output vector 116. Therefore, one or more elements can be obtained from output vector 116 with corresponding probabilities. Where input vector 106 may specify one or more sequences of elements, output vector 116 may specify one or more elements corresponding to the sequence of elements specified by input vector 106. Figure 9 In the example, the element associated with vector 118 corresponds to the input vector with a 71% probability. Additional or alternative elements correspond to the input vector as indicated by the output vector with a lower probability. By defining a threshold that can be compared with the probability, the selection of corresponding elements can be customized according to the user's needs. The elements generated by the model, including embedding layer 102 and output layer 104, can refer to the most probable element indicated by output vector 116. Therefore, Figure 9 The model described in the paper can generate elements associated with vector 118, which has a confidence score of 71%.

[0160] Figure 9The model can be a Continuous Bag-of-Words (CBOW) model. The CBOW model can be trained based on a training dataset comprising multiple input vectors and corresponding output vectors. Since the training dataset may be unlabeled, the training of the CBOW model can be considered self-supervised. Before training the CBOW model, it can be initialized with random values ​​for the weights assigned to neurons. During the training of the CBOW model, the input vectors can be passed through the initialized embedding and output layers, and the loss can be determined by comparing the output vector obtained by passing the input vector 106 through the model with the output vector corresponding to the input vector 106 specified by the training dataset. Based on the determined loss, backpropagation can be applied to determine the gradients associated with the neurons in the embedding layer 102 and the output layer 104 to reduce the loss. Based on the determined gradients, the weights of the neurons can be updated using a gradient descent algorithm. If the CBOW model achieves the predetermined loss, training can be terminated, and a trained CBOW model can be obtained. Based on the trained CBOW model, the embedding layer 102 can be adapted to embed input data comprising one or more elements. This embedding layer 102 can be used in other machine learning architectures that require an embedding layer 102, such as in Figure 10A , Figure 10B and Figure 10C The context describes the transformer encoder, transformer decoder, or transformer encoder-decoder architecture. To train these architectures, a trained embedding layer 102 may be required. Therefore, a model, such as the CBOW model, can be trained before training the transformer encoder, transformer decoder, or transformer encoder-decoder architecture.

[0161] Figure 10A An implementation scheme of a converter encoder architecture is illustrated. The converter encoder includes an encoder input 278, one or more encoder blocks 274, 214, and an encoder output. The converter encoder architecture can be derived from those known in the art and... Figure 10C The converter encoder-decoder architecture shown is derived. Specifically, the converter encoder can be referred to as an X-former. The converter encoder architecture can correspond to an encoder architecture associated with a converter encoder-decoder architecture having additional encoder outputs, rather than a decoder directly connected to the converter encoder-decoder architecture. Multiple converter encoder architectures may exist in the art, such as the bidirectional encoder representation (BERT) from the converter.

[0162] Input data can be received at encoder input 278. Input embedding 202 can be applied to encoder input 278. Applying input embedding 202 can refer to passing the input data through an embedding layer, for example, as... Figure 9As described in the context. Furthermore, position coding 204 can be applied to encoder input 278. Applying position coding 204 can refer to adding position factors to the embedded input obtained via input embedding. Preferably, the input data can specify a sequence of elements. Position factors It can indicate the position of an element within a sequence. For example, the position factor can be obtained based on the following equation. :

[0163]

[0164]

[0165] Where pos can refer to the position of an element in the sequence, i can refer to the dimension associated with the input embedding, and d can refer to the dimension of the model, such as a transformer decoder, transformer encoder, or transformer encoder-decoder. This can be referred to as absolute position embedding. Alternatively, position encoding can be based on Rotated Position Embedding (RoPE). Position encoding is advantageous because it can handle sequential data without requiring an additional dimension to indicate the position of each element. Subsequently, position encoding 204 reduces the computational resources required to embed the input data. By passing the input data through the encoder input, the input data can be transformed into a second-order tensor representing the sequence of elements. This second-order tensor can be referred to as embedded input data. This embedded input data can be processed by the encoder block. This embedded input data can be provided to layer normalization 208 via residual connections. Multi-head self-attention 206 can be applied to this embedded input data. Multi-head self-attention 206 can include two components: multi-head and self-attention. Self-attention can be understood as a filter applied to the embedded input data. By applying the filter to the embedded input data, elements associated with the embedded input data that contribute to the output data to be generated can be identified to generate the output data. Therefore, a filter can represent the degree of contribution of the elements associated with the embedded input data to the output data to be generated. Applying a filter can be referred to as weighting the elements associated with the embedded input data. This is particularly advantageous for long sequences of elements. Filters can be learned and improved during training by learning to identify the contributions of the elements associated with the embedded input data. For example, in a partial sentence “I went to the bakery to buy a”, the last word can be generated by a data-driven model such as a transformer encoder. Self-attention can cause the transformer encoder to focus on the main words “bakery” and “buy” to generate the word “bread”. Self-attention can refer to attention generated based on the input data. Therefore, a filter can be determined based on the input data (preferably, the embedded input data). The embedded input data can be used as a query Q, keyword K, and value V regarding the self-attention operation. Self-attention can refer to attention based on the received input data. Therefore, a filter can be calculated based on the following formula by inserting the corresponding tensor based on the embedded input data:

[0166]

[0167] in The dimensions corresponding to keywords.

[0168] To further improve the efficiency of the converter encoder, multiple heads are used to apply filters, resulting in a multi-head self-attention 206. The multi-head self-attention 206 can include applying filters to two or more portions of the embedded input data. Therefore, the tensor can be divided into two or more portions, and filters can be applied to two or more portions separately by two or more heads according to the following equation:

[0169]

[0170] The parameter matrix is Where i can refer to the number of heads. , and It can refer to the value, keywords, and query dimensions.

[0171] The results of two or more heads can be cascaded according to the following equation: ,in The number of points can be indicated by h.

[0172] The embedded input data can be transformed into a context tensor via multi-head self-attention 206. This context tensor can represent a sequence of elements and the relationship between two or more elements of the input data. The context tensor can be a second-order tensor and / or may include one or more first-order tensors. After multi-head self-attention 206, layer normalization 208 can be applied based on the context tensor and / or the embedded input data from residual connections. Applying layer normalization 208 can refer to normalizing the context tensor. Normalizing the context tensor reduces the values ​​of the entries in the context tensor. This reduces the computational cost associated with processing the context tensor. After layer normalization 208, the context tensor can be passed back to feedforward layer 210, and then layer normalization 212 is performed based on the residual connections to the context tensor and / or the output of feedforward layer 210. Feedforward layer 210 can be a feedforward neural network. The feedforward neural network may include multiple fully connected neurons. Passing the context tensor through the feedforward neural network results in a linear transformation of the context tensor. Alternatively or additionally, the neural network may include one or more activation functions, such as rectified linear unit (ReLU). Thus, the neural network can be configured to perform one or more nonlinear operations on the context tensor and / or nonlinearly transform the context tensor. After the context tensor has been transformed and / or normalized by feedforward layer 210 and layer normalization 212, the context tensor can be provided to one or more additional encoder blocks 214. After passing the context tensor through feedforward layer 210, the context tensor can be adjusted for processing by additional attention layers of one or more additional encoder blocks 214 to apply a self-attention filter, preferably multi-head self-attention 206. The context vector transformed by layer normalization 212 and feedforward layer 210 can be referred to as the hidden state.

[0173] The encoder output 276 includes a linear layer 216 and a softmax layer 218. The linear layer 216 transforms the context vector into a logit vector. The linear layer can be fully connected. The logit vector obtained by passing the context tensor through the linear layer 216 can be passed through the softmax layer 218. Passing the logit vector through the softmax layer 218 means applying the softmax function to the logit vector. Applying the softmax function to the logit vector produces a probability distribution of one or more elements corresponding to the sequence of elements in the input data. One or more elements can be selected based on the probability distribution according to a predefined selection criterion. The one or more selected elements can be referred to as one or more elements generated by the transformer encoder. The one or more generated elements can be provided as encoder input to generate another one or more elements corresponding to the sequence of input data, as well as the one or more elements generated by the transformer encoder, as shown in...Figure 11 Described within the context of [the text].

[0174] Figure 10B An implementation scheme of the converter decoder architecture is shown.

[0175] The converter decoder includes a decoder input 284, one or more decoder blocks 280, 232, and a decoder output 292. The converter decoder architecture can be derived from those known in the art and... Figure 10C The transformer encoder-decoder architecture shown is derived. The transformer decoder can be referred to as the X-former. The transformer decoder architecture can correspond to the decoder architecture associated with the transformer encoder-decoder architecture, regardless of whether one or more hidden states are received from the encoder of the transformer encoder-decoder. Several transformer decoder architectures are available in the art, such as the generalized pre-trained transformer (GPT).

[0176] The decoder input 284 can be applied to, for example, in Figure 10A The input embedding 202 and position encoding 204 described within the context are similar to the input embedding 220 and position encoding 222.

[0177] Decoder block 280 may include layer normalization 226, masked multi-head self-attention 224, feedforward layer 228, and / or layer normalization 230. Embedded input data generated by passing input data through decoder input 284 can be provided to layer normalization 226 via residual connections. Furthermore, masked multi-head self-attention 224 can be applied to the embedded input data. Masked multi-head self-attention 224 corresponds to... Figure 10A The multi-head self-attention 206 described in the context further masks a portion of the embedded input data associated with elements in the sequence that are later than the element to be generated. Alternatively or additionally, portions of the input data associated with elements in the sequence that are later than the element to be generated may not be received and / or transformed into embedded input data. Therefore, the transformer decoder can be adapted to generate subsequent elements of the sequence, while the transformer encoder can be adapted to generate missing elements within a sequence and / or between two or more sequences. Thus, the transformer encoder can be configured for classification tasks. The transformer decoder can be configured for text generation.

[0178] Similar to Figure 10A The transformer encoder described within the context can generate a context tensor by applying masked multi-head self-attention 224 and layer normalization 226. This context tensor can be provided to layer normalization 230 via residual connections. Furthermore, feedforward layer 228 and layer normalization 230 can be similar to... Figure 10AThe context tensor describes the feedforward layer 210 and layer normalization 212. This context tensor can be provided to one or more additional decoder blocks 232.

[0179] The decoder output 292 may include a linear layer 234 and a softmax layer 236. The linear layer 234 and softmax layer 236 can be similar to... Figure 10A The linear layer 216 and softmax layer 218 are described within the context of the above.

[0180] Figure 10C An implementation scheme of a converter encoder-decoder architecture is illustrated. The converter encoder-decoder may include an encoder input 288, one or more encoder blocks 286, 264, a decoder input 294, a decoder block 290, and a decoder output 292. The encoder input 288 may correspond to... Figure 10A The encoder input is 278. One or more encoder blocks 286, 264 can correspond to... Figure 10A One or more encoder blocks 274, 214. Decoder input 294 can correspond to Figure 10B The decoder input is 284.

[0181] Decoder block 290 may include masked multi-head self-attention 270, layer normalization 272, feedforward layer 238, and layer normalization 240, similar to masked multi-head self-attention 224, layer normalization 226, feedforward layer 228, and layer normalization 230 as described in the context of Figure 2B. Decoder block 290 may also include multi-head self-attention 250 and layer normalization 248. Similar to... Figure 10B The description suggests that the context tensor can be obtained from the masked multi-head self-attention 270 and layer normalization 272. Similar to... Figure 10A Multi-head self-attention 206 and multi-head self-attention 250 can be applied to the context vector obtained from layer normalization 272 and the hidden states of one or more encoder blocks 286, 264. Layer normalization 248 can be applied to the context vector obtained from multi-head self-attention 250 and the context vector obtained from layer normalization 272 provided via residual connections. Similar to... Figure 10B The context vector generated by layer normalization 240 can be processed via feedforward layer 238 and layer normalization 248. The context vector generated by layer normalization 240 can be provided to another decoder block 242 similar to decoder block 290. The context vector obtained from one or more decoder blocks 290, 242 can be provided to decoder output 292. Decoder output 292 can correspond to... Figure 10B The decoder outputs 282.

[0182] Using the above architecture, the converter encoder-decoder can receive and process input data at encoder input 288 and one or more encoder blocks 286, 264, and decoder block 290 and decoder output 292. Based on the input data, the converter encoder-decoder can generate output data partially or sequentially. Sequentially generated output data can be provided to decoder input 294, one or more decoder blocks 290, 242, and decoder output 292, and / or can be processed by decoder input 294, one or more decoder blocks 290, 242, and decoder output 292. Preferably, a sequence can be provided to encoder input 288, and after at least a portion of the generated output data has been generated, at least a portion of the elements of the generated output data can be provided to decoder input 294. By doing so, since the converter encoder-decoder can receive more data over time, the next element of the output data can be generated with higher accuracy by considering both the input data and the generated output data.

[0183] Due to the transformer encoder-decoder architecture, the transformer encoder-decoder can be configured to transform a sequence into another representation of the sequence. An example of transforming a sequence into another representation could be translating a sentence into another language. Several transformer encoder-decoders are available in the art, such as BART, T5, etc.

[0184] In one implementation, layer normalization 208, 212 can be applied before the masked multi-head self-attention 224, multi-head self-attention 206, and / or feedforward layer 210 in the transcoder, transcoder, and / or transcoder-decoder. By doing so, the computational resources used to apply multi-head self-attention 206 and / or feedforward layer 210 to the embedded input data and / or context tensors can be reduced, since the entries of the corresponding tensors may be lower after normalization.

[0185] In one implementation, the decoder output 292 may include a classification neural network, additional feedforward layers, convolutional layers, fully connected layers, and so on. For example, the transformer encoder-decoder may be configured to select among multiple options. To this end, the transformer encoder-decoder may be provided with three different input datasets and may classify the context vectors obtained from one or more decoder blocks 290 via one or more linear layers. The architecture may then be extended according to the use case to be addressed. [1]

[0186] Figure 11 An implementation scheme for training and / or deploying a transformer encoder, transformer decoder, and / or transformer encoder-decoder is shown.

[0187] The encoder / decoder / encoder-decoder architecture 302 can correspond to, for example, in Figures 10A to 10C The converter decoder, converter encoder, and / or converter encoder-decoder described within the context of the converter.

[0188] The output data generated by the encoder / decoder / encoder-decoder architecture 302 may include one or more elements, specifically a sequence of elements. Previously generated elements of the output data can be provided as input for the next element in the sequence used to generate the output data.

[0189] exist Figure 11In the example, the input data may include N elements, specifically input tokens. Input tokens may be tokens specifically designed for input into a data-driven model (such as a transformer decoder, transformer encoder, or transformer encoder-decoder). The output data to be generated may include M elements. The encoder / decoder / encoder-decoder architecture 302 may generate one element of the output data based on elements received from the input data and optionally previously generated output data at a certain time step. Therefore, M time steps are required to generate M elements. The time steps include providing inputs 310, 312, and 314 to the encoder / decoder / encoder-decoder architecture 302 and receiving output data 304, 308, and 306 from the encoder / decoder / encoder-decoder architecture 302. In the first time step, input 310 may include N input tokens. The N input tokens may be associated with, for example, N words, stems, or word endings. Preferably, the N input tokens may specify a question. One or more input tokens may specify the start and / or end of a token sequence. Input 310 can be processed by encoder / decoder / encoder-decoder architecture 302. Based on input 310, at least a portion of output data 304 can be generated. At least a portion of the output data may include a first output token. In the next time step, the generated first output token may be provided together with input 312. Specifically, if input 312 can be received by converter encoder-decoder, the input token may be received at encoder input 288, and the first output token may be received at decoder input 294. If input 312 can be received by converter encoder, input 312 may be received by encoder input 278, and similarly by converter decoder and decoder input 284. Based on input 312, output data 308 including a first output token and a second output token can be generated. Generating output data 308 based on input 312 may refer to generating a second token based on the first token and N input tokens, where the first token may have already been generated based on the N input tokens. This process can be repeated until the last token in the sequence of output data 306 can be generated. Preferably, the last token may be an end token. An end token can terminate the generation of another output token.

[0190] Similarly, for data processing during the deployment of the encoder / decoder / encoder-decoder architecture 302, the encoder / decoder / encoder-decoder architecture 302 can be trained. The training dataset can include multiple sequences, each comprising multiple elements. Sequences can be associated with input data and / or output data. Alternatively or additionally, sequences can be independent of the input data and / or output data. For example, where the input and output data can refer to chemical components represented via text, the training dataset can include sequential text data independent of chemical components. In this example, the training dataset can include word sequences derived from a dialogue. In one embodiment, the training dataset can at least partially include the input dataset and / or the output dataset.

[0191] Training can be initialized by initializing the encoder / decoder / encoder-decoder architecture 302. In one implementation, the parameters associated with the encoder / decoder / encoder-decoder architecture 302 can be randomly initialized. Alternatively or additionally, the input embeddings of the encoder / decoder / encoder-decoder architecture 302 can be obtained by training a CBOW model or a skip-gram model, as in... Figure 9 As described in the context, a trained embedding layer can be used during training. The parameters associated with the embedding layer can remain constant and / or can be updated after a predefined number of training epochs. By doing so, the number of parameters to be updated is reduced, resulting in faster training with less computational resource consumption. Furthermore, the accuracy associated with the embedding layer can be constant and / or can be increased by avoiding error compensation associated with the newly initialized encoder / decoder / encoder-decoder architecture 302.

[0192] During training of the encoder / decoder / encoder-decoder architecture 302, at least a portion of the sequence of the training dataset can be provided to the encoder / decoder / encoder-decoder architecture 302 one by one, and one or more elements can be generated based on the sequence of the training dataset one by one. Elements generated based on the sequence may follow elements of the portion of the sequence that may have been provided to the encoder / decoder / encoder-decoder architecture 302. The generated one or more elements can be compared with one or more elements following at least a portion of the sequence provided to the encoder / decoder / encoder-decoder architecture 302 as specified by the training dataset. Therefore, during training, the encoder / decoder / encoder-decoder architecture 302 can generate a guess about the next element, and the guess about the next element in the sequence can be compared with the baseline truth specifying the actual next element according to the training dataset. Based on the guess about the next element and the baseline truth, a loss can be determined. The loss can define the similarity between the guess about the next element and the baseline truth. The loss can be determined by forming a vector dot product between the tokens associated with one or more elements and the tokens associated with the baseline truth. A loss that is not equal to zero can lead to an update of the parameters associated with the encoder / decoder / encoder-decoder architecture 302. Preferably, the parameters associated with the encoder / decoder / encoder-decoder architecture 302 can be independent of the embedding layer. For example, the parameters associated with the encoder / decoder / encoder-decoder architecture 302 can be the weights of the neurons in the encoder / decoder / encoder-decoder architecture 302.

[0193] Based on the determined loss, backpropagation can be applied to determine the gradients of the parameters associated with the encoder / decoder / encoder-decoder architecture 302 to reduce the loss. According to the determined gradients, the parameters associated with the encoder / decoder / encoder-decoder architecture 302, preferably the weights of the neurons associated with the encoder / decoder / encoder-decoder architecture 302, can be updated using a gradient descent algorithm.

[0194] The training dataset can be unlabeled. The sequence of elements within the training dataset can inherently include benchmark truth for determining the loss about one or more elements generated during training of the encoder / decoder / encoder-decoder architecture 302. Therefore, the encoder / decoder / encoder-decoder architecture 302 can be trained self-supervised. This is advantageous because it saves time and resources used to create labeled training datasets. Furthermore, it enables the use of large training datasets associated with sizes in the trillions of bytes. Therefore, the data-driven model can be accurate in generating the elements of the sequence. Moreover, large training datasets enable a small number of predictions or even zero predictions. Thus, the data-driven model trained as described above is versatile and helps save the resources required to train and / or host multiple purpose-driven models such as CNNs. The training described above can be referred to as pre-training. The data-driven model can be configured to perform a small number of predictions or even zero predictions about multiple use cases after pre-training. The performance of the data-driven model can be further improved through additional training, known as fine-tuning.

[0195] Figure 12 An implementation scheme for input embedding is illustrated. Where the sequence of elements associated with the input data (preferably included in the input data) can be of one type, methods such as those used in... Figures 10A to 10C The input embeddings described in the context are 202, 220, 252, and 266. For example, the type of input data can be text, where elements can be associated with at least a portion of a word, punctuation marks, start tokens specifying the beginning of one or more sequences associated with the input data, and / or end tokens. In another example, the input data can be at least partially numeric. Therefore, the input data can include multiple numbers. Numerical input data can be, for example, tabular data. Tabular data can specify one or more rows and / or one or more columns. Therefore, tabular data can include one or more cells, where cells can be associated with one or more numeric values.

[0196] Numeric input data may require different embeddings than text input data. Input embeddings for numeric input data can include token embeddings, position embeddings, column embeddings, row embeddings, or combinations thereof.

[0197] Applying token embedding to one or more elements (specifically, tokens associated with input data) can produce a machine-processable representation associated with one or more elements (specifically, tokens). Applying token embedding to one or more elements can refer to passing one or more elements through an embedding layer, for example, as in... Figure 9As described in the context of [the previous section], token embeddings can specify one or more elements, specifically tokens in a machine-processable representation. For example, token embeddings can transform numerical values ​​into vectors. This is advantageous because the representation can be enriched with more information, such as the position of the token within the sequence and / or the position of the token within a table associated with the token sequence. Positional embeddings can be similar to [the previous section]... Figure 9 , Figures 10A to 10C The location embedding is described within the context of the table. Column embedding can be applied when the input data can be tabular data. Applying column embedding to one or more elements (particularly tokens associated with the input data) can produce a machine-processable representation specifying the position of one or more elements within table 402 (preferably within columns of table 402). Applying column embedding can refer to adding column factors to the input data embedded via token embedding, particularly embedded input data. Column factors can be the same for elements associated with the same column, and / or different for two or more elements associated with different columns. Similarly, row embedding can be applied when the input data can be tabular data. Applying row embedding to one or more elements (particularly tokens associated with the input data) can produce a machine-processable representation specifying the position of one or more elements within table 402 (preferably within rows of table 402). Applying row embedding can refer to adding column factors to the input data embedded via token embedding, particularly embedded input data. Row factors can be the same for elements associated with the same row, and / or different for two or more elements associated with different rows.

[0198] In one implementation, the input data may be at least partially numerical and at least partially textual. Therefore, the input data may include two or more types of data. The data type may refer to a modality. Next, different embeddings can be applied to the input data. For the portion of the input data that includes text, embeddings can be applied... Figure 9 , Figures 10A to 10CThe input embeddings referenced herein. For the portion of input data that is embedded as a numeric token, positional embedding, column embedding, and row embedding can be applied. Furthermore, segmented embeddings can be applied to the input data, regardless of the type of input data. Segmented embeddings can specify the type of input data that can be associated with one or more elements. For example, if the input data includes both text and numbers, the input data can include both types of input data. Applying segmented embeddings to input data can refer to adding segmentation factors to the input data, preferably embedded input data and / or input data after applying token embeddings. Segmentation factors can specify the type of data associated with one or more elements. Segmentation factors can be the same for one or more elements associated with the same type of input data, and / or can differ between two or more elements associated with different types of input data.

[0199] Applying token embedding, position embedding, segment embedding, column embedding, row embedding, or a combination thereof can produce embedded input data and / or can be the output of any of encoder inputs 278, 284, 288 or decoder inputs 284, 294. Data obtained by applying token embedding, position embedding, segment embedding, column embedding, row embedding, or a combination thereof can be processed by encoder blocks 274, 286, decoder blocks 280, 290, encoder output 276, and decoder outputs 292, 282.

[0200] Training data 1013 based on the factory, such as in Figures 9 to 12 The untrained / pre-trained transformer-based models described in the context of and / or as in Figure 2a , Figure 2c , Figure 6 and Figure 8 The trained transformer-based models described in the context can be stored in a database, on an electronic data carrier, or in the cloud. For models suitable for training pre-trained transformer-based models, such as those described in... Figures 9 to 12 The untrained / pre-trained transformer-based models described in the context and / or as in Figure 2a , Figure 2c , Figure 6 and Figure 8Access to the factory-based training data 1013 of the trained transformer-based model described in the context can be granted to the user as a computer-readable token. The computer-readable token can be an authorization key generated by a computer processor based on a request from a requesting computation node associated with the user. This request can be sent to an authorization engine comprising at least one computer processor authorized to grant one or more authorization keys to access one or more of the factory-based training data and / or untrained, pre-trained, trained, retrained, or fine-tuned transformer-based models, as described in [the context of the previous sentence]. Figure 2a , Figure 2c , Figure 6 , Figure 8 and Figures 9 to 12 As described in the context.

[0201] Users can access training data via a token to process the data and receive processing results without receiving the actual training data. Users can also receive the actual factory-based training data or a portion of the factory-based training data. Depending on access permissions, users (upon receiving a token) can use one or more of untrained, pre-trained, and / or trained transformer-based models to achieve their technical objectives. Users can use the claimed factory-based training data or their own training data to train / retrain one or more models they are licensed to use. The access token for using the claimed factory-based training data can be used in conjunction with the token for using... Figure 2a , Figure 2c , Figure 6 , Figure 8 and Figures 9 to 12 The access tokens for one or more models described in the context may be the same or different. Each of the data products (i.e., based on the factory's training data) or data services has a separate access token (i.e., using a token such as in...). Figure 2a , Figure 2c , Figure 6 , Figure 8 and Figures 9 to 12 One or more models described in the context can enhance the security of using data products and / or data services. Access tokens may include user credentials, one or more generation algorithms, user authentication, two-factor authentication, etc.

[0202] The following list of subcategories describes one possible way to implement the disclosure of this application.

[0203] project:

[0204] Project 1. A computer-implemented method for deploying a trained transformer-based model to control and / or monitor a distributed production environment, the distributed production environment comprising one or more pieces of equipment for producing products, the method comprising:

[0205] Provides factory-based training data associated with one or more production operations via a computer interface;

[0206] A pre-trained converter-based model, including at least converter components, is provided via the computer interface.

[0207] The pre-trained transformer-based model is suggested to be fine-tuned or retrained using the factory-based training data;

[0208] The trained transformer-based model is deployed for one or more production operations in the distributed production environment.

[0209] Project 2. A computer-implemented method for controlling and / or monitoring a distributed production environment using a transformer-based model trained according to Project 1, the distributed production environment comprising one or more pieces of equipment for producing products, the method comprising:

[0210] Access to the trained transformer-based model is received by the computer processor;

[0211] Plant-based data received via a computer interface in connection with one or more production operations;

[0212] The trained converter-based model is prompted to analyze the input factory-based data and provide operational instructions for production.

[0213] Project 3. The computer-implemented method according to Project 1, wherein the raw factory-based data is preprocessed by a preprocessing engine to provide the factory-based training data.

[0214] Project 4. The computer-implemented method according to Project 2, wherein the raw factory-based data is preprocessed by a preprocessing engine to provide input factory-based data.

[0215] Project 5. The computer-implemented method according to Project 2 or 4, wherein the input factory-based data includes data provided to the operator via a computer interface.

[0216] Project 6. A computer-implemented method according to any one of Projects 2, 4, or 5, the method further comprising prompts from an operator, the prompts prompting the trained transformer-based model to provide machine-readable instructions based on queries provided to the model by the operator via a computer interface.

[0217] Item 7. A computer-implemented method according to any one of Items 2 and 4 to 6, wherein the prompting includes one or more contexts relating to one or more production operations at the distributed production environment.

[0218] Project 8. A computer-implemented method according to any one of Projects 1 to 7, wherein the factory-based training data is factory-based data in one or more languages, and / or the operation instructions used for production are in one or more languages.

[0219] Project 9. A computer-implemented method according to any one of Projects 1 to 8, wherein the method further comprises updating the factory-based training data and prompting the trained transformer-based model to fine-tune or retrain based on the updated factory-based training data.

[0220] Item 10. The computer-implemented method according to Item 9, wherein the fine-tuning or retraining of the trained transformer-based model is any one of scheduled fine-tuning or retraining, continuous fine-tuning or retraining, or trigger-based fine-tuning or retraining.

[0221] Item 11. The computer-implemented method according to Item 10, wherein the triggering is based on a threshold and score associated with the operation instruction provided by the trained transformer-based model, and / or optionally, based on the number of fine-tuning or retraining cycles.

[0222] Item 12. A computer-implemented method according to any one of Items 2 to 11, wherein the operating instructions include machine-readable instructions for controlling and / or monitoring production.

[0223] Item 13. A computer program product comprising computer-readable instructions that, when executed on a computer, cause the computer to perform the steps according to any one of items 1 to 12.

[0224] Item 14. A computer-readable storage medium storing computer-readable instructions that, when executed on a computer, cause the computer to perform the steps according to any one of items 1 to 12.

[0225] Item 15. A computer product comprising a computer-readable token for accessing factory-based training data and / or a trained or pre-trained model according to any one of Items 1 to 12.

[0226] The following example implementation plans should also be made public:

[0227] Clause 1:

[0228] A computer-implemented method for controlling and / or monitoring a distributed production environment using a trained transformer-based model, the distributed production environment comprising one or more pieces of equipment for producing products, the method comprising:

[0229] - Access to the trained transformer-based model is received by the computer processor;

[0230] - Receives plant-based data via a computer interface that is associated with one or more production operations;

[0231] - The trained transformer-based model is prompted to analyze the input factory-based data and provide operational instructions for the one or more production operations in the distributed production environment.

[0232] Clause 2:

[0233] A computer-implemented method for generating the trained transformer-based model used according to Clause 1, the method comprising:

[0234] - Provides factory-based training data associated with one or more production operations via a computer interface;

[0235] - Provide a pre-trained converter-based model, including at least converter components, via the computer interface;

[0236] - This suggests that the pre-trained transformer-based model can be retrained or fine-tuned using the factory-based training data;

[0237] - Publish the trained transformer-based model for one or more production operations in the distributed production environment according to Clause 1.

[0238] Clause 3:

[0239] According to the computer-implemented method described in Clause 2, the raw factory-based data is preprocessed by a preprocessing engine to provide the factory-based training data.

[0240] Clause 4:

[0241] The computer-implemented method according to Clause 1, wherein the raw factory-based data is preprocessed by a preprocessing engine to provide input factory-based data.

[0242] Clause 5:

[0243] The computer-implemented method according to Clause 1 or 4, wherein the input factory-based data includes data provided to the operator via a computer interface from the operator.

[0244] Clause 6:

[0245] The computer-implemented method according to any one of clauses 1, 4, or 5 further includes prompts from an operator that prompt the trained transformer-based model to provide machine-readable instructions based on queries provided to the model by the operator via a computer interface.

[0246] Clause 7:

[0247] The computer-implemented method according to any one of Clauses 1 and 4 to 6, wherein the prompt includes one or more contexts relating to one or more production operations at the distributed production environment.

[0248] Clause 8:

[0249] The computer-implemented method according to any one of Clauses 2 to 7, wherein the factory-based training data is factory-based data in one or more languages, and / or the operating instructions used for production are in one or more languages.

[0250] Clause 9:

[0251] The computer-implemented method according to any one of Clauses 2 to 8, wherein the method further includes updating the factory-based training data and prompting the trained transformer-based model for retraining or fine-tuning based on the updated factory-based training data.

[0252] Clause 10:

[0253] The computer-implemented method according to Clause 9, wherein the retraining of the trained transformer-based model is any one of scheduled retraining, continuous retraining or fine-tuning, or trigger-based retraining or fine-tuning.

[0254] Clause 11:

[0255] The computer-implemented method according to Clause 10, wherein the triggering is based on a threshold and score associated with the operating instructions provided by the trained transformer-based model, and / or optionally, on the number of retraining or fine-tuning cycles.

[0256] Clause 12:

[0257] The computer-implemented method according to any one of Clauses 1 to 11, wherein the operating instructions include machine-readable instructions for controlling and / or monitoring production.

[0258] Clause 13:

[0259] A computer program product comprising computer-readable instructions that, when executed on a computer, cause the computer to perform any one of the steps pursuant to Articles 1 to 12.

[0260] Clause 14:

[0261] A computer-readable storage medium storing computer-readable instructions that, when executed on a computer, cause the computer to perform any one of the steps pursuant to Articles 1 to 12.

[0262] Clause 15:

[0263] A computer product comprising a computer-readable token for accessing factory-based training data and / or a trained or pre-trained or fine-tuned model according to any one of Clauses 1 to 12.

[0264] It is understood that, unless otherwise disclosed in this application, the features of the above embodiments are combinable.

[0265] The prior art publication; No. 684; paragraphs

[1000] to

[8005] ; ISSN: 2198-4786; publication date: February 12, 2024, shall be considered as reference RF1, the full text of which is incorporated herein by reference. Preferably, the (chemical) product is the product described in paragraphs

[1000] to

[8005] of reference RF1. Preferably, the method / process described herein is further a method / process for producing the product.

[0266] The conversion steps for obtaining the product preferably include one or more steps as described below, and can be performed by conventional methods well known to those skilled in the art. The conversion steps preferably include one or more steps selected from the following:

[0267] Recovery, preferably through depolymerization, gasification, pyrolysis, and / or steam cracking; and / or

[0268] Purification, preferably crystallization, (solvent) extraction, distillation, evaporation, hydrogenation, absorption, adsorption, and / or ion exchange using ion exchangers; and / or

[0269] Assembly, preferably foaming, synthesis, chemical transformation, chemical conversion, polymerization and / or compounding; and / or

[0270] Molding, preferably foaming, extrusion and / or molding; and / or finishing, preferably coating and / or smoothing.

[0271] Furthermore, one or more steps are described in detail in paragraphs

[1000] through

[8005] of reference RF1.

[0272] This disclosure has also described in conjunction with preferred embodiments and examples. However, those skilled in the art can understand and implement other variations of the claimed subject matter through a study of the accompanying drawings, this disclosure, and the claims. It is noteworthy that, in particular, any steps presented can be performed in any order; that is, this disclosure is not limited to a specific order of these steps. Furthermore, it is not necessary to perform different steps at a specific location in the distributed system or at a single node; that is, each of these steps can be performed at different nodes using different equipment / data processing.

[0273] The order of all the method steps presented above is not mandatory, and alternative orders are possible. However, the specific order of the method steps shown as examples in the accompanying drawings should be considered as one possible order of the method steps, for example, for the corresponding embodiments described in the corresponding drawings or embodiments that include at least some of the steps described in the corresponding drawings.

[0274] In this specification, any connections presented in the described embodiments should be understood in a manner that allows the components involved to be operatively coupled. Therefore, connections can be direct or indirect, have any number or combination of intermediate elements, and may exist solely as functional relationships between components.

[0275] The indefinite article “a” or “an” should not be interpreted as “one”, that is, using the expression “one element” does not exclude the presence of other elements. A single element or other unit may perform the function of several entities or items recited in the claims. The fact that certain measures are recited only in mutually different dependent claims does not mean that combinations of these measures cannot be used in advantageous embodiments or that additional elements may be included.

[0276] The expressions “A and / or B” and “at least one of A or B” are considered interchangeable and are intended to include any one of the following three cases: (i) A, (ii) B, (iii) A and B. More generally, the expressions “at least one of the following: ” and “at least one of ” and similar wording (where the list of two or more elements is connected by “and” or “or”) mean at least one element, or at least any two or more elements, or at least all elements.

[0277] The provision within the scope of this disclosure may include any interface configured to provide data. This may include application programming interfaces, human-machine interfaces (such as displays), and / or software module interfaces. The provision may include communication of data or submission of data to an interface, particularly displaying data to a user or using data by a receiving entity.

[0278] Acquisition within the scope of this disclosure may include any interface configured to acquire or receive data. This may include application programming interfaces, human-machine interfaces (such as displays), and / or software module interfaces. Acquisition may include transferring or submitting data from the interface, particularly data used by the receiving entity. Any acquisition of data, data structures, datasets, etc., may include receiving data, data structures, datasets, etc., from a server that provides (e.g., hosts) a database containing data, data structures, datasets, etc.

[0279] Various units, circuits, entities, nodes, or other computing components may be described as being “configured to” perform one or more tasks. “Configured to” means that a structure “has” a “circuit” that performs one or more tasks during operation. Units, circuits, entities, nodes, or other computing components may be configured to perform tasks even when the unit / circuit / component is not operating. Units, circuits, entities, nodes, or other computing components forming a structure corresponding to “configured to” may include hardware circuitry and / or memory storing executable program instructions to perform the operation. For ease of description, units, circuits, entities, nodes, or other computing components may be described as performing one or more tasks. Such descriptions should be interpreted as including the phrase “configured to”. Any expression “configured to” is explicitly intended not to invoke the interpretation of 35 USC § 112(f).

[0280] Generally speaking, the methods, apparatus, systems, computer elements, nodes, or other computing components described herein may include memory, software components, and hardware components. Memory may include volatile memory (such as static or dynamic random access memory) and / or non-volatile memory (such as optical disc or magnetic disk storage devices, flash memory, programmable read-only memory, etc.). Hardware components may include any combination of combinational logic circuits, clock storage devices (such as flip-flops, registers, latches, etc.), finite state machines, memory (such as static random access memory or embedded dynamic random access memory), custom-designed circuits, programmable logic arrays, etc.

[0281] In this specification, any connections presented in the described embodiments should be understood in a manner that allows the components involved to be operatively coupled. Therefore, connections can be direct or indirect, have any number or combination of intermediate elements, and may exist solely as functional relationships between components.

[0282] Furthermore, any methods, processes, and actions described or illustrated herein may be implemented using executable instructions in a general-purpose or special-purpose processor and stored on a computer-readable storage medium (e.g., a disk, memory, etc.) for execution by such a processor. The reference to "computer-readable storage medium" should be understood to encompass special-purpose circuitry, such as signal processing apparatus and other means.

[0283] Any disclosure and implementation described herein relates to the methods, systems, apparatuses, and computer program elements listed above, and vice versa. Advantageously, the benefits provided by any implementation and example also apply to all other implementations and examples, and vice versa.

[0284] Unless otherwise specified, all terms and definitions used herein should be understood broadly and have their general meanings.

[0285] It should be understood that all presented embodiments are merely examples, and any feature presented for a particular example embodiment may be used alone with any aspect, or in combination with any feature presented for the same or another particular example embodiment, and / or in combination with any other feature not mentioned. In particular, the example embodiments presented in this specification should also be understood as being disclosed in all possible combinations between each other, provided that it is technically reasonable and the example embodiments are not alternatives to each other. It should also be understood that any feature presented for example embodiments in a particular category (method / apparatus / computer program / system) may also be used in a corresponding manner in example embodiments of any other category. It should also be understood that the presence of a feature in a presented example embodiment does not necessarily mean that the feature is essential and cannot be omitted or replaced.

[0286] Reference List

[0287] Factory 1001;

[0288] Equipment 1002;

[0289] 1003 sensor;

[0290] 1004 Analysis Engine;

[0291] Factory data (e.g., sensor data, analytics data);

[0292] 1006 Control and / or monitoring engine;

[0293] 1007 and 1008 are machine-readable instructions;

[0294] Historical data of Factory 1010;

[0295] 1011 Input data;

[0296] 1012 operator data;

[0297] 1013 training data;

[0298] 7001 prompts the AI ​​engine by providing hints such as text prompts (e.g., resolving anomalies in factory operations).

[0299] 7002 Add context to the prompt (e.g., abnormal temperature);

[0300] 7003 Optional: Add another context (e.g., data points from the preprocessing engine, such as the plant's identifier in the metadata, the plant's typical operating temperature, etc.).

[0301] 7004 uses AI to generate instructions based on prompts for monitoring and / or controlling production (e.g., instructions on how to resolve temperature anomalies).

Claims

1. A method for controlling and / or monitoring a distributed production environment, the distributed production environment comprising one or more pieces of equipment for producing products, the method comprising: - Receives plant-based data via a computer interface that is associated with one or more production operations; - Determine operational instructions for the one or more production operations in the distributed production environment, wherein the operational instructions include providing the prompt to at least one generative data-driven model in response to receiving a prompt, the at least one generative data-driven model having been trained to generate the operational instructions; - Provide the operation instructions.

2. A method for generating the at least one generative data-driven model for use according to claim 1, the method comprising: - Provides factory-based training data associated with one or more production operations via a computer interface; - Provide a pre-trained generative data-driven model via the computer interface; - Use the factory-based training data to fine-tune the pre-trained generative data-driven model; - Publish the trained generative data-driven model for one or more production operations in the distributed production environment of claim 1.

3. The method of claim 2, wherein the raw factory-based data is preprocessed by a preprocessing engine to provide the factory-based training data.

4. The method of claim 1, wherein the raw plant-based data is preprocessed by a preprocessing engine to provide the input plant-based data.

5. The method of claim 1 or 4, wherein the input plant-based data includes data provided by the operator via a computer interface.

6. The method according to any one of claims 1, 4 or 5, wherein the prompt comprises instructions that provide machine-readable instructions to the trained at least one generative data-driven model.

7. The method of any one of claims 1 to 6, wherein the prompt includes one or more contexts relating to one or more production operations at the distributed production environment.

8. The method according to any one of claims 2 to 7, wherein the factory-based training data is factory-based data in one or more languages, and / or the operation instructions used for production are in one or more languages.

9. The method according to any one of claims 2 to 8, wherein the method further comprises updating the factory-based training data and fine-tuning or retraining the trained generative data-driven model based on the updated factory-based training data.

10. The method of claim 9, wherein the retraining or fine-tuning of the trained generative data-driven model is any one of scheduled retraining or fine-tuning, continuous retraining or fine-tuning, or trigger-based retraining or fine-tuning.

11. The method of claim 10, wherein the triggering is based on a threshold and score associated with the operational instructions provided by the trained generative data-driven model, and / or optionally, on the number of retraining or fine-tuning cycles.

12. The method according to any one of claims 1 to 11, wherein the operating instructions include machine-readable instructions for controlling and / or monitoring production.

13. A computer program product comprising computer-readable instructions that, when executed on a computer, cause the computer to perform the steps according to any one of claims 1 to 12.

14. An apparatus comprising corresponding components for performing or carrying out the steps of any one of claims 1 to 12, or comprising at least one processor and at least one memory storing instructions, the instructions causing the apparatus to perform at least the steps of the method according to any one of claims 1 to 12 when executed by the at least one processor.

15. A computer product comprising a computer-readable token for accessing factory-based training data and / or a trained or pre-trained or fine-tuned model according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Determining operating conditions in chemical production plants

    WO2020165045A1

  • Manufacturing system for monitoring and / or controlling one or more chemical plant(s)

    WO2021116123A1

  • Industrial plant monitoring

    WO2021156157A1