Method for anomalous mitigation in production cycles for distributed chemical production

By using generative data-driven models to identify and mitigate anomalies early in distributed chemical production, the problem of declining product quality has been solved, and resource conservation and efficiency improvement have been achieved.

CN121569256APending Publication Date: 2026-02-24BASF SE
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202480048843.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-26
Filing Date
2024-07-22
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

In distributed chemical production environments, existing technologies struggle to detect and mitigate anomalous behavior early, leading to decreased product quality and resource waste.

Method used

By adopting a generative data-driven model, the system receives analysis instructions, obtains reference and datasets to be analyzed, generates anomaly mitigation instructions, and identifies and mitigates anomalies in the production process as early as possible.

Benefits of technology

It reduces product waste, saves resources, improves chemical production efficiency, and enables early detection of abnormal behavior to restore product quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121569256A_ABST
    Figure CN121569256A_ABST
Patent Text Reader

Abstract

The present disclosure may relate to mitigating climate changes through advanced manufacturing in industrial environments of chemical production. The disclosure also relates to early detection of anomalies in the production cycle and removal of sources of anomalies in the production cycle to restore the quality of the product in the production cycle. The detection and removal of anomalies may be provided by using one or more pre-trained transducer-based models that are trained using one or more training data sets associated with one or more production operations of distributed chemical production. The removal of sources of anomalies in a production cycle may allow for a reduction in the waste of products as well as the waste of production resources associated with producing the products.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a method for mitigating anomalies in the production cycle of distributed chemical production. Background Technology

[0002] For example, the distributed production environment of a chemical plant is highly complex, in which various chemical processes are carried out under the control of various chemical devices. These processes may involve the handling, storage, and conversion of different chemicals, including hazardous, flammable, and / or toxic chemicals. Completing a production cycle for a chemical product in such an environment can take anywhere from days to weeks, for example, when performed in batches. When abnormal behavior occurs in a production plant (such as one or more plants discussed in WO2020165045 (A1), WO2021116123 (A1), and WO2021156157 (A1)), such abnormal behavior can be detected when the quality of the final product is measured at the end of the production cycle. Summary of the Invention

[0003] According to the first aspect, a method for mitigating anomalies in the production cycle of a distributed production environment is disclosed, comprising:

[0004] - Receive analysis instructions related to detecting one or more anomalies in factory-based data, wherein the analysis instructions include a reference indicator indicating at least one reference dataset and an analysis indicator indicating at least one dataset to be analyzed;

[0005] - Based on the analysis instructions, obtain at least one reference dataset and at least one dataset to be analyzed;

[0006] - Determine at least one anomaly mitigation instruction, and determine at least one anomaly detection result by: providing a task instruction to at least one generative data-driven model based on the analysis instruction, the at least one generative data-driven model being trained to generate at least one anomaly mitigation instruction based on at least one reference dataset and at least one dataset to be analyzed in response to receiving the task instruction;

[0007] - Provide at least one exception mitigation instruction.

[0008] According to another source, the relevant equipment, systems, and uses have been disclosed.

[0009] Implementation Plan

[0010] For example, the distributed production environment of a chemical plant is highly complex, where various chemical processes are carried out under the control of various chemical devices. These processes may involve the handling, storage, and conversion of different chemicals, including hazardous, flammable, and / or toxic chemicals. Producing chemical products in such an environment can take days to weeks, for example, using a batching approach. The aspects, embodiments, and examples provided in this disclosure can allow for the early detection of anomalies in the production environment, which can, for example, lead to a decline in product quality. The aspects, embodiments, and examples provided in this disclosure can allow for the (near) real-time identification of anomalous behavior and its sources within a production cycle (e.g., within hours before the end of an affected production cycle). This can allow for mitigation of the decline in the quality of the final product in the affected production cycle and the next consecutive cycle. Therefore, waste generation can be reduced, for example, by reducing the production of products of unacceptable quality. In addition to reducing product waste, resources spent on producing said products can also be saved.

[0011] The aspects, implementation methods, and examples provided in this disclosure can allow for improvements in the efficiency of chemical production to reduce waste caused by unacceptable quality of the produced products, thereby mitigating climate change.

[0012] According to the first aspect, a method (particularly computer-implemented) for mitigating anomalies in the production cycle of distributed chemical production is disclosed, comprising:

[0013] - Receive (e.g., from a plant operator) analysis instructions related to detecting one or more anomalies in plant-based data, wherein the analysis instructions include a reference indicator indicating at least one reference dataset and an analysis indicator indicating at least one dataset to be analyzed;

[0014] - Based on the analysis instructions, obtain at least one reference dataset and at least one dataset to be analyzed (e.g., from a database);

[0015] - Determine at least one anomaly mitigation instruction (e.g., for removing one or more sources of one or more anomalies), and determine at least one anomaly detection result by: providing task instructions to at least one generative data-driven model (e.g., a transformer-based model) based on analysis instructions, the at least one generative data-driven model being trained to generate at least one anomaly mitigation instruction based on at least one reference dataset and at least one dataset to be analyzed in response to receiving task instructions;

[0016] - Provide at least one anomaly mitigation instruction (e.g., to plant operators to mitigate any identified anomalies in the (ongoing) production process).

[0017] Reference indicators and / or analysis indicators can allow the identification of specific datasets, such as a reference dataset or a dataset to be analyzed that includes production data from a specific production line within a specific time range. For example, a specific dataset could be a given range of data (e.g., from date 1 to date 2) and associated with the name of a factory, production line, or equipment. Reference indicators and / or analysis indicators can also include a corresponding dataset, such that the corresponding dataset may not need to be retrieved and can be obtained directly from the analysis instructions. A reference dataset can be a dataset that includes data from the factory's or production line's historical or past operations and can be used as a training dataset for fine-tuning generative data-driven models. Reference datasets can be stored in a database and can be retrieved (automatically) using a (database) query to identify the dataset. A dataset to be analyzed can include data from the current or ongoing production cycle, such as data measured at the end of a production line within a production cycle. The dataset to be analyzed can also be stored in a database or on another computer storage device such as memory. The dataset to be analyzed can be data measured and (automatically) stored during a production cycle.

[0018] Obtaining at least one reference dataset and at least one dataset to be analyzed (e.g., from a database) may include generating a retrieval query against the database by at least one generative data-driven model or another generative data-driven model, and retrieving the dataset from the database based on the retrieval query.

[0019] Task instructions can be determined based on analysis instructions, for example by including at least a portion of the analysis instructions and additional data such as context, or task instructions can be generated by at least one generative data-driven model or another generative data-driven model.

[0020] According to the example implementation of any aspect, the task instruction includes at least a portion of at least one reference dataset and / or at least a portion of at least one dataset to be analyzed.

[0021] For example, a task instruction may include at least a portion of an analysis instruction, at least a portion of at least one reference dataset, and at least a portion of at least one dataset to be analyzed (e.g., as context). Including a portion of at least one reference dataset in the task instruction can be used for one-time training of a generative data-driven model.

[0022] According to example implementations of either aspect, at least one reference dataset (e.g., in operation) is used to fine-tune or retrain the generative data-driven model, and the fine-tuned or retrained generative data-driven model is used as the generative data-driven model when determining at least one anomaly mitigation instruction.

[0023] When using a fine-tuned generative data-driven model, the task instructions for fine-tuning the generative model may include at least a portion of the dataset to be analyzed, but may exclude, for example, at least one reference dataset.

[0024] According to the example implementation of any aspect, obtaining at least one reference dataset and / or at least one dataset to be analyzed includes: determining whether at least one reference dataset and / or at least one dataset to be analyzed is available in the database;

[0025] When it is determined that at least one reference dataset and / or at least one dataset to be analyzed is available in the database: retrieve at least one reference dataset and / or at least one dataset to be analyzed from the database.

[0026] For example, when it is determined that at least one reference dataset is not available in the database: request the user to provide at least one reference dataset; and receive at least one reference dataset via a computer interface.

[0027] For example, when it is determined that at least one dataset to be analyzed is not available in the database: request the user to provide at least one dataset to be analyzed; and receive at least one dataset to be analyzed via a computer interface.

[0028] According to an example implementation of the method of the first aspect, the method further includes:

[0029] The production cycle of distributed chemical production is operated and / or controlled based on at least one anomaly mitigation instruction, particularly for the production of chemical products. Operators can influence production line parameters based on at least one anomaly mitigation instruction, thereby improving the quality of the produced products compared to a production line operating with unchanged parameters.

[0030] According to an example implementation of any aspect, the task instruction includes instructions for generating a technical report, which includes a description indicating the presence of one or more anomalies or indicating the absence of anomalies, wherein, optionally, in the case that one or more anomalies are detected, the technical report also includes an indication of one or more sources of the one or more anomalies.

[0031] According to example implementations of any aspect, a generative data-driven model is fine-tuned or retrained using a technical training dataset (e.g., a first training dataset) comprising one or more plant technical documents, and wherein the fine-tuned or retrained generative data-driven model is used as a generative data-driven model when determining at least one anomaly mitigation instruction. For example, this fine-tuning may complement fine-tuning using at least one reference dataset (e.g., as a second training dataset), which includes plant-based data associated with one or more production operations on one or more production lines of distributed chemical production. For example, the generative data-driven model can be fine-tuned using the technical training dataset in a first training cycle, and then the generative data-driven model fine-tuned in the first cycle can be further fine-tuned using the reference dataset (in a second cycle).

[0032] According to an example implementation of any aspect, the technical training dataset includes data points associated with one or more production operations and with standard operating conditions of the production line, wherein the standard operating conditions are defined by a predefined range of one or more parameters associated with one or more production operations of the production line, wherein optionally, the one or more parameters are any of the following: time period, factory identifier, production line identifier, production cycle identifier, one or more operating parameters associated with one or more production operations, product identifier, factory, one or more parameters associated with one or more attributes of the product, data type, such as temperature, pressure, flow rate, quality.

[0033] According to an example implementation of the method of the first aspect, the method further includes:

[0034] Determine (e.g., identify) whether access credentials and / or authorization credentials are required to access one or more databases, including at least one reference dataset, at least one dataset to be analyzed, and / or a technical dataset.

[0035] When it is determined that access credentials and / or authorization credentials are required, (e.g., by a generative data-driven model) a request is generated for the user to provide authorization credentials and / or access credentials for accessing one or more databases.

[0036] (For example, by querying a database) (e.g., from a database) at least one reference dataset, at least one dataset to be analyzed, and / or a technical dataset, and optionally, at least one reference dataset and / or technical dataset is used to fine-tune or retrain the generative data-driven model; and wherein the fine-tuned or retrained generative data-driven model is used as the generative data-driven model when determining at least one anomaly mitigation instruction.

[0037] Whether to use at least one reference dataset and / or technical dataset to fine-tune or retrain the generative data-driven model can be decided by the operator via a computer interface, or can be done automatically after retrieving at least one reference dataset and / or technical dataset.

[0038] According to an example implementation of the method of the first aspect, the method further includes:

[0039] Access to at least one generative data-driven model is provided via a computer interface, wherein at least one anomaly mitigation instruction is determined using the accessed generative data-driven model.

[0040] According to the example implementation of any aspect, obtaining at least one reference dataset, at least one dataset to be analyzed, and / or a technical dataset includes:

[0041] A retrieval query is generated by at least one generative data-driven model to obtain at least one reference dataset, at least one dataset to be analyzed, and / or a technical dataset.

[0042] The generated search query is provided to a database that includes at least one reference dataset, at least one dataset to be analyzed, and / or a technical dataset, to obtain at least one reference dataset, at least one dataset to be analyzed, and / or a technical dataset.

[0043] Based on the provided search query, obtain at least one reference dataset, at least one dataset to be analyzed, and / or a technical dataset from the database.

[0044] According to a second aspect, an apparatus is disclosed that includes corresponding components for performing or carrying out steps of the method according to the first aspect, or includes at least one processor and at least one memory for storing instructions that, when executed by at least one processor, cause the apparatus to perform at least the steps of the method according to the first aspect.

[0045] According to another aspect, the use of anomaly mitigation instructions generated by the method according to the first aspect or by the equipment according to the second aspect is disclosed, for displaying anomaly mitigation instructions to operators of chemical plants and / or for the production of chemical products.

[0046] According to another example aspect, a system for operating a chemical plant (e.g., to produce chemical products) is disclosed, the system comprising: an apparatus according to any aspect; a database (server) providing at least one reference dataset, at least one dataset to be analyzed, and / or a technical dataset, the database being communicatively coupled to the apparatus; and a server providing at least one generative data-driven model, the server being communicatively coupled to the apparatus, together performing at least the steps of the method according to the first aspect (and / or any implementation or example of the method and combinations thereof).

[0047] According to another example aspect, a computer element is disclosed that includes instructions that, when executed by a processor or computing device, perform or carry out steps according to the methods disclosed herein or defined by the devices disclosed herein.

[0048] According to another example aspect, a computer program or computer program product is disclosed that, when executed by a processor, causes a device, such as a server, to perform and / or control actions according to any aspect of the method.

[0049] According to another example aspect, a (e.g., tangible and / or non-transitory) computer-readable storage medium is disclosed, which includes a computer program that, when executed by a processor, causes a device, such as a server, to perform and / or control actions according to any aspect of the method.

[0050] Generative data-driven models (e.g., transformer-based models) can be trained on large datasets (e.g., several gigabytes of non-purposeful, unstructured text and / or image data). Trained generative data-driven models (such as trained transformer-based models) can have improved capabilities for predicting data patterns (such as patterns presented in natural language). This improved capability can be attributed to the large number of parameters acquired through the training. For example, transformer-based models (such as OpenAI GPT) can include 117 million parameters (GPT-1), 1.5 billion parameters, 175 billion parameters (GPT-3), and 170 trillion parameters (GPT-4). These parameters allow the GPT model to generate improved data outputs compared to other models that do not include transformer components (such as recurrent neural networks (RNNs) or long short-term memory (LSTM) networks).

[0051] The paper “Attention Is All You Need”, published by Vaswani et al. at the 31st Neural Information Processing Systems Conference (NIPS2017) in Long Beach, California (December 6, 2017, arXiv:1706.03762v5), describes a mechanism in machine learning that includes a transformer component (a transformer-based model), which is incorporated herein by reference.

[0052] ISO / IEC 23053:2022(en), ISO / IEC TR 24372:2021(en), ISO / IEC 22989, and ISO / IEC 23053 define standards in the fields of artificial intelligence (AI) and machine learning (ML). Big data is specified in ISO / IEC 20546:2019(en) Information technology—Big data. Data quality can be specified, for example, in ISO / IEC 20546:2019(en), ISO 8000-66:2021(en) / Data quality, and ISO / IEC DIS 5259-1(en).

[0053] Generative data-driven models, such as transformer-based architectures, can allow for capturing long-range dependencies and parallelization of computations. Furthermore, generative data-driven models such as transformer-based architectures can be pre-trained on large text-based datasets and fine-tuned (or retrained) for task-specific datasets with smaller (labeled) datasets. Fine-tuning can be a process of taking a pre-trained generative data-driven model (e.g., trained on a large dataset) and further training it on a smaller, specific dataset, which can allow the knowledge learned by the pre-trained model to be transferred to a specific task. During fine-tuning, the model's weights can be updated based on the provided specific dataset, where the pre-trained weights serve as a starting point, and, for example, only a small number of additional training steps are performed.

[0054] The fine-tuning or retraining allows the superior analytical capabilities of pre-trained transformer-based models to analyze data patterns beyond those found in natural language.

[0055] This disclosure relates to the use of generative transformer-based models when analyzing data patterns using a pre-trained transformer-based model for distributed production, such as chemical production. Data suitable for retraining or fine-tuning a pre-trained transformer-based model for said production can be obtained by providing historical production data accumulated over more than 150 years.

[0056] The amount of training data used to fine-tune (or retrain) a pre-trained transformer-based model can depend on the specific task, the model's complexity, and the desired performance level. In many cases, a smaller amount of specialized training data can be used to fine-tune (or retrain) a transformer model compared to the size of the pre-trained dataset. However, if the application-specific data differs significantly from the pre-trained data, retraining or fine-tuning a pre-trained transformer-based model may require a much larger dataset compared to scenarios where the application-specific data is similar to the pre-trained data. For example, text classification or sentiment analysis may require hundreds of megabytes to gigabytes of labeled data. Machine translation may require tens to hundreds of gigabytes of text data for retraining or fine-tuning. Question answering may require several gigabytes of training data to retrain a pre-trained transformer-based model.

[0057] Utilizing higher-quality training data can reduce the amount of data required for training (e.g., the term "data quality" is defined in ISO / IEC 20546:2019(en), ISO 8000-66:2021(en) / Data quality, and ISO / IEC DIS 5259-1(en)). Data preprocessing for generating high-quality training data can involve labeling data, removing noise, and removing irrelevant information. Therefore, preprocessing data to generate high-quality training data can facilitate more effective retraining or fine-tuning.

[0058] Larger amounts of training data allow for fine-tuning or retraining of pre-trained transformer-based models with more parameters. For example, fine-tuning the DistilBERT model may require less training data compared to fine-tuning the GPT-4 model. The larger number of parameters can also provide improved data analysis capabilities (e.g., GPT-4 is more powerful than GPT-3 in analyzing data patterns).

[0059] Generative data-driven models can be transformer-based models, such as TinyBERT, DistilBERT, Llama 7B, Mistral 7B, GPT-Neo, or larger GPT variants. Furthermore, pre-trained transformer-based models can be, for example, ChatGPT (GPT-2, GPT-3, GPT-3-5-turbo, GPT-4), Davinci, BERT (bidirectional encoder representation from a transformer), DistilBERT, Transformer-XL, XLNet (Extreme Language Understanding Network), T5 (text-to-text transfer transformer), RoBERTa (robustly optimized BERT method), ELECTRA (efficient encoder for learning accurate token substitutions), Reformer, Longformer, and DeBERTa (decoding-enhanced BERT with disentangled attention). Due to differences in architecture and / or pre-trained datasets, the properties of these transformer-based models, and therefore the output data, may differ for the same input data. Therefore, one or more of the transformer-based models (GPT-2, GPT-3, GPT-3-5-turbo, GPT-4, Davinci, BERT, DistilBERT, Transformer-XL, XLNet, T5, RoBERTa, ELECTRA, Reformer, Longformer, DeBERTa) can be used as an alternative to a pre-trained transformer-based model for technical purposes, or one or more of the models can be combined for technical purposes to provide multiple data outputs for supplementary data analysis.

[0060] The time required for fine-tuning or retraining can vary depending on the size of the generative data-driven model and the size of the dataset used for fine-tuning or retraining. Generative data-driven models with 1 million to 2 billion parameters (such as TinyBERT with 4.4 million parameters, DistilBERT with 66 million parameters, and GPT-Neo with 125 million parameters) can be fine-tuned or retrained in a few hours (e.g., 1 to 2 hours) using a dataset several gigabytes in size (e.g., corresponding to 10 years of production). This is less than the time required for a full production cycle (e.g., 15 days or more) and allows for rapid adaptation to any parameters downstream of the production cycle, such as those affecting subsequent production lines. Larger generative data-driven models (e.g., those with up to 10 billion parameters, such as Llama 7B and Mistral 7B) may take much longer (e.g., 1 to 2 days) to fine-tune on datasets of the same size.

[0061] In the context of this application, a generative data-driven model (e.g., a transformer-based model) (generative artificial intelligence AI model) can refer to a model that may include transformer components such as Figures 9 to 12 The basic model described in the text (machine learning ML model).

[0062] Generative artificial intelligence (AI) can refer to computer programs that can generate outputs, as described in ISO / IEC 23053:2022(en), ISO / IEC 23053:2022(en), ISO / IEC TR 24372:2021(en), ISO / IEC 22989, ISO / IEC 23053, ISO / IEC DIS 5259-1(en), and ISO / IEC 24661:2023(en). Generative AI programs can include ML models, such as transformer-based models (generative pre-trained transformer GPT models). Generative pre-trained transformer models (or simply transformer-based models) can also be referred to as base models.

[0063] The definition of big data is given in ISO / IEC 20546:2019(en) Information technology—Big data. For example, the definition of “data quality” is given in ISO / IEC 20546:2019(en), ISO 8000-66:2021(en) / Data quality, and ISO / IEC DIS 5259-1(en).

[0064] The paper “Attention Is All You Need”, published by Vaswani et al. at the 31st Neural Information Processing Systems Conference (NIPS2017) in Long Beach, California (December 6, 2017, arXiv:1706.03762v5), describes a mechanism in machine learning that includes a transformer component (a transformer-based model), which is incorporated herein by reference.

[0065] By using a pre-trained transformer-based model retrained on production data, early detection and removal of anomalies during the production cycle can allow for the recovery of product quality. In particular, early detection of anomalies during the affected production cycle (where anomalies occur) can allow the quality of intermediate products produced to be recovered during the affected production cycle or during the next production cycle, thereby obtaining a final product with acceptable quality at the end of the affected production cycle or at the end of the next production cycle.

[0066] Because the proposed generative artificial intelligence (AI) system has the speed and data analysis capabilities to determine deviations from data patterns, it can be used for anomaly detection in factory production data to detect deviations from standard (reference) data patterns obtained during standard factory operations.

[0067] In addition to detecting anomalies early in affected production cycles, aspects, embodiments, and examples of this disclosure provide a more user-friendly interface for anomaly detection. Furthermore, the efficiency of anomaly detection can be further improved by combining both plant operational documentation and production data. These aspects, embodiments, and examples can serve multiple different plants in a distributed production environment and can be used by plant operators without data science knowledge to discover anomalies in plant operations. Therefore, these aspects, embodiments, and examples provide improved methods and systems with enhanced user interfaces to effectively guide users (i.e., plant operators) through the process of identifying and removing anomalies associated with operations in the production cycle.

[0068] Optional continuous monitoring and notification systems can provide further advantages in improving production monitoring.

[0069] In production contexts, the term "real-time" can mean time intervals of several hours, depending on factors such as the amount of training data and the size of the chosen model. For example, analyzing 10 years of production data, resulting in, for instance, hundreds of gigabytes, might take several hours. In this example, anomalies might be detected within a few hours. In another example, analyzing 10 days of data, resulting in, for instance, several gigabytes, might take several minutes. In that example, anomaly detection might take several minutes. The amount of data can also depend on the complexity and size of the plant or part of the plant being analyzed. Attached Figure Description

[0070] In the following sections, embodiments of this disclosure will be outlined by way of examples. It should be understood that this disclosure is not limited to the described embodiments and / or examples. This disclosure is also described in conjunction with preferred embodiments and examples. However, by studying the accompanying drawings, the content of this disclosure, and the claims, those skilled in the art will understand and implement other variations of the claimed invention.

[0071] Figure 1 An implementation scheme for detecting anomalies in the production cycle is shown.

[0072] Figure 2 The implementation scheme for the production line is shown.

[0073] Figure 3A shows Figure 1 The operating system for the production line.

[0074] Figure 3B shows the situation caused by... Figure 1 The pre-trained transformer-based model is trained on the production data generated by the production line.

[0075] Figure 3C illustrates an implementation scheme that uses a trained transformer-based model to detect and remove anomalies in the production cycle.

[0076] Figure 4 The preprocessing engine of the operating system shown in Figure 3 is illustrated.

[0077] Figure 5 This shows production data associated with one or more production operations on a production line in a distributed production environment.

[0078] Figure 6 An implementation scheme for the user interface is shown.

[0079] Figure 7 Other aspects of training a pre-trained transformer-based model are shown.

[0080] Figure 8 The contextualization of the prompt is shown.

[0081] Figure 9 An implementation scheme for training the embedding layer is shown.

[0082] Figure 10A An implementation scheme of the converter encoder architecture is shown.

[0083] Figure 10B An implementation scheme of the converter decoder architecture is shown.

[0084] Figure 10C An implementation scheme of the converter encoder-decoder architecture is shown.

[0085] Figure 11 An implementation scheme for training and / or deploying a transformer encoder, transformer decoder, and / or transformer encoder-decoder is shown.

[0086] Figure 12 An implementation scheme for input embedding is shown. Detailed Implementation

[0087] The following embodiments are merely examples for implementing the methods, systems, devices, or application apparatuses disclosed herein and should not be considered limiting. The following description is intended to enhance understanding and should be understood as supplementing and reading in conjunction with the descriptions provided in the foregoing Summary and Embodiments section of this specification. Some aspects may have different terminology than, for example, that is provided in the foregoing description. However, those skilled in the art will understand that these terms refer to the same subject matter, for example, in a more specific manner.

[0088] Unless a specific order is disclosed, any steps presented herein can be performed in any order. The methods disclosed herein are not limited to a specific order of these steps unless a specific order is disclosed (e.g., pre-training of the transformer-based model precedes training; training precedes retraining; obtaining data for training precedes training the model; and / or what is directly and explicitly disclosed in this application).

[0089] There is no need to perform different steps at a specific location in a distributed system or in a specific computing engine; that is, each step can be performed at a different computing engine using different equipment / data processing.

[0090] As used herein, "determine" also includes "initiate or cause determination," "generate" also includes "initiate and / or cause generation," and "provide" also includes "initiate or cause determination, generation, selection, sending, and / or receiving." "Initiate or cause execution of an action" includes any processing signal that triggers a computing node or device to perform a corresponding action.

[0091] In the claims and in the description, the words “comprising,” “including,” or similar wording do not exclude other elements or steps, and the indefinite articles “a” or “an” do not exclude multiple. A single element or other unit may perform the function of several entities or items recited in the claims. The fact that certain measures are recited only in mutually different dependent claims does not mean that combinations of these measures cannot be used in advantageous embodiments.

[0092] Any disclosure and embodiments described herein relate to the methods, systems, apparatuses, and computer program elements listed above, and vice versa. Advantageously, the benefits provided by any embodiments and examples also apply to all other embodiments and examples, and vice versa.

[0093] All terms and definitions used in this document should be understood broadly and have their general meaning.

[0094] Figures 2 to 8 In the context of “engine”, it includes at least one computer processor.

[0095] In the context of this disclosure, "computer interface" can be, for example, a graphical user interface, an application programming interface, or a web-based interface.

[0096] For example, a pre-trained transformer-based model can be pre-trained for a first purpose (e.g., analyzing non-specific text-based data) and retrained / fine-tuned for a second purpose (e.g., production). The pre-trained transformer-based model can be further retrained / fine-tuned for the second purpose to better suit the second purpose, thereby improving the output data provided by the model.

[0097] A “distributed production environment” or “factory” can refer to, but is not limited to, any technological infrastructure used for industrial purposes of manufacturing, producing, or processing one or more products (i.e., manufacturing or production processes or processing performed by a distributed production environment). A distributed production environment can be a “factory” (infrastructure) with distributed units for production.

[0098] A distributed production environment can be a technological infrastructure (factory) comprising one or more production lines. The one or more production lines can include one or more production operations. These one or more production operations can be distributed.

[0099] A distributed production environment can be more than one plant distributed in space and / or for distributed operations. A distributed production environment can be one or more of the following: chemical plant, processing plant, pharmaceutical plant, fossil fuel processing facility (such as oil wells and / or natural gas wells), refinery, petrochemical plant, cracking plant, etc. A distributed production environment can even be any of the following: distillery, processing plant, or recycling plant. A distributed production environment can be any of the examples given above or a combination of their analogues.

[0100] A “product” produced at the end of a production cycle by one or more production lines in a distributed production environment can be, for example, any physical product, such as chemicals, biological products, pharmaceuticals, food, nutritional products, beverages, textiles, metals, plastics, semiconductors, cosmetics, or even any combination thereof. Additionally, or alternatively, the product can be a service product, such as through recycling or waste disposal (e.g., recycling), or chemical processing (e.g., decomposition or dissolution into one or more chemical products). Some non-limiting examples of chemical products are organic or inorganic compositions, monomers, polymers, foams, pesticides, herbicides, fertilizers, feed, nutritional products, precursors, pharmaceutical or therapeutic products, or any or more of their components or active ingredients. In some cases, a chemical product can be a product that can be used by an end user or consumer, such as a cosmetic or pharmaceutical composition. A chemical product can be a product that can be used to further manufacture one or more products; for example, a chemical product can be a synthetic foam that can be used to manufacture shoe soles or a coating that can be used on the exterior of automobiles. Chemical products can be in any form, such as solid, semi-solid, paste, liquid, emulsion, solution, granules, particles, or powder.

[0101] One or more production lines in a distributed production environment may include equipment or process units such as any one or more of the following: heat exchangers, towers (such as fractionation towers), furnaces, reaction chambers, cracking units, storage tanks, extruders, granulators, settlers, agitators, mixers, cutters, curing tubes, evaporators, filters, sieves, pipes, chimneys, valves, actuators, mills, transformers, conveying systems, circuit breakers, machinery (e.g., heavy rotating equipment such as turbines, generators, crushers, compressors, industrial fans, pumps), conveying elements (such as conveyor systems), motors, etc.

[0102] Furthermore, one or more production lines in a distributed production environment typically include multiple sensors and at least one control system for controlling at least one parameter or process parameter related to the production process in the plant. Such control functions are typically performed by the control system or controller in response to at least one measurement signal from at least one of the sensors. The plant's controller or control system can be implemented as a distributed control system (DCS) and / or a programmable logic controller (PLC). Multiple sensors can be distributed throughout the distributed production environment for monitoring and / or control purposes. Such sensors can generate large amounts of data. Sensors may or may not be considered part of the equipment. Therefore, production such as chemical and / or service production can be a data-intensive environment. Distributed production environments can generate large amounts of process-related data.

[0103] The sensors can be used to measure one or more process parameters and / or to measure the operating conditions of the equipment or parameters related to the equipment or process unit. For example, sensors can be used to measure process parameters such as flow rate in a pipe, liquid level in a tank, furnace temperature, chemical composition of a gas, etc., and some sensors can be used to measure vibration of a crusher, fan speed, valve opening, pipe corrosion, voltage across a transducer, etc. The differences between these sensors are not only based on the parameters they sense, but can even be based on the sensing principles used by the respective sensors. Some examples of sensors based on the parameters sensed by the sensor can include: temperature sensors, pressure sensors, radiation sensors such as light sensors, flow sensors, vibration sensors, displacement sensors, and chemical sensors such as those for detecting specific substances such as gases. Examples of sensors that differ in the sensing principles employed can be, for example: piezoelectric sensors, piezoresistive sensors, thermocouples, impedance sensors such as capacitive sensors and resistive sensors, etc.

[0104] A distributed production environment (including one or more production lines) can be multiple distributed production environments. Multiple distributed production environments can be coupled such that they can share one or more of their value chains, extracts, and / or products. Multiple distributed production environments can also be referred to as a complex, a composite site, an "integrated" or "integrated site." Such an integrated site or chemical industrial park can be or may include one or more distributed production environments, where products manufactured in at least one distributed production environment can serve as raw materials for another distributed production environment.

[0105] "Production" refers to any industrial process that provides an output product different from the input product when used on or applied to an input component. Therefore, production can be any manufacturing or processing process or a combination of processes to obtain the aforementioned product. A production process can even include the packaging and / or stacking of one or more products.

[0106] Production processes can be continuous or periodic; for example, a batch chemical production process can be used when a catalyst that needs to be recovered. A key difference between these production types lies in the frequency of occurrence of data generated during production. For example, in a batch process, production data extends from the beginning of the process to the last batch of different batches produced during that run. In a continuous setup, the data is more continuous due to potential changes in production operations and / or maintenance-driven downtime. Therefore, the required data analysis can vary depending on the data stream (batch or continuous). For example, periodic retraining or fine-tuning of a trained converter-based model may be advantageous for batch data streams, while continuous retraining or fine-tuning of a trained converter-based model may be advantageous for continuous data streams.

[0107] The terms "production data," "plant data," or "plant-based data" are used interchangeably and can refer to product data (e.g., product attributes), process data (process parameters), and operating conditions. Plant-based data can refer to data that includes values ​​(e.g., numerical or binary signal values) measured, for example, via one or more sensors during a production process. Process data can be time-series data of one or more process parameters and / or equipment operating conditions. Typically, plant-based data can include temporal information about process parameters and / or equipment operating conditions; for example, the data includes timestamps for at least some data points related to process parameters and / or equipment operating conditions. Plant-based data can include temporal-spatial data, i.e., time data and location or data associated with one or more physically separated equipment areas, allowing temporal-spatial relationships to be derived from the data.

[0108] "Process parameters" can refer to any variable related to the production process, such as any or more of the following: temperature, pressure, time, level, etc., which are related to the production of the product as defined above.

[0109] The above definitions of distributed production environment, products produced by distributed production environment, production process, data generated by production environment, and production control are merely examples and should not be construed as restrictive. It is understood that the systems and methods of the claimed invention can be applied to any kind of production that produces products and generates multi-parameter data streams related to production. Any type of factory-based data can be preprocessed by a preprocessing unit to generate the required factory-based training data (labeled data, (pre-)structured data, filtered data, data in digital and / or text formats, etc.) or factory-based input data, respectively suitable for retraining or fine-tuning a pre-trained transformer-based model or for production using a trained transformer-based model. Therefore, a distributed production environment should be broadly interpreted as a technological environment for producing products (physical products and / or services associated with the products; products can even be data products), and, during the production of said products, generating multi-parameter production-related data streams (factory-based data) through said technological environment.

[0110] Intermediate products from one production line can be used as input for the next production line in the production cycle. The final products produced at the end of a production cycle can be used as input for plants with different production processes and / or those producing different products. Therefore, the production processes and production lines in distributed manufacturing facilities are interconnected and influence the final product, posing challenging requirements for controlling and monitoring distributed production.

[0111] Figure 1 An implementation scheme for detecting anomalies in the production cycle is shown.

[0112] like Figure 1 As shown, chemical production may include one or more consecutive production cycles 1, 2…m. At the end of the one or more production cycles, the final product may be produced. The quality of intermediate / final products can be monitored to ensure compliance with standard quality requirements. When one or more quality parameters of the product (e.g., color, intensity, etc.) are within acceptable ranges (predefined ranges for one or more parameters), the product meets quality control (quality inspection). When one or more product parameters are outside the predefined ranges for one or more parameters, the product quality may be unacceptable. When the quality of a product (intermediate or final product) is unacceptable, it may need to be recycled or discarded, resulting in waste, and further, a waste of the production resources used to produce the product.

[0113] The one or more production cycles may include one or more interconnected production lines 1, 2, ..., n. The front end of one or more production lines may receive intermediate products from the rear end of the same production line. The front end of the first production line may receive raw materials. At the end of each of the one or more production lines, the quality of the output product can be measured according to the same principle as described above. At the end of each of the one or more production lines 1, 2...(n-1), intermediate products can be produced, which can become input products for continuous production lines. At the end of the production cycle m, which ends at the last production line n, the final product can be produced. The quality of the final product can be measured or quantified according to the principles described above.

[0114] When an anomaly occurs in a production cycle (e.g., affected production line 2 of affected production cycle 1), the quality of the product (e.g., intermediate product, quality check 2) may be unacceptable. If the unacceptable quality is not corrected, supplying intermediate products from the back end of the affected production line (production line 2 of production cycle 1) may result in the final product also being of unacceptable quality (e.g., the final product of production line n of production cycle 1). If the source of the anomaly is not identified and removed at the beginning of a subsequent production cycle (e.g., production cycle 2 in this example), the quality of intermediate or final products in the subsequent production cycle (cycle 2) may also be compromised, leading to an accumulation of product waste and a waste of production resources until the source of the anomaly is finally corrected.

[0115] During the production cycle, the operating parameters of each production line 1, 2…n can be measured. Based on the measurements, the operation of each production line 1, 2…n can be monitored and / or controlled. In other words, based on the measurements, the operating parameters of each production line 1, 2…n can be adjusted.

[0116] The measured production data (factory data) can be used to analyze typical data patterns associated with the standard operating conditions of the production line.

[0117] To ensure and maintain high-quality production cycles, plant operators (users, production line operators) can leverage their operational knowledge associated with one or more production operations on the production line. Plant operators can further analyze the measured production data (plant data) to identify anomalies in data patterns (e.g., deviations of one or more data points from predefined ranges (e.g., standard / normal ranges) of one or more parameters). Monitoring anomalies in the data can reduce the risk of product quality degradation during affected and continuous production cycles.

[0118] The production process of the final chemical product can typically be carried out on several production lines 1, 2…n (dividing the entire process into sub-processes). Quality checks are performed at the end of each production line. Quality checks mean that the product quality at the end of each production line is within acceptable ranges for one or more parameters (e.g., strength, color, and other properties) measured on the produced product. Furthermore, a final quality check (such as strength, color, and other properties) can be performed on the final product at the end of the production cycle.

[0119] If an abnormal behavior occurs in one or more production lines 1, 2…n, due to the features claimed in this invention (i.e., using a trained transformer-based model to analyze data patterns in the production data), the abnormality can be resolved as it occurs, without waiting until the end of the entire cycle. For example, it may be until… Figure 1 The anomaly is only identified after the affected cycle 1 has ended. Cycle 1 may include one or more interconnected production lines 1, 2, ..., n, where each of the one or more production lines may include one or more interconnected production operations.

[0120] The interconnected production lines, including interconnected production operations, result in complex chemical production processes. These interconnected production lines and operations influence each other, affecting the quality of the final product, complicating processes, and consequently placing high demands on data analysis, monitoring, and control in chemical production.

[0121] For example, on a production line (e.g., affected production line 2), the temperature of a sensor might rise, for instance, by 10 degrees Celsius (°C) above the temperature it should be heated to under standard operating conditions. In this example, the operator can adjust the operating conditions in successive steps so that the quality of the final or intermediate product in the affected or subsequent production cycles is not compromised. Detecting this abnormal temperature rise and notifying the operator of this anomaly as early as possible (e.g., at the end of the affected cycle or before the start of the next cycle) allows the operator to adjust the operating conditions to restore the quality of the final product in the affected cycle or the next production cycle.

[0122] Another example could be a production line where quality checks on one or more product parameters fail to meet standards (one or more product parameters may exceed predetermined ranges). Once such an anomaly in one or more product parameters is detected and notified to the plant operators, the operators can adjust one or more production operations. For example, the operator can decide what needs adjustment instead of adding the output product of the affected production line to the overall process (e.g., purified / distilled / synthesized raw materials or intermediate components). For instance, the operator can correct the issue and restore the quality to an acceptable level, or if restoration is not possible, the operator can decide whether the cycle should continue. This reduces resource waste.

[0123] The aspects, implementation methods, and examples provided herein can relate to providing systems and methods applicable to different plants without requiring separate investigations on a case-by-case basis. By using the systems and methods of the claimed invention, operators may not need to log anomalies in plant data. Operators may not need to provide case studies to a data scientist team to investigate how to remove the source of the anomaly.

[0124] The features of this disclosure provide operators (e.g., human operators) with a user-friendly computer interface to trigger the analysis of production data. Aspects, embodiments, and examples of this disclosure can allow for the restoration of production quality by detecting anomalies and identifying their sources.

[0125] The user interface may include interface elements that provide functionality to translate natural language into database queries for retrieving production data. The user interface may be integrated across one or more plants in a distributed production environment. The user interface can provide portable services across various production lines, operations, and facilities producing a variety of (interrelated) products, thereby improving the integration of anomaly detection systems in a distributed production environment.

[0126] The aspects, embodiments, and examples of this disclosure can allow for the rapid detection and identification of sources of anomalies in the production cycle, thereby allowing for the restoration of the quality of intermediate / final products in the production cycle. By removing the source of the anomaly as early as possible, waste of product and production resources can be reduced. The advanced manufacturing technologies described in this application can eliminate the time-consuming work required to eliminate anomalies and reduce waste. Therefore, the advanced manufacturing technologies described in this application can provide climate change mitigation technologies for chemical production.

[0127] Figure 2 The implementation scheme for the production line is shown.

[0128] A distributed production environment may include equipment 1002 and sensors 1003 that generate one or more sensor-related data streams. The distributed production environment may produce one or more products as defined above, wherein attributes of the one or more products can be measured, extracted, or calculated to generate one or more product-related data streams (e.g., color, intensity, etc., as shown in...). Figure 1 (As described in the context). Plant data 1005 (plant-based data or production data) may include data obtained from each of the one or more data streams.

[0129] Equipment 1002 can be any equipment in a distributed production environment, such as pumps, heat exchangers, valves, reaction vessels, separation chambers, etc.

[0130] Sensor 1003 can be any type of sensor in a distributed production environment, such as temperature sensors, flow sensors, pressure sensors, etc.

[0131] One or more products produced by one or more production lines can be any type of intermediate or final product, as described in the general description above. Product properties can be measured, for example, by gas chromatography.

[0132] Factory data (production data) 1005 can be stored in a database, for example, as historical production data. Factory data 1005 can be provided to a preprocessing engine, which preprocesses the data and provides factory-based input data to an analysis engine 1004. This analysis engine can analyze the factory-based input data and generate machine-readable instructions 1007 for controlling and / or monitoring the engine 1006. The generation of machine-readable instructions 1007 can be performed automatically by the analysis engine 1004 based on the analysis of the input factory data (i.e., without operator intervention). For example, the analysis engine 1004 can continuously receive input factory data and analyze the data in a continuous mode. When an anomaly occurs, the analysis engine can generate machine-readable instructions 1007 for the control and / or monitoring engine to remove the anomaly. The analysis engine can determine solutions to improve production efficiency by analyzing factory input data in the context of the production process and can send machine-readable instructions 1007 to the control and / or monitoring engine to improve production (e.g., machine-readable instructions 1007 regarding how to remove the source of the anomaly). The control and / or monitoring engine 1004 can display push notifications to the operator to review instructions 1007, and based on this review, can generate machine-readable instructions 1008. Based on the machine-readable instructions 1008, the control system can change the operating parameters of one or more pieces of equipment 1002. Reviewing the machine-readable instructions can allow trained transducer-based models to be safely integrated into distributed production environments such as chemical production.

[0133] Alternatively, in addition to the machine-readable instructions 1007 being automatically generated by the analysis engine, the operator may prompt the analysis engine 1004 to provide machine-readable instructions 1007 based on prompts.

[0134] Machine-readable instruction 1007 can be used by an operator to control and / or monitor one or more production operations in a distributed production environment.

[0135] Machine-readable instructions 1007 may include operational instructions for production, such as machine-readable instructions for controlling equipment 1002 and / or machine-readable instructions for monitoring equipment 1002 and / or sensors 1003. Machine-readable instructions 1007 may also include operational instructions for operators controlling and / or monitoring a distributed production environment, wherein the operator may be a human-based operator, a computer-based operating system, or a hybrid system including human operators and computer-assisted systems.

[0136] A control system for a distributed production environment may include one or more computing units that can manipulate one or more parameters related to the production process by controlling one or more actuator or switch and / or end effector units, such as by manipulating one or more equipment operating conditions. Control is typically performed in response to one or more signals retrieved from the equipment.

[0137] The control and monitoring engine may include one or more computer processors for modifying machine-readable instructions 1007 and generating machine-readable instructions 1008 for a control system used in distributed production. The control system for distributed production can adjust equipment operating conditions based on the machine-readable instructions 1008, such that the adjusted process parameters and / or equipment operating conditions produce a controlled product (such as a chemical product) with one or more desired or predetermined properties or performance parameters. Therefore, production can be controlled in real time while ensuring that equipment operating conditions adapt to undesirable changes in process parameters.

[0138] It is understood that the control and monitoring of a distributed production environment generally involves controlling the equipment and / or production lines used to produce products by sending machine-readable instructions to the production environment. "Product" should be interpreted broadly as described above.

[0139] As in Figure 3a , Figure 3b , Figures 9 to 12The transformer-based model (first ML model) operated by the AI ​​engine, as described in the context, can be integrated with another computer program or a second ML model via a computer interface. The second ML can include algorithms different from the first ML model. For example, the second ML can be classic ML that is not based on a transformer architecture; a variant of the first model based on a transformer architecture, etc. Alternatively, the second ML can include the same architecture as the first ML model, but can be trained on a different dataset compared to the first ML model. Training on different datasets can provide different properties to the first and second ML models. These different properties can allow the first and second ML models to be used in combination or as alternatives.

[0140] For example, the second ML model can be a data-driven model, such as the one disclosed in WO2021156157 (A1), which can be integrated with the first model (the transformer-based model) via a computer interface. The second ML model can be used to preprocess the raw factory-based data to generate factory-based training data. Preprocessing of the raw data may involve noise removal, data filtering, data labeling, data sorting, converting the raw data to different formats, converting operator data to different formats more suitable for the first ML model (the transformer-based model), and so on. For example, preprocessing may include converting the training data into text and / or numbers, and using the text and / or numbers for training, such as... Figures 9 to 12 As described in the context. Alternatively, instead of or in addition to a second ML model, preprocessing can be performed by classical computer algorithms that do not involve ML (e.g., data labeling may not require ML). Preprocessing of the raw factory-based data can improve the quality of the training or input data based on the raw data and can reduce the computational power required to train / use the first ML model (transformer-based model).

[0141] Figure 3A shows Figure 1 and Figure 2 The operating system for the production line.

[0142] The operating system includes (re)training or fine-tuning a pre-trained converter-based model and using the trained converter-based model to control and / or monitor production, such as chemical production. The training of the pre-trained converter-based model is further described in the context of Figure 3B. The purpose of using the trained converter-based model is further described in the context of Figure 3C.

[0143] Figure 3B shows the situation caused by... Figure 1 and Figure 2The pre-trained transformer-based model is trained on the production data generated by the production line.

[0144] Such as in Figures 9 to 12 As described in the context, training (fine-tuning, retraining) a pre-trained converter-based model includes accessing the pre-trained converter-based model via a computer interface (GUI, API, web-based interface) and receiving factory history data 1010 (as an example of a reference dataset) from a preprocessing engine 1009 via the interface. An operator retraining or fine-tuning the pre-trained converter-based model may prompt the pre-trained converter-based model to access the factory history data via the computer interface, or alternatively, the operator may upload historical factory data from a database to an AI engine via the computer interface. The AI ​​engine may include at least one processor that operates the pre-trained converter-based model. The operator retraining or fine-tuning the pre-trained model may be a human operator, an automated operating system including a computer processor, or a hybrid operating system including both human operators and computer-aided operating systems, prompting the pre-trained model to train, fine-tune, or retrain on factory-based data.

[0145] Factory historical data 1010 can be based on factory data 1005. Factory data 1005 and factory historical data can be stored in one or more databases (e.g., provided by one or more servers). Factory data 1005 may include, for example, data from... Figure 5 This describes multiple production data points within the context of [the data]. During training, historical factory data can be embedded via an embedding layer, such as [example data]. Figure 9 The context described above. Embedded factory data can generate embedded factory data.

[0146] The above examples of factory data are merely illustrative of possible ways to implement various aspects and embodiments of this disclosure. Factory data should be interpreted broadly and should be understood as any kind of data associated with the production of products by the technical infrastructure. Any type or format of data associated with production can be preprocessed by a preprocessing engine to be suitable for training or production using transformer-based models, for example, preprocessed into text and / or numbers.

[0147] At the end of the training cycle, the analytics engine 1004 can output (publish, provide, generate) a trained transformer-based model suitable for use in distributed production environments such as chemical production. The published trained model can be stored in a database for purposes such as version control. The published trained model can be a computer program product. Access to the published trained model can be provided to users as a data service to aid production.

[0148] (For example, in a reference dataset or technical training dataset) Factory-based training data can be factory-based data in one or more languages. Training a pre-trained transformer-based model on factory-based training data in one or more languages ​​allows for scaling up the factory-based training dataset and, more importantly, allows for providing operational instructions for production in one or more languages. Providing instructions in one or more languages ​​can improve user interaction with the trained transformer-based model.

[0149] Figure 3C illustrates an implementation scheme that uses a trained transformer-based model to detect anomalies in the production cycle and remove the sources of anomalies in the production cycle.

[0150] Factory data 1005 can be provided to preprocessing engine 1004 for preprocessing. Preprocessed data 1005 may include, for example, data from... Figure 4 The steps described in the context. Preprocessing may also involve other steps required to provide training / input data based on factory data suitable for training or using a transformer-based model, such as labeling data, removing noise, constructing unstructured data, converting data into different formats, etc., for production (e.g., text / number).

[0151] Preprocessed data from preprocessing engine 1004 can be used as factory-based input data for a trained transformer-based model for production.

[0152] A trained transducer-based model can receive factory-based input data and, for example, predict anomalies in the factory-based input data. The trained transducer-based model can identify the source of the anomaly. The trained transducer-based model can notify the operator. The notification may include instructions to the operator. Alternatively or additionally, the analysis engine 1004 can generate machine-readable instructions 1007 for controlling and / or monitoring the engine 1006.

[0153] The analytics engine 1004, which operates a pre-trained transformer-based model, may have at least one computer interface for operating the transformer-based model (uploading data, providing prompts such as prompts intended to review the output "Response" generated by the model, etc.). The transformer-based model (untrained, pre-trained, or trained) may be stored in a database or in the cloud. The model may be operated by the AI ​​engine via at least one computer interface (e.g., API; GUI).

[0154] An operator of a trained transformer-based model can provide input, such as text and / or audio queries, to the model via a computer interface. The operators of the trained transformer-based model can additionally provide, for example, extracted data points from data provided by a preprocessing engine. These extracted data points may, for example, relate to outlier data points in a distributed production environment, typical or optimal operating parameters. These extracted data points, along with the operator's query, can form part of the input data to the trained transformer-based model.

[0155] The trained transformer-based model can process input data and provide solutions to user queries, for example, in the form of machine-readable instructions 1007.

[0156] The control and / or monitoring engine can review machine-readable instructions 1007.

[0157] Following review, the control and / or monitoring engine 1006 can generate machine-readable instructions 1008, which may be identical to, partially based on, or different from machine-readable instructions 1007. The operator of the control and / or monitoring engine can use the machine-readable instructions 1007 solely for monitoring production, or can forward the machine-readable instructions 1007 into machine-readable instructions 1008 for controlling production. The operator can be a human operator and / or an operating system including the processor, and optionally a human operator.

[0158] The control and / or monitoring engine can trigger retraining or fine-tuning of the trained transformer-based model. Retraining or fine-tuning can be triggered by an operator such as a human operator or by an automated system including a computer processor. Alternatively or concurrently, retraining or fine-tuning can be scheduled or continuous.

[0159] A trained generative data-driven model (e.g., a transformer-based model) can be trained on factory-based training data in one or more languages. The trained generative data-driven model (e.g., a transformer-based model) can provide operation instructions for production in one or more languages. Providing operation instructions in one or more languages ​​may be advantageous, for example, because the training data may be more available in one language than in another. A trained generative data-driven model (e.g., a transformer-based model) trained on factory-based training data in one or more languages ​​can be configured to provide operation instructions in the language where most of the training data is available. Alternatively, the trained generative data-driven model (e.g., a transformer-based model) can be configured to provide operation instructions for production in a user-selected language, making the model more user-friendly. Alternatively, the trained generative data-driven model (e.g., a transformer-based model) can be requested to provide operation instructions for production in more than one language to cross-check the operation instructions and select the most suitable instruction for improved production.

[0160] Trained generative data-driven models (e.g., transformer-based models) can be used to continuously monitor and notify the system. When an anomaly is automatically detected by an analytics engine operating on the trained generative data-driven model (e.g., a transformer-based model), the analytics engine can send a notification to the operator informing them of the detected anomaly.

[0161] Figure 4 The preprocessing engine of the operating system is shown in Figure 3.

[0162] As in Figure 4 and Figure 5 The raw factory data 1005 described in the context can be provided to the preprocessing engine 1009.

[0163] The preprocessing engine 1009 can preprocess raw factory-based data to provide factory-based training data and / or factory-based input data to generative data-driven models (e.g., transformer-based models). Preprocessing steps may include selecting desired parameters, merging / aggregating, computing factory-based training data (e.g., computing derived parameters), removing outliers, etc. Preprocessing may include filtering data, removing noise, labeling data, sorting data, and converting data from formats unsuitable for training / using the model to formats suitable for training / using the model, etc. The output data of the preprocessing engine can be stored in a database and used as factory-based training data for retraining or fine-tuning pre-trained generative data-driven models (e.g., transformer-based models), as shown in Figure 3. Figure 7 and Figures 9 to 12As described in the context, the output data of the preprocessing engine can also be used as input data for trained generative data-driven models, such as transformer-based models.

[0164] Figure 5 This shows production data associated with one or more production operations on a production line in a distributed production environment.

[0165] It can be done via a computer interface, such as in Figure 6 The user interface described in the context receives plant data from the distributed production environment. Alternatively, operators may prompt the model to access data via alternative computer interfaces such as APIs or web-based interfaces.

[0166] Factory data can include different categories of factory data, such as sensor data, operational data, factory metadata, and analytical data.

[0167] Sensor data can refer to quantities that are available for measurement in a production plant using installed sensors (such as temperature sensors, pressure sensors, flow rate sensors, etc.).

[0168] The analytical data may involve quantities provided by analytical measurements of samples taken from any point in the production plant, such as the composition of reactants, starting materials, products and / or byproducts determined, for example, by gas chromatography from samples taken from different stages of the production process (e.g., before or after the catalytic reactor).

[0169] Operational data may involve raw data (basic data, unprocessed analytical data, and / or sensor data) or processed or derived parameters (derived directly or indirectly from raw data).

[0170] Plant metadata can indicate the physical plant layout and can include plant-specific quantities that describe attributes such as reactors. These plant-specific quantities are predefined by the physical plant layout and can be related to plant or reactor performance.

[0171] Factory data 1005 can include text and / or numbers (structured data). Factory data can also be unstructured. Unstructured data (such as scans, data tables including images, QR codes, etc.) can be preprocessed by a preprocessing engine and transformed into a format suitable for training, such as... Figures 9 to 12 The text and / or numbers described in the context of a pre-trained transformer-based model.

[0172] Figure 6 An implementation scheme for the user interface is shown.

[0173] A user interface for operating the method according to the first aspect (e.g., including operating a transformer-based model) may include elements of a user interface that provide the user with the following functionalities: selecting / uploading training data (as an example of a reference dataset); selecting / uploading new data; uploading engineering documents; detecting anomalies; generating machine-readable instructions (“Response”) for monitoring and / or removing anomalies; and outputting the “Response” as text. Selecting / uploading training / new data and / or engineering documents (as an example of a technical training dataset) can be implemented through one or more user prompts that guide the transformer-based model to select / upload training / new data and / or engineering documents. The transformer-based model may be prompted to analyze new data (as an example of a dataset to be analyzed) and provide a “Response” (as an example of analysis instructions) to user queries, such as whether / once anomalies exist in the new data, and if so, how to remove the source of the anomalies in the production cycle. In response to user prompts, the transformer-based model can provide a “Response” that includes identifiers of one or more anomalies, identifiers of one or more sources of the one or more anomalies, and / or machine-readable instructions for removing the sources of said one or more anomalies. The machine-readable instructions may include text and / or control instructions for monitoring and / or controlling equipment and / or sensors on one or more production lines. Control instructions may include machine-readable instructions 1007 (operation instructions), such as instructions for changing operating parameters of equipment 1002, predictive maintenance (e.g., replacing sensor 1003), etc. Text instructions may include operator-specific instructions, such as reviewing anomalies in data, reviewing one or more sources of anomalies, and instructions on how to remove the source of the anomaly, etc.

[0174] Figure 7 Other aspects of training pre-trained generative data-driven models (e.g., transformer-based models) are shown.

[0175] Pre-trained generative data-driven models (e.g., transformer-based models) that have been pre-trained on text and / or numbers can be trained based on production-related training data. Figures 9 to 12 (As shown) for use in production. Training data based on historical factory data (as an example of a reference dataset) can be converted into text / numbers by a computer processor (e.g., a preprocessing engine).

[0176] The (re)training or fine-tuning of a pre-trained generative data-driven model can include one or more training epochs. These one or more training epochs can use one or more training datasets. For example... Figure 7As shown, the training data associated with production may include training data 10131 for a first training cycle and training data 10132 for a second training cycle. Training data 10131 may include factory documentation data (e.g., engineering documents, operation manuals, technical reports, definitions of normal factory conditions, e.g., ranges of one or more parameters associated with one or more production operations as an example of a technical training dataset). Training data 10132 (as an example of a reference dataset) may include historical factory data (e.g., as in... Figure 5 (The production data described in the context of 1005).

[0177] For the first and second training cycles, the following can be followed: Figure 2 b and Figures 9 to 12 The steps described within the context. In this case, used for training in Figures 9 to 12 In the context of the generative data-driven model described, the generic text / numerical data can be replaced by training data 10131 and / or 1032 (which can be converted to text and / or numeric formats). During training (the first training epoch and / or the second training epoch), training data 10131 and / or training data 10132 can be embedded in the embedding layer, as shown in... Figure 9 As described in the context.

[0178] After training in the first training cycle, the trained generative data-driven model can be trained (parameterized) to analyze the plant data context, such as the plant's normal / standard operating conditions.

[0179] The trained generative data-driven model trained during the first training cycle can be further trained in the second training cycle using training data 10132. Training data 10132 may include data when the plant operates under normal conditions (including conditions without predefined ranges of operating parameters). At the end of the second training cycle, a trained generative data-driven model can be generated, wherein the model becomes suitable for analyzing production data patterns under normal / standard operating conditions (i.e., conditions including one or more parameters related to one or more production operations, wherein the one or more parameters are within a predetermined range, such as a standard / normal range).

[0180] Example 1 (Natural language prompts are converted into database queries).

[0181] As in Figures 9 to 12 The pre-trained transformer model described in the context can be trained on a set of database queries for one or more database query languages ​​to provide a trained transformer-based model suitable for generating database queries based on requests.

[0182] During the first and / or second training cycles, a factory operator training a pre-trained transformer-based model can select training data 10131, 10132 via natural language prompts. The operator's natural language can be interpreted as a database query by the trained transformer-based model (trained using a set of database queries) to retrieve training data 10131, 10132 from one or more databases.

[0183] Distributed production facilities (such as chemical production) may include one or more databases (e.g., SQL databases, EDL-enterprise data lakes, etc.) associated with one or more database systems. The one or more database systems may connect to connectors / APIs to retrieve data from the one or more databases. Optionally, access to the one or more databases may include an authorization process requesting user / operator authorization. Operators or users may require access permissions to access the one or more databases.

[0184] Transformer-based models can be trained to interpret natural language instructions, such as “get the production data of factory area keyword Z between time X and Y”.

[0185] When trained to interpret natural language to identify which database to access, a transformer-based model can generate appropriate database queries (DB queries) for, for example, an EDL production database. The connector / API can allow the execution of DB queries and the retrieval of requested data from one or more database systems. DB queries and data retrieval can be extended to different types of databases using different query languages. In this case, a pre-trained or trained transformer-based model (trained in a first training cycle and / or a second training cycle) can be trained in one or more query languages ​​to analyze DB queries and retrieve data from one or more database systems including said one or more database query languages. In this case, the training data used to train the pre-trained or trained transformer-based model (trained in the first training cycle and / or the second training cycle) can be a collection of DB queries in one or more database query languages.

[0186] Therefore, after pre-training on text data (general text) and additionally training on one or more DB query languages, a trained transformer-based model (or generative data-driven model) can be able to interpret the operator's natural language into database queries to access relevant databases, thus eliminating the need for the operator to write relevant DB queries and setting up API connections to one or more databases. When the trained transformer-based model, pre-trained on general text and trained on DB query languages, is further trained on the first training data 10131 and / or the second training data 10132, the trained transformer-based model can not only retrieve one or more production-based data (10131, 10132; 1005; 1011) requested by the operator, but the model can also analyze the data.

[0187] When access / authorization is required to access one or more databases, a trained transformer-based model trained on a DB query set can prompt the operator / user to authorize / grant access to the database, where the requested data may be stored. The trained transformer-based model can then be automatically retrieved by a computing processor, without the user having to write DB queries or set up connectors to access the database. Automated data retrieval performed by the computer processor can accelerate the retrieval of one or more production-based data sets (10131, 10132; 1005; 1011) to train a pre-trained transformer-based model and / or use the trained transformer-based model, for example, to detect anomalies occurring in the production cycle, and optionally help the operator remove the source of the anomaly. For example, a trained transformer-based model can retrieve data requested by prompting "Retrieve production data for factory area keyword Z between 10:10 and 18:00 on April 7, 2023". The user / operator can then prompt the model to analyze the data and identify anomalies within it. Users / operators can further prompt the model to identify the source of any detected anomalies. Users / operators can also prompt the model to provide instructions on how to remove the source of the anomaly. Accelerated data retrieval (since users / operators do not need to write DB queries and set up connectors / APIs) can speed up anomaly detection and the removal of anomaly sources in the production cycle, leading to more efficient chemical production, reduced waste, and mitigation of climate change.

[0188] Example 2 (data sorting, labeling, selection for training, defining data quality for effective training of a pre-trained transformer-based model).

[0189] Operators can define additional constraints or definitions for selecting training data (10132). For example, an operator can select training data within a predefined standard range associated with standard operating conditions of the production line (e.g., when one or more parameters associated with one or more production operations of the production line are within a predefined range). The operator can label data within the predefined standard range as "high-quality data." An operator can select data outside the predefined standard range and label that data as "low-quality data." An operator can select data associated with extreme operating conditions (e.g., under extreme operating conditions, a piece of equipment 1003 might explode, for example, when one or more operating parameters exceed predefined extreme thresholds) and label that data as "extreme data." Sort, label, classify, and select categories of data used for training can improve training efficiency by, for example, reducing computational resources, increasing training speed, and improving model properties, enabling more effective and earlier identification and removal of sources of anomalies during the production cycle.

[0190] Training a pre-trained converter-based model on “extreme data” can improve the safety of controlling and / or monitoring distributed production. When a pre-trained converter-based model is trained on extreme data, the operator can prompt the trained model to generate one or more operational parameters associated with one or more production operations, where these parameters do not reach extreme thresholds. Upon prompting, the trained converter-based model can generate machine-readable instructions that do not violate the safety range of one or more parameters associated with one or more production operations in distributed production (chemical production), thereby improving the safety of the production environment and simultaneously providing a trained converter-based model for safe deployment in chemical production.

[0191] Using “high-quality data” to train a pre-trained transformer-based model can improve the model’s ability to identify anomalies in production data.

[0192] Production cycles in a manufacturing plant can be continuous processes. For example, the rate of certain raw materials flowing through a reactor can be continuously increased until an optimal rate (full load) is reached. The operator can sort, label, and / or select "high-quality data" by choosing data associated with this optimal rate (e.g., the mass flow rate of raw material A = 1000 kg / h). The operator can also sort, label, and / or select "high-quality data" by choosing data points associated with one or more predefined thresholds (e.g., the mass flow rate of raw material A > 1000 kg / h > 800 kg / h). At the end of the production cycle, the rate slowly decreases to zero or residual flow. To select production cycles operating at full load and optimal conditions, the operator can specify data selection criteria (e.g., acquiring data after the desired flow rate of 1000 kg / h is reached), thereby selecting data labeled as "high-quality data" for training.

[0193] Example 3 (Upload training or input data).

[0194] The operator can upload data from one or more data categories (technical documents, plant-based data, etc.) as described in the context of Implementation Scheme 2 to train the model. The operator can also upload data (e.g., production-based input data 1011, 1005 over a period of time) for analysis by the trained transformer-based model. Having the operator upload data, rather than prompting the transformer-based model to access data in storage, allows for faster data analysis because an authorization process for data access may not be necessary (e.g., the operator may have data already prepared for upload).

[0195] The operator's data upload function can be implemented as a supplement to or alternative to the function of providing data via database query, as described in the context of implementation scheme 1.

[0196] In situations where data is not stored in a robust storage system (e.g., the API / connector is not fully developed); pre-trained (or trained in the first and / or second training cycles) transformer-based models may not yet be trained on the relevant DB query language; users may encounter issues with data access control (licenses); or rapid analysis of small amounts of data may be sufficient or necessary, making the feature of uploading data via a computer interface advantageous. Specifically, in such cases, the operator can manually upload training data via the upload feature of the computer interface, thereby eliminating the aforementioned problems associated with accessing storage devices or databases.

[0197] Once the pre-trained transformer-based model has been trained in the first and second training epochs, it acquires an embedding layer based on the context and data patterns associated with the training datasets 10131 and 10132. Due to this embedding layer, the trained transformer-based model gains the ability to generate data distributions associated with, for example, standard operating conditions, extreme operating conditions, and suboptimal operating conditions associated with “poor-quality data.” Therefore, the trained transformer-based model can compare these standard, extreme, and suboptimal data distributions with real-time production data to identify any anomalies.

[0198] Detected anomalies can be responded to to factory operators via a user interface in natural language (including, for example, text strings and / or audio signals). Outputting these responses to operators in natural language can provide improved human-machine interaction, as operators may face time and psychological pressure when dealing with extreme situations related to abnormal behavior in the production environment. Providing operators with instructions in natural language enables them to act more quickly (e.g., compared to scenarios where the output is provided as error message codes, etc., when the operator needs to explain the error or find a solution) to remove the source of the anomaly, thereby restoring the quality of intermediate or final products in the production cycle.

[0199] Alternatively, or as an alternative to providing a response in natural language (text or audio), the response may include generating a technical report that includes, for example, a description of one or more anomalies, and optionally a description of the source of the one or more anomalies, and optionally operational instructions for removing the one or more sources of the anomalies in a production environment. The report may also be generated in a planned manner or upon request from the operator.

[0200] At the end of each training cycle, the model's trainer (the operator can be a human or a computer processor) can, as follows: Figure 8 The trained model is tested as described in the context, and based on the test, the trainer can run additional training cycles or release the trained model for production, for example, as... Figure 3a , Figure 3c and Figure 7 , Figure 8 As described in the context.

[0201] The published, trained transformer-based model can be stored in a database or in the cloud. The published model can be operated by computing units / nodes, including computer processors (such as analytics engines). A copy of the trained transformer-based model can be made available to users for use and / or further training. Alternatively, users can be given only access to operate the trained transformer-based model.

[0202] Replacement Figures 9 to 12 In the context described, the pre-trained transformer-based model can be used by the analytics engine 1004 with another base model, such as ChatGPT (GPT-2, GPT-3, GPT-3-5-turbo, GPT-4 or higher / similar), Davinci, BERT (bidirectional encoder representation from transformers), DistilBERT, Transformer-XL, XLNet (Extreme Language Understanding Network), T5 (text-to-text transfer transformer), RoBERTa (robustly optimized BERT method), ELECTRA (efficient encoder for learning accurate classification of token replacements), Reformer, Longformer, DeBERTa (decoding-enhanced BERT with disentangled attention), or any other large-scale language model pre-trained on big data such as general text, images, and videos.

[0203] Depending on the availability of factory-based training data, the availability of computing resources, and the required accuracy for analyzing the data, users can choose a base model with fewer or more parameters. Figures 9 to 12 The transformer-based model shown can be pre-trained to publish a pre-trained model with the required number of parameters to meet the user's technical objectives. Figure 2 b、 Figures 9 to 12 In the context of [the context], pre-trained models can be further improved to release models with even fewer parameters, thereby increasing computational speed and reducing computational resources.

[0204] exist Figure 7 And / or the retraining or fine-tuning described in the context of Figure 3B can be continuous. Retraining or fine-tuning can be performed in the context of ongoing production without stopping / interrupting production.

[0205] Figure 8 It shows how to use such Figure 7 The multi-step task process for chemical production described in the context of the model trained in the first and second training cycles.

[0206] Used in Figure 7The model trained in the context may include the sequence of steps 1 through 11 described below.

[0207] Step 1: The plant operator may intend to detect one or more anomalies in production data 1005 of plant A.

[0208] Step 2: The operator can log in to the system (e.g., Analysis Engine 1004).

[0209] Step 3: The operator can use prompts including natural language, such as “Select production data for factory area keyword ‘Factory A’ from April 1, 2022 to May 31, 2023, where the mass flow rate of F123X is >1000kg / h”, to select training data based on factory historical data 1010.

[0210] Step 4: The system (e.g., analytics engine 1004) can analyze the prompts and generate database queries for the production database PIMS, for example:

[0211] ProductionData

[0212] | where Areakey=="Plant A"

[0213] and time between (datetime(2022-04-01).. datetime(2023-05-31))

[0214] Step 5: The system (e.g., analytics engine 1004) can execute the generated query and retrieve data from the PIMS (as prompted by the step 3 request).

[0215] Step 6: The operator can be prompted to upload factory-specific documents to generate training data for additional training of the model.

[0216] Step 7: The system (e.g., analytics engine 1004) can use the training data and the uploaded documents to generate an anomaly detection model, where the anomaly detection model refers to the trained transformer-based model trained in the context of steps 1 through 6 above.

[0217] Step 8: Simultaneously, the operator can select new data (input data 1011 based on production data 1005) for anomaly detection by providing a prompt to the trained transformer-based model as described in the context of Step 7, where the prompt can include natural language, such as “select production data of factory A from June 1, 2023, based on the factory area keyword 'Factory A'”.

[0218] Step 9: The system (e.g., analytics engine 1004) can use the database query generated by PIMS in step 4 to retrieve the new data requested in step 8.

[0219] Step 10: The system (e.g., analysis engine 1004) can compare the data distribution under normal / standard operating conditions (embedded in the trained transformer-based model during, for example, the second training cycle as described above) with the new data (requested in step 8) to detect any anomalies in the new data.

[0220] Step 11: The system (e.g., analytics engine 1004) can generate a notification (“Response”) for the operator, such as: “On June 5, 2023, the temperature of sensor T7235 suddenly dropped by 10 degrees and remained at a similar temperature for the rest of the time. The value of F6324 remained normal, but suddenly increased by 12% more than expected between 16:04 and 17:15 on June 17, 2023, before stabilizing. The average pressure of P7692 was 30% higher than expected during this period.”

[0221] Figure 9 An implementation scheme for training the embedding layer is shown.

[0222] Embedding layers can be obtained by training, for example, a Continuous Bag-of-Words (CBOW) model or a skip-gram model. Embedding layers are suitable for generating embedded input data based on input data. Generating embedded input data can refer to embedding input data. Embedding input data can produce a representation associated with the input data. Therefore, embedded input 114 can be a representation associated with the input data. The input data can include one or more elements. One or more elements can be represented by input vector 106. In particular, embedded input 114 and / or input vector 106 can be machine-readable and / or processable by a processor. For this purpose, embedded input 114 and / or input vector 106 can be tensors, particularly first-order tensors. Specifically, input vector 106 can be a one-hot vector or the sum of multiple one-hot vectors. A one-hot vector can be a vector with one entry that is not equal to zero. Examples of one-hot vectors can be 108, 110, and 112. The non-zero entries in the one-hot vector and / or input vector 106 can indicate elements. For example, a lookup table can define the relationship between the position of a non-zero entry and the element indicated by the one-hot vector. A lookup table can specify multiple distinct elements. The number of distinct elements can be equal to the number of entries in a one-hot vector. The number of distinct elements can be referred to as the vocabulary. In one example, elements can be represented by tokens, and a sequence of elements can refer to at least a portion of a sentence. At least a portion of a sentence can be represented by multiple tokens. Tokens can represent at least a portion of an element and / or a word. For example, in the case where an element will be associated with only one word, words such as “embeddings,” “embedding,” or “embed” will constitute distinct elements. The first token can represent the stem “embed,” and the suffixes that typically appear in multiple words can be represented by the second, third, and fourth tokens. The second, third, and fourth tokens can be used to represent other words, such as “look,” “looking,” etc., preferably together with a fifth token representing the stem “look.” Ultimately, this tokenization of elements associated with multiple stems and multiple suffixes results in using fewer tokens to represent multiple elements, and therefore using fewer computational resources.

[0223] A lookup table specifying, for example, a subset of the English vocabulary can include 10,000 or more words. The embedded input 114 can be a lower-dimensional representation than the input vector 106. For example, a typical embedded input 114 can include hundreds of different entries. Next, the embedded input 114 uses fewer computational resources to construct a dense representation of one or more elements. Furthermore, the embedded input 114 can represent relationships between two or more elements. For example, the words “Italy” and “Germany” can be similar or more closely related because they both define European countries, while the word “embodiment” can be quite different from both. The smaller the dot product between two embedded inputs 114, the more similar the two elements associated with the embedded input 114 can be. Therefore, the embedded input 114 can accurately represent one or more elements and produce accurate results based on processing the embedded input 114.

[0224] To transform input vector 106 into embedded input 114, the embedding layer may include a number of neurons equal to the number of entries in embedded input 114. Based on embedded input 114, the output layer may generate output vector 116. The output vector may be a vector and / or may indicate one or more elements. Output vector 116 may indicate one or more elements different from input vector 106 and / or the one-hot vector associated with input vector 106. For this purpose, the output layer may include a number of neurons equal to the number of entries in input vector 106 and / or output vector 116. The output layer may apply a softmax function to embedded input 114. By doing so, the output vector may include non-zero probabilities associated with elements of entries associated with output vector 116. Therefore, one or more elements can be obtained from output vector 116 with corresponding probabilities. Where input vector 106 may specify one or more sequences of elements, output vector 116 may specify one or more elements corresponding to the sequence of elements specified by input vector 106. Figure 9 In the example, the element associated with vector 118 corresponds to the input vector with a 71% probability. Additional or alternative elements correspond to the input vector as indicated by the output vector with a lower probability. By defining a threshold that can be compared with the probability, the selection of corresponding elements can be customized according to the user's needs. The elements generated by the model, including embedding layer 102 and output layer 104, can refer to the most probable element indicated by output vector 116. Therefore, Figure 9 The model described in the paper can generate elements associated with a vector 118 with a confidence score of 71%.

[0225] Figure 9The model can be a Continuous Bag-of-Words (CBOW) model. The CBOW model can be trained based on a training dataset comprising multiple input vectors and corresponding output vectors. Since the training dataset may be unlabeled, the training of the CBOW model can be considered self-supervised. Before training the CBOW model, it can be initialized with random values ​​for the weights assigned to neurons. During the training of the CBOW model, the input vectors can be passed through the initialized embedding and output layers, and the loss can be determined by comparing the output vector obtained by passing the input vector 106 through the model with the output vector corresponding to the input vector 106 specified by the training dataset. Based on the determined loss, backpropagation can be applied to determine the gradients associated with the neurons in the embedding layer 102 and the output layer 104 to reduce the loss. Based on the determined gradients, the weights of the neurons can be updated using a gradient descent algorithm. If the CBOW model achieves the predetermined loss, training can be terminated, and a trained CBOW model can be obtained. Based on the trained CBOW model, the embedding layer 102 can be adapted to embed input data comprising one or more elements. This embedding layer 102 can be used in other machine learning architectures that require an embedding layer 102, such as in Figure 10A , Figure 10B and Figure 10C The context describes the transformer encoder, transformer decoder, or transformer encoder-decoder architecture. To train these architectures, a trained embedding layer 102 may be required. Therefore, a model, such as the CBOW model, can be trained before training the transformer encoder, transformer decoder, or transformer encoder-decoder architecture.

[0226] Figure 10A An implementation scheme of the converter encoder architecture is shown.

[0227] The converter encoder includes an encoder input 278, one or more encoder blocks 274, 214, and an encoder output. The converter encoder architecture can be derived from those known in the art and... Figure 10C The converter encoder-decoder architecture shown is derived. Specifically, the converter encoder can be referred to as an X-former. The converter encoder architecture can correspond to an encoder architecture associated with a converter encoder-decoder architecture having additional encoder outputs, rather than a decoder directly connected to the converter encoder-decoder architecture. Multiple converter encoder architectures may exist in the art, such as the bidirectional encoder representation (BERT) from the converter.

[0228] Input data can be received at encoder input 278. Input embedding 202 can be applied to encoder input 278. Applying input embedding 202 can refer to passing the input data through an embedding layer, for example, as... Figure 9As described in the context. Furthermore, position coding 204 can be applied to encoder input 278. Applying position coding 204 can refer to adding a position factor to the embedded input obtained via input embedding. Preferably, the input data can specify a sequence of elements. Position factor p pos It can indicate the position of an element within a sequence.

[0229] For example, the position factor p can be obtained based on the following equation. pos :

[0230]

[0231] Where pos can refer to the position of an element in the sequence, i can refer to the dimension associated with the input embedding, and d can refer to the dimension of the model, such as a transformer decoder, transformer encoder, or transformer encoder-decoder. This can be referred to as absolute position embedding. Alternatively, position encoding can be based on Rotated Position Embedding (RoPE). Position encoding is advantageous because it can handle sequential data without requiring an additional dimension to indicate the position of each element. Subsequently, position encoding 204 reduces the computational resources required to embed the input data. By passing the input data through the encoder input, the input data can be transformed into a second-order tensor representing the sequence of elements. This second-order tensor can be referred to as embedded input data. This embedded input data can be processed by the encoder block. This embedded input data can be provided to layer normalization 208 via residual connections. Multi-head self-attention 206 can be applied to this embedded input data. Multi-head self-attention 206 can include two components: multi-head and self-attention. Self-attention can be understood as a filter applied to the embedded input data. By applying the filter to the embedded input data, elements associated with the embedded input data that contribute to the output data to be generated can be identified to generate the output data. Therefore, a filter can represent the degree of contribution of the elements associated with the embedded input data to the output data to be generated. Applying a filter can be referred to as weighting the elements associated with the embedded input data. This is particularly advantageous for long sequences of elements. Filters can be learned and improved during training by learning to identify the contributions of the elements associated with the embedded input data. For example, in a partial sentence “I went to the bakery to buy a”, the last word can be generated by a data-driven model such as a transformer encoder. Self-attention can cause the transformer encoder to focus on the main words “bakery” and “buy” to generate the word “bread”. Self-attention can refer to attention generated based on the input data. Therefore, a filter can be determined based on the input data (preferably, the embedded input data). The embedded input data can be used as a query Q, keyword K, and value V regarding the self-attention operation. Self-attention can refer to attention based on the received input data. Therefore, a filter can be calculated based on the following formula by inserting the corresponding tensor based on the embedded input data:

[0232]

[0233] Where d k The dimensions corresponding to keywords.

[0234] To further improve the efficiency of the converter encoder, multiple heads are used to apply filters, resulting in a multi-head self-attention 206. The multi-head self-attention 206 can include applying filters to two or more portions of the embedded input data. Therefore, the tensor can be divided into two or more portions, and filters can be applied to two or more portions separately by two or more heads according to the following equation:

[0235]

[0236] The parameter matrix is Where i can refer to the number of heads, dv, dk, and d Q It can refer to the value, keywords, and query dimensions.

[0237] The results of two or more heads can be cascaded according to the following equation:

[0238]

[0239] in And h can refer to the number of heads.

[0240] The embedded input data can be transformed into a context tensor via multi-head self-attention 206. This context tensor can represent a sequence of elements and the relationship between two or more elements of the input data. The context tensor can be a second-order tensor and / or may include one or more first-order tensors. After multi-head self-attention 206, layer normalization 208 can be applied based on the context tensor and / or the embedded input data from residual connections. Applying layer normalization 208 can refer to normalizing the context tensor. Normalizing the context tensor reduces the values ​​of the entries in the context tensor. This reduces the computational cost associated with processing the context tensor. After layer normalization 208, the context tensor can be passed back to feedforward layer 210, and then layer normalization 212 is performed based on the residual connections to the context tensor and / or the output of feedforward layer 210. Feedforward layer 210 can be a feedforward neural network. The feedforward neural network may include multiple fully connected neurons. Passing the context tensor through the feedforward neural network results in a linear transformation of the context tensor. Alternatively or additionally, the neural network may include one or more activation functions, such as rectified linear unit (ReLU). Thus, the neural network can be configured to perform one or more nonlinear operations on the context tensor and / or nonlinearly transform the context tensor. After the context tensor has been transformed and / or normalized by feedforward layer 210 and layer normalization 212, the context tensor can be provided to one or more additional encoder blocks 214. After passing the context tensor through feedforward layer 210, the context tensor can be adjusted for processing by additional attention layers of one or more additional encoder blocks 214 to apply a self-attention filter, preferably multi-head self-attention 206. The context vector transformed by layer normalization 212 and feedforward layer 210 can be referred to as the hidden state.

[0241] The encoder output 276 includes a linear layer 216 and a softmax layer 218. The linear layer 216 transforms the context vector into a logit vector. The linear layer can be fully connected. The logit vector obtained by passing the context tensor through the linear layer 216 can be passed through the softmax layer 218. Passing the logit vector through the softmax layer 218 means applying the softmax function to the logit vector. Applying the softmax function to the logit vector produces a probability distribution of one or more elements corresponding to the sequence of elements in the input data. One or more elements can be selected based on the probability distribution according to a predefined selection criterion. The one or more selected elements can be referred to as one or more elements generated by the transformer encoder. The one or more generated elements can be provided as encoder input to generate another one or more elements corresponding to the sequence of input data, as well as the one or more elements generated by the transformer encoder, as shown in... Figure 11 Described within the context of [the text].

[0242] Figure 10B An implementation scheme of the converter decoder architecture is shown.

[0243] The converter decoder includes a decoder input 284, one or more decoder blocks 280, 232, and a decoder output 292. The converter decoder architecture can be derived from those known in the art and... Figure 10C The transformer encoder-decoder architecture shown is derived. The transformer decoder can be referred to as the X-former. The transformer decoder architecture can correspond to the decoder architecture associated with the transformer encoder-decoder architecture, regardless of whether one or more hidden states are received from the encoder of the transformer encoder-decoder. Several transformer decoder architectures are available in the art, such as the generalized pre-trained transformer (GPT).

[0244] The decoder input 284 can be applied to, for example, in Figure 10A The input embedding 202 and position encoding 204 described within the context are similar to the input embedding 220 and position encoding 222.

[0245] Decoder block 280 may include layer normalization 226, masked multi-head self-attention 224, feedforward layer 228, and / or layer normalization 230. Embedded input data generated by passing input data through decoder input 284 can be provided to layer normalization 226 via residual connections. Furthermore, masked multi-head self-attention 224 can be applied to the embedded input data. Masked multi-head self-attention 224 corresponds to... Figure 10A The multi-head self-attention 206 described in the context further masks a portion of the embedded input data associated with elements in the sequence that are later than the element to be generated. Alternatively or additionally, portions of the input data associated with elements in the sequence that are later than the element to be generated may not be received and / or transformed into embedded input data. Therefore, the transformer decoder can be adapted to generate subsequent elements of the sequence, while the transformer encoder can be adapted to generate missing elements within a sequence and / or between two or more sequences. Thus, the transformer encoder can be configured for classification tasks. The transformer decoder can be configured for text generation.

[0246] Similar to Figure 10A The transformer encoder described within the context can generate a context tensor by applying masked multi-head self-attention 224 and layer normalization 226. This context tensor can be provided to layer normalization 230 via residual connections. Furthermore, feedforward layer 228 and layer normalization 230 can be similar to... Figure 10AThe context tensor describes the feedforward layer 210 and layer normalization 212. This context tensor can be provided to one or more additional decoder blocks 232.

[0247] The decoder output 292 may include a linear layer 234 and a softmax layer 236. The linear layer 234 and softmax layer 236 can be similar to... Figure 10A The linear layer 216 and softmax layer 218 are described within the context of the above.

[0248] Figure 10C An implementation scheme of a converter encoder-decoder architecture is illustrated. The converter encoder-decoder may include an encoder input 288, one or more encoder blocks 286, 264, a decoder input 294, a decoder block 290, and a decoder output 292. The encoder input 288 may correspond to... Figure 10A The encoder input is 278. One or more encoder blocks 286, 264 can correspond to... Figure 10A One or more encoder blocks 274, 214. Decoder input 294 can correspond to Figure 10B The decoder input is 284.

[0249] Decoder block 290 may include masked multi-head self-attention 270, layer normalization 272, feedforward layer 238, and layer normalization 240, similar to... Figure 2 The masking multi-head self-attention 224, layer normalization 226, feedforward layer 228, and layer normalization 230 described in the context of B. Decoder block 290 may also include multi-head self-attention 250 and layer normalization 248. Similar to... Figure 10B The description suggests that the context tensor can be obtained from the masked multi-head self-attention 270 and layer normalization 272. Similar to... Figure 10A Multi-head self-attention 206 and multi-head self-attention 250 can be applied to the context vector obtained from layer normalization 272 and the hidden states of one or more encoder blocks 286, 264. Layer normalization 248 can be applied to the context vector obtained from multi-head self-attention 250 and the context vector obtained from layer normalization 272 provided via residual connections. Similar to... Figure 10B The context vector generated by layer normalization 240 can be processed via feedforward layer 238 and layer normalization 248. The context vector generated by layer normalization 240 can be provided to another decoder block 242 similar to decoder block 290. The context vector obtained from one or more decoder blocks 290, 242 can be provided to decoder output 292. Decoder output 292 can correspond to... Figure 10B The decoder outputs 282.

[0250] Using the above architecture, the converter encoder-decoder can receive and process input data at encoder input 288 and one or more encoder blocks 286, 264, and decoder block 290 and decoder output 292. Based on the input data, the converter encoder-decoder can generate output data partially or sequentially. Sequentially generated output data can be provided to decoder input 294, one or more decoder blocks 290, 242, and decoder output 292, and / or can be processed by decoder input 294, one or more decoder blocks 290, 242, and decoder output 292. Preferably, a sequence can be provided to encoder input 288, and after at least a portion of the generated output data has been generated, at least a portion of the elements of the generated output data can be provided to decoder input 294. By doing so, since the converter encoder-decoder can receive more data over time, the next element of the output data can be generated with higher accuracy by considering both the input data and the generated output data.

[0251] Due to the transformer encoder-decoder architecture, the transformer encoder-decoder can be configured to transform a sequence into another representation of the sequence. An example of transforming a sequence into another representation could be translating a sentence into another language. Several transformer encoder-decoders are available in the art, such as BART, T5, etc.

[0252] In one implementation, layer normalization 208, 212 can be applied before the masked multi-head self-attention 224, multi-head self-attention 206, and / or feedforward layer 210 in the transcoder, transcoder, and / or transcoder-decoder. By doing so, the computational resources used to apply multi-head self-attention 206 and / or feedforward layer 210 to the embedded input data and / or context tensors can be reduced, since the entries of the corresponding tensors may be lower after normalization.

[0253] In one implementation, the decoder output 292 may include a classification neural network, additional feedforward layers, convolutional layers, fully connected layers, and so on. For example, the transformer encoder-decoder may be configured to select among multiple options. To this end, the transformer encoder-decoder may be provided with three different input datasets and may classify the context vectors obtained from one or more decoder blocks 290 via one or more linear layers. The architecture may then be extended according to the use case to be addressed. [1]

[0254] Figure 11 An implementation scheme for training and / or deploying a transformer encoder, transformer decoder, and / or transformer encoder-decoder is shown.

[0255] The encoder / decoder / encoder-decoder architecture 302 can correspond to, for example, in Figures 10A to 10C The converter decoder, converter encoder, and / or converter encoder-decoder described within the context of the converter.

[0256] The output data generated by the encoder / decoder / encoder-decoder architecture 302 may include one or more elements, specifically a sequence of elements. Previously generated elements of the output data can be provided as input for the next element in the sequence used to generate the output data.

[0257] exist Figure 11In the example, the input data may include N elements, specifically input tokens. Input tokens may be tokens specifically designed for input into a data-driven model (such as a transformer decoder, transformer encoder, or transformer encoder-decoder). The output data to be generated may include M elements. The encoder / decoder / encoder-decoder architecture 302 may generate one element of the output data based on elements received from the input data and optionally previously generated output data at a certain time step. Therefore, M time steps are required to generate M elements. The time steps include providing inputs 310, 312, and 314 to the encoder / decoder / encoder-decoder architecture 302 and receiving output data 304, 308, and 306 from the encoder / decoder / encoder-decoder architecture 302. In the first time step, input 310 may include N input tokens. The N input tokens may be associated with, for example, N words, stems, or word endings. Preferably, the N input tokens may specify a question. One or more input tokens may specify the start and / or end of a token sequence. Input 310 can be processed by encoder / decoder / encoder-decoder architecture 302. Based on input 310, at least a portion of output data 304 can be generated. At least a portion of the output data may include a first output token. In the next time step, the generated first output token may be provided together with input 312. Specifically, if input 312 can be received by converter encoder-decoder, the input token may be received at encoder input 288, and the first output token may be received at decoder input 294. If input 312 can be received by converter encoder, input 312 may be received by encoder input 278, and similarly by converter decoder and decoder input 284. Based on input 312, output data 308 including a first output token and a second output token can be generated. Generating output data 308 based on input 312 may refer to generating a second token based on the first token and N input tokens, where the first token may have already been generated based on the N input tokens. This process can be repeated until the last token in the sequence of output data 306 can be generated. Preferably, the last token may be an end token. An end token can terminate the generation of another output token.

[0258] Similarly, for data processing during the deployment of the encoder / decoder / encoder-decoder architecture 302, the encoder / decoder / encoder-decoder architecture 302 can be trained. The training dataset can include multiple sequences, each comprising multiple elements. Sequences can be associated with input data and / or output data. Alternatively or additionally, sequences can be independent of the input data and / or output data. For example, where the input and output data can refer to chemical components represented via text, the training dataset can include sequential text data independent of chemical components. In this example, the training dataset can include word sequences derived from a dialogue. In one embodiment, the training dataset can at least partially include the input dataset and / or the output dataset.

[0259] Training can be initialized by initializing the encoder / decoder / encoder-decoder architecture 302. In one implementation, the parameters associated with the encoder / decoder / encoder-decoder architecture 302 can be randomly initialized. Alternatively or additionally, the input embeddings of the encoder / decoder / encoder-decoder architecture 302 can be obtained by training a CBOW model or a skip-gram model, as in... Figure 9 As described in the context, a trained embedding layer can be used during training. The parameters associated with the embedding layer can remain constant and / or can be updated after a predefined number of training epochs. By doing so, the number of parameters to be updated is reduced, resulting in faster training with less computational resource consumption. Furthermore, the accuracy associated with the embedding layer can be constant and / or can be increased by avoiding error compensation associated with the newly initialized encoder / decoder / encoder-decoder architecture 302.

[0260] During training of the encoder / decoder / encoder-decoder architecture 302, at least a portion of the sequence of the training dataset can be provided to the encoder / decoder / encoder-decoder architecture 302 one by one, and one or more elements can be generated based on the sequence of the training dataset one by one. Elements generated based on the sequence may follow elements of the portion of the sequence that may have been provided to the encoder / decoder / encoder-decoder architecture 302. The generated one or more elements can be compared with one or more elements following at least a portion of the sequence provided to the encoder / decoder / encoder-decoder architecture 302 as specified by the training dataset. Therefore, during training, the encoder / decoder / encoder-decoder architecture 302 can generate a guess about the next element, and the guess about the next element in the sequence can be compared with the baseline truth specifying the actual next element according to the training dataset. Based on the guess about the next element and the baseline truth, a loss can be determined. The loss can define the similarity between the guess about the next element and the baseline truth. The loss can be determined by forming a vector dot product between the tokens associated with one or more elements and the tokens associated with the baseline truth. A loss that is not equal to zero can lead to an update of the parameters associated with the encoder / decoder / encoder-decoder architecture 302. Preferably, the parameters associated with the encoder / decoder / encoder-decoder architecture 302 can be independent of the embedding layer. For example, the parameters associated with the encoder / decoder / encoder-decoder architecture 302 can be the weights of the neurons in the encoder / decoder / encoder-decoder architecture 302.

[0261] Based on the determined loss, backpropagation can be applied to determine the gradients of the parameters associated with the encoder / decoder / encoder-decoder architecture 302 to reduce the loss. According to the determined gradients, the parameters associated with the encoder / decoder / encoder-decoder architecture 302, preferably the weights of the neurons associated with the encoder / decoder / encoder-decoder architecture 302, can be updated using a gradient descent algorithm.

[0262] The training dataset can be unlabeled. The sequence of elements within the training dataset can inherently include benchmark truth for determining the loss about one or more elements generated during training of the encoder / decoder / encoder-decoder architecture 302. Therefore, the encoder / decoder / encoder-decoder architecture 302 can be trained self-supervised. This is advantageous because it saves time and resources used to create labeled training datasets. Furthermore, it enables the use of large training datasets associated with sizes in the trillions of bytes. Therefore, the data-driven model can be accurate in generating the elements of the sequence. Moreover, large training datasets enable a small number of predictions or even zero predictions. Thus, the data-driven model trained as described above is versatile and helps save the resources required to train and / or host multiple purpose-driven models such as CNNs. The training described above can be referred to as pre-training. The data-driven model can be configured to perform a small number of predictions or even zero predictions about multiple use cases after pre-training. The performance of the data-driven model can be further improved through additional training, known as fine-tuning.

[0263] Figure 12 An implementation scheme for input embedding is shown.

[0264] When the sequence of elements associated with the input data (preferably included in the input data) can be of one type, methods such as those used in... Figures 10A to 10C The input embeddings described in the context are 202, 220, 252, and 266. For example, the type of input data can be text, where elements can be associated with at least a portion of a word, punctuation marks, start tokens specifying the beginning of one or more sequences associated with the input data, and / or end tokens. In another example, the input data can be at least partially numeric. Therefore, the input data can include multiple numbers. Numerical input data can be, for example, tabular data. Tabular data can specify one or more rows and / or one or more columns. Therefore, tabular data can include one or more cells, where cells can be associated with one or more numeric values.

[0265] Numeric input data may require different embeddings than text input data. Input embeddings for numeric input data can include token embeddings, position embeddings, column embeddings, row embeddings, or combinations thereof.

[0266] Applying token embedding to one or more elements (specifically, tokens associated with input data) can produce a machine-processable representation associated with one or more elements (specifically, tokens). Applying token embedding to one or more elements can refer to passing one or more elements through an embedding layer, for example, as in... Figure 9As described in the context of [the previous section], token embeddings can specify one or more elements, specifically tokens in a machine-processable representation. For example, token embeddings can transform numerical values ​​into vectors. This is advantageous because the representation can be enriched with more information, such as the position of the token within the sequence and / or the position of the token within a table associated with the token sequence. Positional embeddings can be similar to [the previous section]... Figure 9 , Figures 10A to 10C The location embedding is described within the context of the table. Column embedding can be applied when the input data can be tabular data. Applying column embedding to one or more elements (particularly tokens associated with the input data) can produce a machine-processable representation specifying the position of one or more elements within table 402 (preferably within columns of table 402). Applying column embedding can refer to adding column factors to the input data embedded via token embedding, particularly embedded input data. Column factors can be the same for elements associated with the same column, and / or different for two or more elements associated with different columns. Similarly, row embedding can be applied when the input data can be tabular data. Applying row embedding to one or more elements (particularly tokens associated with the input data) can produce a machine-processable representation specifying the position of one or more elements within table 402 (preferably within rows of table 402). Applying row embedding can refer to adding column factors to the input data embedded via token embedding, particularly embedded input data. Row factors can be the same for elements associated with the same row, and / or different for two or more elements associated with different rows.

[0267] In one implementation, the input data may be at least partially numerical and at least partially textual. Therefore, the input data may include two or more types of data. The data type may refer to a modality. Next, different embeddings can be applied to the input data. For the portion of the input data that includes text, embeddings can be applied... Figure 9 , Figures 10A to 10CThe input embeddings referenced herein. For the portion of input data that is embedded as a numeric token, positional embedding, column embedding, and row embedding can be applied. Furthermore, segmented embeddings can be applied to the input data, regardless of the type of input data. Segmented embeddings can specify the type of input data that can be associated with one or more elements. For example, if the input data includes both text and numbers, the input data can include both types of input data. Applying segmented embeddings to input data can refer to adding segmentation factors to the input data, preferably embedded input data and / or input data after applying token embeddings. Segmentation factors can specify the type of data associated with one or more elements. Segmentation factors can be the same for one or more elements associated with the same type of input data, and / or can differ between two or more elements associated with different types of input data.

[0268] Applying token embedding, position embedding, segment embedding, column embedding, row embedding, or a combination thereof can produce embedded input data and / or can be the output of any of encoder inputs 278, 284, 288 or decoder inputs 284, 294. Data obtained by applying token embedding, position embedding, segment embedding, column embedding, row embedding, or a combination thereof can be processed by encoder blocks 274, 286, decoder blocks 280, 290, encoder output 276, and decoder outputs 292, 282.

[0269] The following embodiments describe various aspects of this disclosure and other possible ways of implementing the embodiments for use in industrial environments for chemical production to mitigate climate change through advanced manufacturing technologies.

[0270] Training data 1013 based on the factory, such as in Figures 9 to 12 The untrained / pre-trained transformer-based models described in the context of and / or as in Figure 2 a, Figure 2 c. Figure 6 and Figure 8 The trained transformer-based models described in the context can be stored in a database, on an electronic data carrier, or in the cloud. For models suitable for training pre-trained transformer-based models, such as those described in... Figures 9 to 12 The untrained / pre-trained transformer-based models described in the context and / or as in Figure 2 a, Figure 2 c. Figure 6 and Figure 8Access to the factory-based training data 1013 of the trained transformer-based model described in the context can be granted to the user as a computer-readable token. The computer-readable token can be an authorization key generated by a computer processor based on a request from a requesting computation node associated with the user. This request can be sent to an authorization engine comprising at least one computer processor authorized to grant one or more authorization keys to access one or more of the factory-based training data and / or untrained, pre-trained, trained, retrained, or fine-tuned transformer-based models, as described in [the context of the previous sentence]. Figure 2 a, Figure 2 c. Figure 6 , Figure 8 and Figures 9 to 12 As described in the context.

[0271] Users can access training data via a token to process the data and receive processing results without receiving the actual training data. Users can also receive the actual factory-based training data or a portion of the factory-based training data. Depending on access permissions, users (upon receiving a token) can use one or more of untrained, pre-trained, and / or trained transformer-based models to achieve their technical objectives. Users can use the claimed factory-based training data or their own training data to train / retrain one or more models they are licensed to use. The access token for using the claimed factory-based training data can be used in conjunction with the token for using... Figure 2 a, Figure 2 c. Figure 6 , Figure 8 and Figures 9 to 12 The access tokens for one or more models described in the context may be the same or different. Each of the data products (i.e., based on the factory's training data) or data services has a separate access token (i.e., using a token such as in...). Figure 2 a, Figure 2 c. Figure 6 , Figure 8 and Figures 9 to 12 One or more models described in the context can enhance the security of using data products and / or data services. Access tokens may include user credentials, one or more generation algorithms, user authentication, two-factor authentication, etc.

[0272] Example 4 (Migration of one or more embedded layers for various tasks associated with one or more production operations)

[0273] like Figure 9The data embeddings described in the context can be transferred from one or more trained transformer-based models trained on one or more training datasets to another transformer-based model that has not yet been trained using said one or more training datasets. For example, a first dataset including non-text-specific data can be used to train a transformer-based model to generate a first pre-trained transformer-based model including the first embedded data, as in Figures 9 to 12 The first embedded data can be transferred from a first pre-trained transformer-based model to another transformer-based model. Furthermore, the first pre-trained transformer-based model can be trained on a second dataset including one or more database queries in one or more database query languages ​​to generate a first trained transformer-based model, which is trained to generate database queries based on prompts including instructions provided in natural language from an operator. The first trained transformer-based model can include embedded data based on a second training dataset. This embedded data can be transferred to one or more pre-trained or trained transformer-based models to generate database queries.

[0274] A pre-trained or first-trained transformer-based model can be trained on first training data 10131 and / or second training data 10132 to generate a second-trained transformer-based model and / or a third-trained transformer-based model. These second-trained and / or third-trained transformer-based models are trained to analyze production-related data. The second-trained and / or third-trained transformer-based models may include embedded data based on the first training data 10131 and / or second training data 10132. This embedded data can be transferred to a pre-trained or trained transformer-based model that has not yet been trained using the first training data 10131 and / or second training data 10132.

[0275] The transfer of embedded data between transformer-based models can accelerate training and improve the properties of trained models to perform one or more tasks associated with one or more production operations in a distributed production environment. For example, embedded data from a first training dataset embedded in a first transformer-based model can be transferred to a second transformer-based model that has better properties than the first transformer-based model, but which has not yet been trained on the first training dataset. Therefore, through embedding transfer, the second, better model can acquire the embedded data of the first model in a faster manner. Using the second transformer-based model may be more advantageous due to its initial properties (properties before embedding transfer). Furthermore, the second transformer-based model can acquire the embedded data of the first model, thus combining the properties of the first model (embedded data) with the properties of the second model (initial properties).

[0276] It is understood that, unless otherwise disclosed in this application, the features of the above and below embodiments are combinable.

[0277] The following example implementation plans should also be made public:

[0278] Clause 1:

[0279] A computer-implemented method for mitigating climate change by reducing product waste and production resource waste through detecting one or more anomalies in the production cycle of distributed chemical production, identifying one or more sources of the anomalies, and removing one or more sources of the anomalies, the method comprising:

[0280] Access to the trained transformer-based model is received by the computer processor;

[0281] Receive plant-based data via a computer interface that is associated with one or more production operations of one or more production lines of the distributed chemical production;

[0282] The trained transformer-based model is prompted to analyze the input factory-based data to detect one or more anomalies in the input factory-based data, and when one or more anomalies are present, to identify one or more sources of the anomalies, and to identify one or more instructions for removing one or more sources of the one or more anomalies.

[0283] Clause 2:

[0284] The computer-implemented method according to Clause 1 further includes:

[0285] The trained transformer-based model generates a technical report, which includes descriptions indicating the presence or absence of the one or more anomalies. Optionally, if the one or more anomalies are detected, the technical report also includes indications of one or more sources of the anomalies.

[0286] Optionally, the technical report may also include one or more operational instructions for removing the one or more sources of the one or more anomalies.

[0287] Optionally, the operation instructions further include:

[0288] The computer interface outputs instructions to the operators of the distributed production environment for controlling and / or monitoring the one or more production operations.

[0289] And / or machine-readable instructions for automatically controlling equipment and / or sensors by a computer processor to remove the one or more sources of the one or more anomalies.

[0290] Clause 3:

[0291] A method for providing a computer implementation of a trained transformer-based model used according to clause 1 or 2, the method comprising:

[0292] One or more training datasets are provided via the computer interface;

[0293] The computer interface provides access to a pre-trained transformer-based model that includes at least transformer components; and

[0294] The one or more training datasets are used to prompt the pre-trained transformer-based model to be fine-tuned or retrained to generate the trained transformer-based model used in accordance with Clause 1 or 2.

[0295] Clause 4:

[0296] The computer-implemented method according to Clause 3, wherein the one or more training datasets comprise:

[0297] The first training dataset includes one or more factory technical documents, and

[0298] The second training dataset includes plant-based data associated with one or more production operations on one or more production lines of the distributed chemical production.

[0299] Clause 5:

[0300] The computer-implemented method according to Clause 4, wherein:

[0301] The second training dataset includes data points associated with one or more production operations and with standard operating conditions of the production line, wherein the standard operating conditions are defined by a predefined range of one or more parameters associated with the one or more production operations of the production line;

[0302] Clause 6:

[0303] The computer-implemented method according to Clause 4 or 5, wherein:

[0304] The second training dataset includes data points associated with one or more production operations and extreme operating conditions of the production line, wherein the extreme operating conditions are defined by one or more parameters exceeding the predefined range of one or more parameters associated with one or more production operations of the production line.

[0305] Clause 7:

[0306] The computer-implemented method according to any one of clauses 1 to 6 further comprises:

[0307] The trained transformer-based model receives one or more database query training datasets, including one or more database query languages.

[0308] When a user prompt includes instructions for retrieving one or more data according to any one of clauses 1 to 5, the trained transformer-based model is prompted to use the one or more databases to query the training dataset to retrieve the one or more data from the one or more databases, wherein the instructions include natural language.

[0309] Clause 8:

[0310] The computer-implemented method according to Clause 7 further includes:

[0311] The prompt indicates that the trained transformer-based model described in Clause 7 can retrieve one or more data points according to any one of Clauses 1 to 6.

[0312] The transformer-based model trained according to Clause 7 identifies whether access credentials and / or authorization credentials are required to access one or more databases including the one or more data.

[0313] If access credentials and / or authorization credentials are required, a request for authorization credentials and / or access credentials for accessing the one or more databases is generated by a transformer-based model trained in accordance with Clause 6.

[0314] The one or more data are retrieved by a transformer-based model trained according to Clause 7.

[0315] And optionally, use one or more data according to any one of Clauses 1 to 6.

[0316] Clause 9:

[0317] A computer-implemented method according to any one of Clauses 1 to 8, wherein one or more data according to any one of Clauses 1 to 6 are provided by uploading the one or more data via a user interface.

[0318] Clause 10:

[0319] A computer-implemented method for detecting one or more anomalies in a production cycle of distributed chemical production, the method comprising:

[0320] When an AI engine, which includes at least one computer processor for operating one or more transformer-based models, requests authorization, the user authorizes access to the AI ​​engine via a computer interface.

[0321] Select a transformer-based model from the list of one or more transformer-based models via a computer interface.

[0322] When the AI ​​engine requests authorization credentials and / or access credentials to access the selected transformer-based model, the user provides the authorization credentials and / or access credentials via the computer interface.

[0323] The selected transformer-based model is a pre-trained transformer-based model that includes a transformer component and embedded data based on unstructured data, including at least one or more text data.

[0324] Access the pre-trained transformer-based model via the computer interface.

[0325] The computer interface prompts the pre-trained transformer-based model to access a first training dataset, which includes one or more database queries in one or more database query languages.

[0326] The computer interface prompts the pre-trained transformer-based model to train using the first training dataset to generate a first trained transformer-based model, which is trained to interpret natural language instructions from the operators of the distributed chemical production as database queries.

[0327] The computer interface prompts the first trained transformer-based model to retrieve a second dataset associated with one or more production operations, wherein the prompt includes one or more parameters associated with one or more data characteristics representing the second dataset.

[0328] The second dataset is retrieved by the first trained transformer-based model.

[0329] The first trained transformer-based model is trained using the second dataset to generate a second trained transformer-based model, which includes an embedded data distribution model suitable for interpreting data associated with one or more production operations of the distributed chemical production.

[0330] The computer interface prompts the second trained transformer-based model to retrieve a third dataset associated with one or more production operations, wherein the prompt includes one or more parameters associated with one or more data characteristics characterizing the third training dataset, and wherein the prompt also includes instructions for analyzing the third dataset.

[0331] The third dataset is retrieved by the second trained transformer-based model.

[0332] The second trained transformer-based model analyzes the third dataset according to the instructions for analyzing the third dataset, wherein the analysis includes comparing the third dataset with the embedded data distribution model, and the analysis further includes identifying the deviations of data points in the third dataset from those in the embedded data distribution model.

[0333] A response is generated to the user by the second trained transformer-based model, wherein the response includes instructions in a natural language format, and / or the response includes a technical report characterizing the deviation of the data points of the third dataset from the embedded data distribution model, for removing the one or more sources of the one or more anomalies causing the deviation of the data points.

[0334] Clause 11:

[0335] The computer-implemented method according to Clause 10

[0336] The first trained transformer-based model and the second trained transformer-based model can be transformer-based models with an embedding layer that embeds data from the first training dataset and the second training dataset.

[0337] Alternatively, the first trained transformer-based model and the second trained transformer-based model can be two transformer-based models, each having an embedding layer that embeds the first training dataset or the second training dataset.

[0338] It also includes migrating one or more embedded data, including embedded data based on unstructured data, embedded data based on the first training dataset, and / or embedded data based on the second dataset, to another transformer-based model that does not yet include the one or more embedded data, wherein the other transformer-based model is selected from a list of the one or more transformer-based models, the first transformer-based model, or the second transformer-based model.

[0339] Clause 12:

[0340] Use another transformer-based model as described in Clause 11, instead of any of the pre-trained or trained transformer-based models specified in any of Clauses 1 to 11.

[0341] Clause 13:

[0342] The computer-implemented method according to Clause 12, wherein the one or more parameters associated with the one or more data characteristics representing the data of the second dataset and / or the third dataset are any of the following: time period, factory identifier, production line identifier, production cycle identifier, one or more operation parameters associated with one or more production operations, product identifier, factory, one or more parameters associated with one or more attributes of the product, data type, such as temperature, pressure, flow rate, quality.

[0343] Clause 14:

[0344] The computer-implemented method according to any one of Clauses 11 to 13,

[0345] It also includes providing the first training dataset, the second training dataset, the third dataset, and / or additional datasets associated with one or more production operations of the distributed facility via a user interface.

[0346] Clause 15:

[0347] A computer program product comprising computer-readable instructions that, when executed on a computer, cause the computer to perform steps specified in any one of clauses 1 to 14.

[0348] Clause 16:

[0349] A method for detecting one or more anomalies in a production cycle of distributed chemical production, the method comprising:

[0350] A transformer-based model is selected from a list of one or more transformer-based models via a computer interface. The selected transformer-based model is a pre-trained transformer-based model that includes a transformer component and embedded data based on unstructured data, including at least one or more text data.

[0351] Access the pre-trained transformer-based model via the computer interface.

[0352] The computer interface prompts the pre-trained transformer-based model to access a first training dataset, which includes one or more database queries in one or more database query languages.

[0353] The computer interface prompts the pre-trained transformer-based model to train using the first training dataset to generate a first trained transformer-based model, which is trained to interpret natural language instructions from the operators of the distributed chemical production as database queries.

[0354] The computer interface prompts the first trained transformer-based model to retrieve a second dataset associated with one or more production operations, wherein the prompt includes one or more parameters associated with one or more data characteristics representing the second dataset.

[0355] The second dataset is retrieved by the first trained transformer-based model.

[0356] The first trained transformer-based model is trained using the second dataset to generate a second trained transformer-based model, which includes an embedded data distribution model suitable for interpreting data associated with one or more production operations of the distributed chemical production.

[0357] The computer interface prompts the second trained transformer-based model to retrieve a third dataset associated with one or more production operations, wherein the prompt includes one or more parameters associated with one or more data characteristics characterizing the third training dataset, and wherein the prompt also includes instructions for analyzing the third dataset.

[0358] The third dataset is retrieved by the second trained transformer-based model.

[0359] The second trained transformer-based model analyzes the third dataset according to the instructions for analyzing the third dataset, wherein the analysis includes comparing the third dataset with the embedded data distribution model, and the analysis further includes identifying the deviations of data points in the third dataset from those in the embedded data distribution model.

[0360] A response is generated to the user by the second trained transformer-based model, wherein the response includes instructions in a natural language format, and / or the response includes a technical report characterizing the deviation of the data points of the third dataset from the embedded data distribution model, for removing the one or more sources of the one or more anomalies causing the deviation of the data points.

[0361] The prior art publication; No. 684; paragraphs

[1000] to

[8005] ; ISSN: 2198-4786; publication date: February 12, 2024, shall be considered as reference RF1, the full text of which is incorporated herein by reference. Preferably, the (chemical) product is the product described in paragraphs

[1000] to

[8005] of reference RF1. Preferably, the method / process described herein is further a method / process for producing the product.

[0362] The conversion steps for obtaining the product preferably include one or more steps as described below, and can be performed by conventional methods well known to those skilled in the art. The conversion steps preferably include one or more steps selected from the following:

[0363] Recovery, preferably depolymerization, gasification, pyrolysis and / or steam cracking; and / or purification, preferably crystallization, (solvent) extraction, distillation, evaporation, hydrogenation, absorption, adsorption and / or ion exchange using ion exchangers; and / or assembly, preferably foaming, synthesis, chemical conversion, chemical transformation, polymerization and / or compounding; and / or molding, preferably foaming, extrusion and / or molding; and / or finishing, preferably coating and / or smoothing.

[0364] Furthermore, one or more steps are described in detail in paragraphs

[1000] through

[8005] of reference RF1.

[0365] This disclosure is also described in conjunction with preferred embodiments and examples. However, those skilled in the art can understand and implement other variations of the claimed subject matter through a study of the accompanying drawings, this disclosure, and the claims. It is noteworthy that, in particular, any steps presented can be performed in any order; that is, this disclosure is not limited to a specific order of these steps. Furthermore, it is not necessary to perform different steps at a specific location in the distributed system or on a single node; that is, each step can be performed on different nodes using different equipment / data processing.

[0366] The order of all the method steps presented above is not mandatory, and alternative orders are possible. However, the specific order of the method steps shown as examples in the accompanying drawings should be considered as one possible order of the method steps, for example, for the corresponding embodiments described in the corresponding drawings or embodiments that include at least some of the steps described in the corresponding drawings.

[0367] In this specification, any connections presented in the described embodiments should be understood in a manner that allows the components involved to be operatively coupled. Therefore, connections can be direct or indirect, have any number or combination of intermediate elements, and may exist solely as functional relationships between components.

[0368] The indefinite article “a” or “an” should not be interpreted as “one”, that is, using the expression “one element” does not exclude the presence of other elements. A single element or other unit may perform the function of several entities or items recited in the claims. The fact that certain measures are recited only in mutually different dependent claims does not mean that combinations of these measures cannot be used in advantageous embodiments or that additional elements may be included.

[0369] The expressions “A and / or B” and “at least one of A or B” are considered interchangeable and are intended to include any one of the following three cases: (i) A, (ii) B, (iii) A and B. More generally, the expressions “at least one of the following: ” and “at least one of ” and similar wording (where the list of two or more elements is connected by “and” or “or”) mean at least one element, or at least any two or more elements, or at least all elements.

[0370] The provision within the scope of this disclosure may include any interface configured to provide data. This may include application programming interfaces, human-machine interfaces (such as displays), and / or software module interfaces. The provision may include communication of data or submission of data to the interface, particularly displaying data to a user or using data by a receiving entity.

[0371] Acquisition within the scope of this disclosure may include any interface configured to acquire or receive data. This may include application programming interfaces, human-machine interfaces (such as displays), and / or software module interfaces. Acquisition may include transferring or submitting data from the interface, particularly data used by the receiving entity. Any acquisition of data, data structures, datasets, etc., may include receiving data, data structures, datasets, etc., from a server that provides (e.g., hosts) a database containing data, data structures, datasets, etc.

[0372] Various units, circuits, entities, nodes, or other computing components may be described as being “configured to” perform one or more tasks. “Configured to” means that a structure “has” a “circuit” that performs one or more tasks during operation. Units, circuits, entities, nodes, or other computing components may be configured to perform tasks even when the unit / circuit / component is not operating. Units, circuits, entities, nodes, or other computing components forming a structure corresponding to “configured to” may include hardware circuitry and / or memory storing executable program instructions to perform the operation. For ease of description, units, circuits, entities, nodes, or other computing components may be described as performing one or more tasks. Such descriptions should be interpreted as including the phrase “configured to”. Any expression “configured to” is explicitly intended not to invoke the interpretation of 35 USC § 112(f).

[0373] Generally speaking, the methods, apparatus, systems, computer elements, nodes, or other computing components described herein may include memory, software components, and hardware components. Memory may include volatile memory (such as static or dynamic random access memory) and / or non-volatile memory (such as optical disc or magnetic disk storage devices, flash memory, programmable read-only memory, etc.). Hardware components may include any combination of combinational logic circuits, clock storage devices (such as flip-flops, registers, latches, etc.), finite state machines, memory (such as static random access memory or embedded dynamic random access memory), custom-designed circuits, programmable logic arrays, etc.

[0374] In this specification, any connections presented in the described embodiments should be understood in a manner that allows the components involved to be operatively coupled. Therefore, connections can be direct or indirect, have any number or combination of intermediate elements, and may exist solely as functional relationships between components.

[0375] Furthermore, any methods, processes, and actions described or illustrated herein may be implemented using executable instructions in a general-purpose or special-purpose processor and stored on a computer-readable storage medium (e.g., a disk, memory, etc.) for execution by such a processor. The reference to "computer-readable storage medium" should be understood to encompass special-purpose circuitry, such as signal processing apparatus and other means.

[0376] Any disclosure and embodiments described herein relate to the methods, systems, apparatuses, and computer program elements listed above, and vice versa. Advantageously, the benefits provided by any embodiments and examples also apply to all other embodiments and examples, and vice versa.

[0377] Unless otherwise specified, all terms and definitions used herein should be understood broadly and have their general meanings.

[0378] It should be understood that all presented embodiments are merely examples, and any feature presented for a particular example embodiment may be used alone with any aspect, or in combination with any feature presented for the same or another particular example embodiment, and / or in combination with any other feature not mentioned. In particular, the example embodiments presented in this specification should also be understood as being disclosed in all possible combinations between each other, provided that it is technically reasonable and the example embodiments are not alternatives to each other. It should also be understood that any feature presented for example embodiments in a particular category (method / apparatus / computer program / system) may also be used in a corresponding manner in example embodiments of any other category. It should also be understood that the presence of a feature in a presented example embodiment does not necessarily mean that the feature is essential and cannot be omitted or replaced.

[0379] Reference List

[0380] Production line 1001;

[0381] Equipment 1002;

[0382] 1003 sensor;

[0383] 1004 Analysis Engine;

[0384] 1005 Production data / factory data, such as sensor data and analytical data;

[0385] 1006 Control and / or monitoring engine;

[0386] 1007 and 1008 are machine-readable instructions;

[0387] 1010 production / factory historical data;

[0388] 1011 Input data;

[0389] 1012 Operator Data;

[0390] 1013 training data.

Claims

1. A method for mitigating anomalies in the production cycle of distributed chemical production, the method comprising: - Receive analysis instructions related to detecting one or more anomalies in factory-based data, wherein the analysis instructions include a reference indicator indicating at least one reference dataset and an analysis indicator indicating at least one dataset to be analyzed; - Based on the analysis instructions, obtain the at least one reference dataset and the at least one dataset to be analyzed; - Determining at least one anomaly mitigation instruction, and determining at least one anomaly detection result includes: providing a task instruction to at least one generative data-driven model based on the analysis instruction, wherein the at least one generative data-driven model has been trained to generate the at least one anomaly mitigation instruction based on the at least one reference dataset and the at least one dataset to be analyzed in response to receiving the task instruction; - Provide at least one of the above-mentioned exception mitigation instructions.

2. The method of claim 1, wherein the task instruction includes at least a portion of the at least one reference dataset and / or at least a portion of the at least one dataset to be analyzed.

3. The method of claim 1 or 2, wherein the at least one reference dataset is used to fine-tune or retrain the generative data-driven model, and wherein the fine-tuned or retrained generative data-driven model is used as the generative data-driven model when determining the at least one anomaly mitigation instruction.

4. The method according to any one of claims 1 to 3, wherein obtaining the at least one reference dataset and / or the at least one dataset to be analyzed comprises: Determine whether the at least one reference dataset and / or the at least one dataset to be analyzed is available in the database; When it is determined that the at least one reference dataset and / or the at least one dataset to be analyzed is available in the database: retrieve the at least one reference dataset and / or the at least one dataset to be analyzed from the database.

5. The method according to claim 4, further comprising: When it is determined that at least one reference dataset is not available in the database: The user is requested to provide at least one reference dataset; Receive at least one reference dataset via a computer interface.

6. The method according to claim 4 or 5, further comprising: When it is determined that at least one dataset to be analyzed is not available in the database: The user is requested to provide at least one dataset to be analyzed; Receive at least one dataset to be analyzed via a computer interface.

7. The method according to any one of claims 1 to 6, further comprising: The production cycle of the distributed chemical production is operated and / or controlled based on the at least one anomaly mitigation instruction, particularly for the production of chemical products.

8. The method according to any one of claims 1 to 7, wherein the task instruction includes instructions for generating a technical report, the technical report including a description indicating the presence of the one or more anomalies or indicating the absence of anomalies, wherein, In the event that one or more anomalies are detected, the technical report may optionally include an indication of one or more sources of the anomalies.

9. The method according to any one of claims 1 to 8, wherein the generative data-driven model is fine-tuned or retrained using a technical training dataset comprising one or more factory technical documents, and wherein the fine-tuned or retrained generative data-driven model is used as the generative data-driven model when determining the at least one anomaly mitigation instruction.

10. The method of claim 9, wherein the technical training dataset comprises data points associated with one or more production operations and with standard operating conditions of the production line, wherein the standard operating conditions are defined by a predefined range of one or more parameters associated with the one or more production operations of the production line, wherein optionally, the one or more parameters are any of the following: time period, factory identifier, production line identifier, production cycle identifier, one or more operating parameters associated with one or more production operations, product identifier, factory, one or more parameters associated with one or more attributes of the product, data type, such as temperature, pressure, flow rate, quality.

11. The method according to any one of claims 1 to 10, further comprising: Determine whether access credentials and / or authorization credentials are required for accessing one or more databases including the at least one reference dataset, the at least one dataset to be analyzed, and / or the technical dataset according to claim 9 or 10. When it is determined that access credentials and / or authorization credentials are required, a request is generated to provide the user with authorization credentials and / or access credentials for accessing the one or more databases. Retrieve the at least one reference dataset, the at least one dataset to be analyzed, and / or the technical dataset according to claim 9 or 10, and optionally, use the at least one reference dataset and / or the technical dataset according to claim 9 or 10 to fine-tune or retrain the generative data-driven model; Furthermore, the fine-tuned or retrained generative data-driven model is used as the generative data-driven model when determining the at least one anomaly mitigation instruction.

12. The method according to any one of claims 1 to 11, further comprising: Access to the at least one generative data-driven model is provided via the computer interface, wherein it is determined that the at least one anomaly mitigation instruction is performed using the accessed at least one generative data-driven model.

13. The method according to any one of claims 1 to 12, wherein obtaining the at least one reference dataset, the at least one dataset to be analyzed, and / or the technical dataset according to claim 9 or 10 comprises: The at least one generative data-driven model generates a retrieval query for obtaining the at least one reference dataset, the at least one dataset to be analyzed, and / or the technical dataset; The generated search query is provided to a database including the at least one reference dataset, the at least one dataset to be analyzed, and / or the technical dataset, to obtain the at least one reference dataset, the at least one dataset to be analyzed, and / or the technical dataset; Based on the provided search query, obtain the at least one reference dataset, the at least one dataset to be analyzed, and / or the technical dataset from the database.

14. An apparatus comprising corresponding components for performing or carrying out the steps of any one of claims 1 to 13, or comprising at least one processor and at least one memory storing instructions, the instructions, when executed by the at least one processor, causing the apparatus to perform at least the steps of the method according to any one of claims 1 to 13.

15. Use of the method according to any one of claims 1 to 13 or the anomaly mitigation instruction generated by the device according to claim 14 for displaying the anomaly mitigation instruction to an operator of a chemical plant and / or for the production of chemical products.

Citation Information

Patent Citations

  • Determining operating conditions in chemical production plants

    WO2020165045A1

  • Manufacturing system for monitoring and / or controlling one or more chemical plant(s)

    WO2021116123A1

  • Industrial plant monitoring

    WO2021156157A1