Systems and methods for base model functionality of time series data of operational processes with industrial context information for industrial applications

By transforming the time-series data of industrial plant operations into time-unit sequences and adding contextual information, and training a generative AI model, the problems of low efficiency and poor reliability of ML solutions in new factory applications are solved, achieving rapid adaptation and accurate prediction.

CN122452809APending Publication Date: 2026-07-24ABB (SCHWEIZ) AG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ABB (SCHWEIZ) AG
Filing Date
2026-01-21
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies for applying machine learning (ML) solutions to industrial plants, especially new plants, suffer from low efficiency, poor reliability, and long processing times.

Method used

By transforming the time-series data of industrial plant operations into time unit sequences and adding industrial context information, a context lexicalization method is used to generate context lexical sequences. Generative AI models are then trained to fine-tune and predict operational processes for different plants.

Benefits of technology

It enables efficient, reliable, and time-saving generative AI model fine-tuning and forecasting, can quickly adapt to new factories, provides zero-sample prediction capabilities, and improves the accuracy and reliability of predictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122452809A_ABST
    Figure CN122452809A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to systems and methods for a base model functionality for time series data of operational processes with industrial context information for industrial applications. The method comprises: transforming time series data of an operational process associated with at least a first industrial plant into a sequence of time cells associated with respective segments of the operational process; appending time cells in the sequence of time cells with industrial context information; obtaining a sequence of context tokens from the appending by means of a tokenization method, wherein one context token represents one time cell; training a generative AI model on the sequence of context tokens using a loss function; and using the trained generative AI model with the base model functionality for at least one of: i) fine-tuning the trained generative AI model for a second industrial plant different from the first industrial plant; ii) making a forecast for a further progress of an operational process of the second industrial plant.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for obtaining the functionality of a basic model for an operational process with industrial context information for industrial applications. The invention also relates to data processing apparatus, data processing systems, computer-readable media, computer program products, and trained generative AI models with this basic model functionality used in industrial environments. Background Technology

[0002] For example, when a new industrial plant is built and put into operation, two requirements need to be met to provide reasonable analytics and machine learning (ML)-based solutions (e.g., operation monitoring, anomaly detection, maintenance due date prediction, etc.) for this new industrial plant. First, long-term data logging of the new industrial plant's operation is required to train ML based on this data. Second, this data-driven training actually takes a long time before deploying and effectively operating the actual ML-based solutions.

[0003] Therefore, in the context of industrial applications, there is a question: how to make ML-based solutions more efficient, more reliable, and less time-consuming in industrial plants (especially new industrial plants)?

[0004] Therefore, there is still room for improvement and a need for further development regarding the application of ML-based solutions in industrial plants (especially new ones). Summary of the Invention

[0005] In view of the above, the purpose of this disclosure is to at least overcome some of the shortcomings of applying ML-based solutions to industrial plants (especially new industrial plants).

[0006] Therefore, to address one or more of these deficiencies, a computer-implemented method is provided in a first aspect for obtaining the functionality of a basic model for time-series data of an operational process with industrial context information for an industrial sector. The method includes transforming the time-series data of an operational process associated with at least a first industrial plant into a sequence of time units associated with corresponding segments of the operational process. The method further includes appending industrial context information to the time units in the time unit sequence, wherein segments of the industrial context information are associated with segments of the operational process, and wherein concurrent segments of the industrial context information are appended to the time units. The method also includes obtaining a sequence of context lexical units from the appending using a lexicalization method, where each context lexical unit represents a time unit appended with concurrent segments of industrial context information. The method further includes training a generative artificial intelligence (AI) model on the context lexical sequence using a loss function. The method also includes using the trained generative AI model with basic model functionality for:

[0007] i) Fine-tune the trained generative AI model for a second industrial plant that differs from the first industrial plant, and / or

[0008] ii) To forecast further progress in the operation of the second industrial plant.

[0009] It's important to note that the term "base model functionality" refers to the functionality of a ML-based model. More specifically, this means that the ML-based model possesses extensive knowledge and understanding of the subject of interest and is generally ready to use out of the box without any additional work (e.g., further dedicated, resource-intensive data-driven training, calibration, tuning, or optimization), exhibiting good applicability and usability. In other words, a ML-based model with base model functionality can, for example, possess extensive knowledge and understanding of the operational processes of starting and / or optimizing one or more specific industrial plants (i.e., the subject of interest). For example, these one or more industrial plants can be understood as representing one or more industrial plants used in a specific industrial sector (such as in the mining industry or the chemical industry). Therefore, if the industrial plant is newly incorporated or embedded in such a specific industrial sector, for example, then a ML-based model with base model functionality is already quite suitable—that is, ready to use out of the box and well applied to the newly incorporated, embedded, or installed industrial plant.

[0010] Therefore, the term "at least one industrial plant" includes the transformation of time-series data on the operation processes of one or more industrial plants. These one or more industrial plants may be associated with the same or different industrial sectors. For example, all of these one or more industrial plants may be used in the mining industry, or a first set of these one or more industrial plants may be used in the mining industry, while a second set of these one or more industrial plants may be used in the chemical industry. Furthermore, these one or more industrial plants may be associated with the same or different operation processes. For example, for each of these one or more industrial plants, time-series data on the same operation process or the same type of operation process may have been recorded (e.g., an operation process for bath production in the chemical industry). Alternatively, for each of the first sets of these one or more industrial plants, time-series data on a first operation process or a first type of operation process may have been recorded (e.g., an operation process for bath production in the chemical industry), and for each of the second sets of these one or more industrial plants, time-series data on a second operation process or a second type of operation process may have been recorded (e.g., an operation process for extracting raw materials of interest from ore).

[0011] Furthermore, it should be understood that these one or more industrial plants do not include a second industrial plant. However, a second industrial plant may be associated with an industrial sector (e.g., associated with a target industrial sector), and the second industrial plant may perform the target operating process or a target type of operating process. A second industrial plant may be associated with one or more industrial plants. For example, at least one of the one or more industrial plants may be associated with the same industrial sector as the target industrial sector. Additionally or alternatively, for example, at least one operating process or at least one type of operating process performed by one or more industrial plants may be the same as the target operating process or target type of operating process performed by the second industrial plant.

[0012] The term "time-series data of operational processes with industrial context information" refers to the time-series observation or recording of one or more components, quantities, and / or events in an industrial plant, along with their context and / or environmental conditions. For example, these components may include distillation columns, level controllers, reactors, or storage tanks. Quantities may include physical quantities, tank levels, sensor values, or key performance indicators (KPIs). For example, events may include logs or alarms. For example, context and / or environmental conditions may include plant topology, specific industrial context, seasonality, weather, date, temperature, or video recordings.

[0013] The phrase "time units associated with corresponding segments of the operation process" clarifies that each segment of the operation process is associated with a specific time period or point in time. For example, for ease of explanation, we can assume that the operation process comprises a sequence of three segments: the first segment, the second segment, and the third segment. Each of these segments can begin at a certain point in time and last for a specific time period. For instance, the first segment can begin at a first time and last for a first time period; the second segment can begin at a second time and last for a second time period; and the third segment can begin at a third time and last for a third time period. Therefore, the first, second, and third segments can all be associated with specific points in time and / or time periods. Thus, by observing specific points in time and / or time periods, the corresponding segments of the operation process can be identified.

[0014] The statement "attaching industrial context information to time units in a time unit sequence, wherein segments of industrial context information are associated with segments of the operation process, and time units are attached with concurrent segments of industrial context information" implies, in other words, that time units are linked to corresponding (i.e., concurrent) industrial context information. For example, it can be assumed that a time unit represents a target time period, and that this time unit is attached to or linked to industrial context information that is available, exists, or has occurred within that target time period. For example, for ease of explanation, it can be assumed that the industrial context information includes segments of a first industrial context information, a second industrial context information, and a third industrial context information. The first industrial context information segment may include information about the industrial plant tank during operation. The second industrial context information segment may include temperature information about the environment surrounding the industrial plant tank during operation. The third industrial context information segment may include pressure at the valves of the industrial plant tank during shutdown. Therefore, the first and second industrial context information segments may refer to the same time period. That is, during operation (e.g., during tank startup), the state of the tank is monitored, and the temperature around the tank is measured. Therefore, the start-up time unit corresponding to tank startup can be appended with fragments of first and second industrial context information. In this example, the fragments of first and second industrial context information can be understood as concurrent fragments of industrial context information. Furthermore, if the operation process is not tank startup but a shutdown process (i.e., tank shutdown), then the shutdown time unit corresponding to tank shutdown can be appended with fragments of first, second, and third industrial context information. In this example, the fragments of first, second, and third industrial context information can be understood as concurrent fragments of industrial context information.

[0015] In the context of generative AI, tokenization is a common natural language processing technique that breaks down paragraphs and sentences into smaller units that are more easily assigned semantic meaning. A similar operation can be performed on temporal sequences: a sequence of values ​​(time-stamped, like words in a sentence) can be decomposed into segments with specific patterns, such as monotonically decreasing, increasing, rising followed by rapid decreasing, parts of a sine / cosine curve, or even something resembling a heartbeat. These typical patterns (like words or parts of words) can be considered tokens. Furthermore, when evaluating temporal sequences in context, the different parts possess additional characteristics. For example, a monotonically decreasing temporal segment might correspond to different conditions, such as the section of the plant in which it was recorded (e.g., decreasing temperature in a distillation column versus decreasing pressure in a sealed storage tank), or the weather conditions at the time of recording (compared to other conditions). These unique combinations are then assigned unique tokens. Therefore, in the context of generative AI, the term "lexicalization" describes the process of breaking down large amounts of sequential data (e.g., documents, text, or sequences of timestamped values) into smaller units with additional contextual information. These smaller units are called lemmas. Depending on the type of lexicalization, lemmas can be defined by specific character sequences, punctuation marks, or other definitions. Doing so ultimately makes it easier for machines to process documents, body text, time series, or time series with context (where the context can be textual content or values).

[0016] It is important to note that the aforementioned start-up time unit (the fragment appended with the first industrial context information and the second industrial context information) can be understood as representing the first context term or the start-up context term. The close-up time unit (the fragment appended with the first industrial context information, the second industrial context information, and the third industrial context information) can be understood as representing the second context term or the close-up context term. Generally, a context term can be a term that captures the time unit (e.g., a timing period or timing point) along with industrial context information.

[0017] It is important to note that the statement "training a generative AI model on a sequence of context lexical units" means, in other words, using the sequence of context lexical units (i.e., the data that constitutes the sequence of context lexical units) as input data to train the generative AI model.

[0018] Furthermore, it's important to note that the trained generative AI model with basic model functionality can be used in three ways. First, the trained generative AI model can be fine-tuned for the second industrial plant. Second, the trained generative AI model can be used to predict further progress in the operation of the second industrial plant without fine-tuning. Third, the fine-tuned trained generative AI model can be used to predict the operation of the second industrial plant.

[0019] The approach based on the first aspect has advantages because it can participate in the fine-tuning of generative AI models for specific industrial applications in an efficient, reliable, and time-saving manner. Furthermore, it also contributes to the implementation of zero-shot hints for prediction or forecasting.

[0020] More specifically, trained generative AI models can be understood as foundational models of time series data with context in industrial applications. A key advantage of these foundational models is that they are base models (i.e., trained on large time series datasets), allowing for zero-shot predictions or forecasts. Therefore, foundational models require only a minimal amount of input data (i.e., a sequence of time series data with "context") to enable rapid startup, for example, for a newly installed industrial plant, and to predict the future of such a plant using well-performing and readily available multimodal models (e.g., time series and topology-based). By sampling multiple future trajectories, a probability forecast distribution and uncertainty interval can be obtained, thereby improving confidentiality. Based on a language model architecture, trained generative AI models promise to exhibit next-word prediction capabilities as impressive as language models themselves. Thanks to the integration of industrial context information disclosed in this paper, existing techniques for pure time-series forecasting are extended, and trained generative AI models can hopefully provide a solid foundation for delivering fairly good and rapid outputs, especially given the significant resources required for training (the cost of time, computation, and expertise needed to train specialized models, and the cost of data). As described, once a trained generative AI model is available, it can be used as a model to generate effective forecasts using only a very small amount of input data.

[0021] Furthermore, this benefits operators. Regarding new topologies and “new industrial plants,” for example, from only one week's worth of recorded data, leveraging “base model knowledge” and zero-shot capability, a reasonably reasonable prediction of the operation of a “new industrial plant” can be obtained after fine-tuning a trained generative AI model (i.e., after fine-tuning a specific base model using this week's worth of recorded data), whereas no specific model could be trained earlier or better using only one week's worth of recorded data. Therefore, trained generative AI models are particularly valuable in scenarios where a rapid acquisition of time-series prediction models is required; for any use case, whether it's pattern analysis, anomaly detection, or operational pattern classification. Regarding topology modifications and “modified industrial plants,” the reliability of a trained generative AI model may decrease compared to a model trained only on a specific “modified industrial plant,” but the accuracy of a trained generative AI model will still be higher than a model trained using only one week's worth of data.

[0022] Furthermore, this also benefits dashboard users. For dashboard users who work with algorithm-based dashboards that observe or monitor processes in operations based on time-series data, and if the dashboard user has an AI / ML model or algorithm to analyze this time-series data, the dashboard user will be able to set up dashboards more quickly using trained generative AI models.

[0023] In addition, trained generative AI models are valuable when one or more operational processes in an industrial plant require urgent prediction, and / or when the industrial plant does not have a dedicated model available or cannot afford one. Reliability can be improved by sampling several inference runs and thus obtaining a probability forecast distribution.

[0024] According to several examples of this disclosure, the transformation may include: transforming time series data into a time unit sequence by mean scaling and / or quantizing the time series data to obtain a time unit sequence.

[0025] Therefore, transformations can be performed in an efficient manner.

[0026] According to several examples of this disclosure, the appending may include appending a context vector to each time unit in the time unit, the context vector being associated with a segment of industrial context information, which is associated with a corresponding time unit.

[0027] Therefore, time units can be efficiently attached with corresponding industrial context information, and the connection or relationship between time units and corresponding industrial context information can be traced by using context vectors.

[0028] According to several examples of this disclosure, at least a first industrial plant may include multiple plant components. Timing data may include plant component timing data of plant component operation processes associated with the multiple plant components. The plant component operation process may consist of segments of the operation process, and each segment of the plant component timing data may correspond to one or more corresponding plant components among the multiple plant components. Segments of industrial context information may include plant component industrial context information associated with the plant component operation process.

[0029] It is important to note that, in other words, considering multiple factory components aims to decompose the timing data of the operational process into individual factory components. The operational process of an industrial plant can be understood as the operational process of factory components, encompassing the industrial plant's individual factory components. Therefore, segments of the industrial plant's operational process can include factory component operational processes. Each factory component can be associated with corresponding factory component timing data; that is, corresponding factory component timing data can be recorded for each factory component. For each factory component, corresponding factory component industrial context information can be identified. Therefore, context terms can represent time units, which can be appended with concurrent segments of factory component industrial context information. Concurrent segments of factory component industrial context information appended to a time unit can correspond to the same factory component or different factory components.

[0030] Therefore, trained generative AI models are also applicable at the factory component level, that is, applicable to individual factory components.

[0031] According to several examples of this disclosure, fragments of industrial context information may also indicate at least one of the following:

[0032] - Environmental data associated with the operational processes of at least the first industrial plant;

[0033] - Environmental data associated with the operation of multiple factory components;

[0034] - An industrial domain in which at least one industrial plant is operated;

[0035] - Alarm and event data associated with at least the first industrial plant;

[0036] - Multimodal data regarding multiple factory components; and

[0037] - A topological context in which at least the first industrial plant is embedded.

[0038] It is important to note that environmental data can include ambient temperature, humidity, and weather conditions. For example, the temperature, humidity, and weather conditions of the area within which an industrial plant operates. Or, the temperature and humidity of the plant itself. An industrial domain can be a specific industrial sector, such as the mining or chemical industry. Multimodal data regarding multiple plant components can mean, for example,: i) textual or numerical characteristics, such as component size, type of processed material, operating temperature, etc.; ii) image sequences; iii) video; and / or iv) sound. Topological context can indicate which components are connected to at least one industrial plant, and how / where they are connected.

[0039] Therefore, because industrial context information covers such a wide range of different types of information, a huge database can be generated for training generative AI models, and because industrial context information covers so many different types of information in detail, trained generative AI models can be obtained and even further improved.

[0040] According to several examples of this disclosure, the training may include performing the training in a self-supervised manner by using at least one of the following tasks: multivariate prediction for context lexical sequences, and introducing perturbations into context lexical sequences.

[0041] Therefore, generative AI models can be trained in an efficient manner.

[0042] According to several examples of this disclosure, fine-tuning may include: - Fine-tuning the trained generative AI model using fine-tuning training data, which is associated with the operation process of the second industrial plant; - Forecasts may include forecasting using a fine-tuned, trained generative AI model.

[0043] Therefore, the application of the trained generative AI model in the second industrial plant has been further improved.

[0044] According to several examples of this disclosure, fine-tuning may include fine-tuning the trained generative AI model for at least one of the following: - Applications that are to be operated in specific types of industries in secondary industrial plants; - Applications that execute specific processes in a secondary industrial plant; and – Applications that need to be operated in specific climate zones at a second industrial plant.

[0045] For example, a specific climate zone may include ambient temperature and ambient humidity within a specific range.

[0046] Therefore, the application of the trained generative AI model in the second industrial plant has been further improved.

[0047] According to several examples of this disclosure, fine-tuning training data may include temporal fine-tuning data and fine-tuning context information associated with the temporal fine-tuning data, wherein the temporal fine-tuning data and fine-tuning context information may be associated with the operation of a second industrial plant and / or with the operation of plant components of the second industrial plant. According to several examples of this disclosure, the temporal fine-tuning data may be specific to an industrial context domain associated with the second industrial plant. According to several examples of this disclosure, the context information may be specific to that industrial context domain.

[0048] Therefore, the application of trained generative AI models in the second industrial plant has been further enhanced.

[0049] According to several examples of this disclosure, time-series data and / or time-series fine-tuning data may include at least one of three types of training data: historical data from actual operational processes, synthetic data from synthetic operational processes, and synthetic data from historical data. Synthetic data from historical data may be synthesized via data manipulation (e.g., by adding noise).

[0050] Therefore, the application of trained generative AI models in the second industrial plant has been further enhanced.

[0051] According to several examples of this disclosure, the forecast may include obtaining a probabilistic forecast of the operational behavior of a second industrial plant by sampling future trajectories of multiple forecasts.

[0052] Therefore, the reliability of the forecast is improved.

[0053] According to several examples in this disclosure, the loss function may be a cross-entropy loss function. Generally, an appropriate loss function should be chosen for classification problems (e.g., cross-entropy) or regression problems (e.g., root mean square error). The time unit may include a point in time and / or a time period. The industry context information may be sparse industry context information, and a sequence of contextual terms can be obtained by appending the sparse industry context information to the time unit.

[0054] Therefore, since the loss function is the cross-entropy loss function, the quality of the training results is improved. Furthermore, because the time unit includes both time periods and time points, time-series data can be divided into individual time units more finely and specifically.

[0055] According to a second aspect, a data processing apparatus is provided. The data processing apparatus includes one or more processors configured to perform the method according to the first aspect.

[0056] The advantage of the data processing device described in the second aspect is that it can participate in the fine-tuning of generative AI models for specific industrial plants in an efficient, reliable, and time-saving manner. Furthermore, it can also participate in zero-shot hint prediction or forecasting.

[0057] More specifically, trained generative AI models can be understood as foundational models of time series data with context in industrial applications. A key advantage of these foundational models is that they are base models (i.e., trained on large time series datasets), allowing for zero-shot predictions or forecasts. Therefore, foundational models require only a minimal amount of input data (i.e., a sequence of time series data with "context") to enable rapid startup, for example, for a newly installed industrial plant, and to predict the future of such a plant using well-performing and readily available multimodal models (e.g., time series and topology-based). By sampling multiple future trajectories, a probability forecast distribution and uncertainty interval can be obtained, thereby improving confidentiality. Based on a language model architecture, trained generative AI models promise to exhibit next-word prediction capabilities as impressive as language models themselves. Thanks to the integration of industrial context information disclosed in this paper, existing techniques for pure time-series forecasting are extended, and trained generative AI models can hopefully provide a solid foundation for delivering fairly good and rapid outputs, especially given the significant resources required for training (the cost of time, computation, and expertise needed to train specialized models, and the cost of data). As described, once a trained generative AI model is available, it can be used as a model to generate effective forecasts using only a very small amount of input data.

[0058] Furthermore, it benefits operators. Regarding new topologies and "new industrial plants," for example, using only one week's worth of recorded data, leveraging "base model knowledge" and zero-shot capability, a reasonably reasonable prediction of the operation of a "new industrial plant" can be obtained after fine-tuning a trained generative AI model (i.e., after fine-tuning a specific base model using this week's worth of recorded data), whereas no specific model could be trained earlier or better using only one week's worth of recorded data. Therefore, trained generative AI models are particularly valuable in scenarios requiring rapid time-series prediction models; for any use case, whether it's pattern analysis, anomaly detection, or operational pattern classification. Regarding topology modifications and "modified industrial plants," the reliability of a trained generative AI model may decrease compared to a model trained only on a specific "modified industrial plant," but its accuracy will still be higher than a model trained using only one week's worth of data.

[0059] Furthermore, this also benefits dashboard users. For dashboard users who work with algorithm-based dashboards that observe or monitor processes in operations based on time-series data, and if the dashboard user has an AI / ML model or algorithm to analyze this time-series data, the dashboard user will be able to set up dashboards more quickly using trained generative AI models.

[0060] In addition, trained generative AI models are valuable when one or more operational processes in an industrial plant require urgent prediction, and / or when the industrial plant does not have a dedicated model available or cannot afford one. Reliability can be improved by sampling several inference runs and thus obtaining a probability forecast distribution.

[0061] According to a third aspect, a data processing system is provided. The data processing system includes the data processing apparatus according to the second aspect. Alternatively or concurrently, the data processing system further includes components for performing the method according to the first aspect.

[0062] The third aspect of the data processing system has the following advantages: it can participate in the fine-tuning of generative AI models for specific industrial plants in an efficient, reliable, and time-saving manner. Furthermore, it can also participate in zero-shot hint prediction or forecasting.

[0063] More specifically, trained generative AI models can be understood as foundational models of time series data with context in industrial applications. A key advantage of these foundational models is that they are base models (i.e., trained on large time series datasets), allowing for zero-shot predictions or forecasts. Therefore, foundational models require only a minimal amount of input data (i.e., a sequence of time series data with "context") to enable rapid startup, for example, for a newly installed industrial plant, and to predict the future of such a plant using well-performing and readily available multimodal models (e.g., time series and topology-based). By sampling multiple future trajectories, a probability forecast distribution and uncertainty interval can be obtained, thereby improving confidentiality. Based on a language model architecture, trained generative AI models promise to exhibit next-word prediction capabilities as impressive as language models themselves. Thanks to the integration of industrial context information disclosed in this paper, existing techniques for pure time-series forecasting are extended, and trained generative AI models promise to provide a solid foundation for delivering fairly good and rapid outputs, especially given the significant resources required for training (the cost of time, computation, and expertise needed to train specialized models, and the data itself). As described, once a trained generative AI model is available, it can be used as a model to generate effective forecasts using only a very small amount of input data.

[0064] Furthermore, it benefits operators. Regarding new topologies and "new industrial plants," for example, using only one week's worth of recorded data, leveraging "base model knowledge" and zero-shot capability, a reasonably reasonable prediction of the operation of a "new industrial plant" can be obtained after fine-tuning a trained generative AI model (i.e., after fine-tuning a specific base model using this week's worth of recorded data), whereas no specific model could be trained earlier or better using only one week's worth of recorded data. Therefore, trained generative AI models are particularly valuable in scenarios requiring rapid time-series prediction models; for any use case, whether it's pattern analysis, anomaly detection, or operational pattern classification. Regarding topology modifications and "modified industrial plants," the reliability of a trained generative AI model may decrease compared to a model trained only on a specific "modified industrial plant," but its accuracy will still be higher than a model trained using only one week's worth of data.

[0065] Furthermore, this also benefits dashboard users. For dashboard users who work with algorithm-based dashboards that observe or monitor processes in operations based on time-series data, and if the dashboard user has an AI / ML model or algorithm to analyze this time-series data, the dashboard user will be able to set up dashboards more quickly using trained generative AI models.

[0066] In addition, trained generative AI models are valuable when one or more operational processes in an industrial plant require urgent prediction, and / or when the industrial plant does not have a dedicated model available or cannot afford one. Reliability can be improved by sampling several inference runs and thus obtaining a probability forecast distribution.

[0067] According to a fourth aspect, an industrial plant is provided, the industrial plant including the data processing apparatus according to the second aspect and / or the data processing system according to the third aspect.

[0068] According to several examples, "industrial plant" means an industrial factory, autonomous industrial plant, or industrial production plant, for example, including one or more pipelines, production lines, and / or assembly lines for transforming one or more raw materials into products and / or assembling one or more components into a final product. According to several examples, it can mean an industrial plant in the oil industry, natural gas industry, mining industry, chemical industry, wind and electricity industry, or food and beverage industry.

[0069] According to the fourth aspect, the advantage of this industrial plant lies in its ability to participate in the fine-tuning of generative AI models for specific industrial plants in an efficient, reliable, and time-saving manner. Furthermore, it can also participate in zero-shot hint prediction or forecasting.

[0070] More specifically, trained generative AI models can be understood as foundational models of time series data with context in industrial applications. A key advantage of these foundational models is that they are base models (i.e., trained on large time series datasets), allowing for zero-shot predictions or forecasts. Therefore, foundational models require only a minimal amount of input data (i.e., a sequence of time series data with "context") to enable rapid startup, for example, for a newly installed industrial plant, and to predict the future of such a plant using well-performing and readily available multimodal models (e.g., time series and topology-based). By sampling multiple future trajectories, a probability forecast distribution and uncertainty interval can be obtained, thereby improving confidentiality. Based on a language model architecture, trained generative AI models promise to exhibit next-word prediction capabilities as impressive as language models themselves. Thanks to the integration of industrial context information disclosed in this paper, existing techniques for pure time-series forecasting are extended, and trained generative AI models can hopefully provide a solid foundation for delivering fairly good and rapid outputs, especially given the significant resources required for training (the cost of time, computation, and expertise needed to train specialized models, and the cost of data). As described, once a trained generative AI model is available, it can be used as a model to generate effective forecasts using only a very small amount of input data.

[0071] Furthermore, it benefits operators. Regarding new topologies and "new industrial plants," for example, using only one week's worth of recorded data, leveraging "base model knowledge" and zero-shot capability, a reasonably reasonable prediction of the operation of a "new industrial plant" can be obtained after fine-tuning a trained generative AI model (i.e., after fine-tuning a specific base model using this week's worth of recorded data), whereas no specific model could be trained earlier or better using only one week's worth of recorded data. Therefore, trained generative AI models are particularly valuable in scenarios requiring rapid time-series prediction models; for any use case, whether it's pattern analysis, anomaly detection, or operational pattern classification. Regarding topology modifications and "modified industrial plants," the reliability of a trained generative AI model may decrease compared to a model trained only on a specific "modified industrial plant," but its accuracy will still be higher than a model trained using only one week's worth of data.

[0072] Furthermore, this also benefits dashboard users. For dashboard users who work with algorithm-based dashboards that observe or monitor processes in operations based on time-series data, and if the dashboard user has an AI / ML model or algorithm to analyze this time-series data, the dashboard user will be able to set up dashboards more quickly using trained generative AI models.

[0073] In addition, trained generative AI models are very valuable when one or more operational processes in an industrial plant require urgent prediction, and / or when the industrial plant does not have a dedicated model available or cannot afford one. Reliability can be improved by sampling several inference runs and thus obtaining a probability forecast distribution.

[0074] According to a fifth aspect, a computer-readable medium including instructions that, when executed by a computing system, cause the computing system to perform the method according to a first aspect. The computer-readable medium may be transient or non-transient, volatile or non-volatile.

[0075] According to the fifth aspect, the advantage of this computer-readable medium is that it can participate in the fine-tuning of generative AI models for specific industrial plants in an efficient, reliable, and time-saving manner. Furthermore, it can also participate in zero-shot hint prediction or forecasting.

[0076] More specifically, trained generative AI models can be understood as foundational models of time series data with context in industrial applications. A key advantage of these foundational models is that they are base models (i.e., trained on large time series datasets), allowing for zero-shot predictions or forecasts. Therefore, foundational models require only a minimal amount of input data (i.e., a sequence of time series data with "context") to enable rapid startup, for example, for a newly installed industrial plant, and to predict the future of such a plant using well-performing and readily available multimodal models (e.g., time series and topology-based). By sampling multiple future trajectories, a probability forecast distribution and uncertainty interval can be obtained, thereby improving confidentiality. Based on a language model architecture, trained generative AI models promise to exhibit next-word prediction capabilities as impressive as language models themselves. Thanks to the integration of industrial context information disclosed in this paper, existing techniques for pure time-series forecasting are extended, and trained generative AI models can hopefully provide a solid foundation for delivering fairly good and rapid outputs, especially given the significant resources required for training (the cost of time, computation, and expertise needed to train specialized models, and the cost of data). As described, once a trained generative AI model is available, it can be used as a model to generate effective forecasts using only a very small amount of input data.

[0077] Furthermore, it benefits operators. Regarding new topologies and "new industrial plants," for example, using only one week's worth of recorded data, leveraging "base model knowledge" and zero-shot capability, a reasonably reasonable prediction of the operation of a "new industrial plant" can be obtained after fine-tuning a trained generative AI model (i.e., after fine-tuning a specific base model using this week's worth of recorded data), whereas no specific model could be trained earlier or better using only one week's worth of recorded data. Therefore, trained generative AI models are particularly valuable in scenarios requiring rapid time-series prediction models; for any use case, whether it's pattern analysis, anomaly detection, or operational pattern classification. Regarding topology modifications and "modified industrial plants," the reliability of a trained generative AI model may decrease compared to a model trained only on a specific "modified industrial plant," but its accuracy will still be higher than a model trained using only one week's worth of data.

[0078] Furthermore, this also benefits dashboard users. For dashboard users who work with algorithm-based dashboards that observe or monitor processes in operations based on time-series data, and if the dashboard user has an AI / ML model or algorithm to analyze this time-series data, the dashboard user will be able to set up dashboards more quickly using trained generative AI models.

[0079] In addition, trained generative AI models are very valuable when one or more operational processes in an industrial plant require urgent prediction, and / or when the industrial plant does not have a dedicated model available or cannot afford one. Reliability can be improved by sampling several inference runs and thus obtaining a probability forecast distribution.

[0080] According to a sixth aspect, a computer program product including instructions, which, when executed by a computing system, cause the computing system to perform the method according to the first aspect, are provided. The computer program product may include a computer-readable medium that includes the instructions of the computer program product.

[0081] The computer program product in the sixth aspect has the following advantages: it can participate in the fine-tuning of generative AI models for specific industrial plants in an efficient, reliable, and time-saving manner. Furthermore, it can participate in zero-shot hint prediction or forecasting.

[0082] More specifically, trained generative AI models can be understood as foundational models of time series data with context in industrial applications. A key advantage of these foundational models is that they are base models (i.e., trained on large time series datasets), allowing for zero-shot predictions or forecasts. Therefore, foundational models require only a minimal amount of input data (i.e., a sequence of time series data with "context") to enable rapid startup, for example, for a newly installed industrial plant, and to predict the future of such a plant using well-performing and readily available multimodal models (e.g., time series and topology-based). By sampling multiple future trajectories, a probability forecast distribution and uncertainty interval can be obtained, thereby improving confidentiality. Based on a language model architecture, trained generative AI models promise to exhibit next-word prediction capabilities as impressive as language models themselves. Thanks to the integration of industrial context information disclosed in this paper, existing techniques for pure time-series forecasting are extended, and trained generative AI models can hopefully provide a solid foundation for delivering fairly good and rapid outputs, especially given the significant resources required for training (the cost of time, computation, and expertise needed to train specialized models, and the cost of data). As described, once a trained generative AI model is available, it can be used as a model to generate effective forecasts using only a very small amount of input data.

[0083] Furthermore, it benefits operators. Regarding new topologies and "new industrial plants," for example, using only one week's worth of recorded data, leveraging "base model knowledge" and zero-shot capability, a reasonably reasonable prediction of the operation of a "new industrial plant" can be obtained after fine-tuning a trained generative AI model (i.e., after fine-tuning a specific base model using this week's worth of recorded data), whereas no specific model could be trained earlier or better using only one week's worth of recorded data. Therefore, trained generative AI models are particularly valuable in scenarios requiring rapid time-series prediction models; for any use case, whether it's pattern analysis, anomaly detection, or operational pattern classification. Regarding topology modifications and "modified industrial plants," the reliability of a trained generative AI model may decrease compared to a model trained only on a specific "modified industrial plant," but its accuracy will still be higher than a model trained using only one week's worth of data.

[0084] Furthermore, this also benefits dashboard users. For dashboard users who work with algorithm-based dashboards that observe or monitor processes in operations based on time-series data, and if the dashboard user has an AI / ML model or algorithm to analyze this time-series data, the dashboard user will be able to set up dashboards more quickly using trained generative AI models.

[0085] In addition, trained generative AI models are very valuable when one or more operational processes in an industrial plant require urgent prediction, and / or when the industrial plant does not have a dedicated model available or cannot afford one. Reliability can be improved by sampling several inference runs and thus obtaining a probability forecast distribution.

[0086] According to the seventh aspect, use is provided for at least one of the trained generative AI model according to the first aspect, the data processing apparatus according to the second aspect, the data processing system according to the third aspect, the industrial plant according to the fourth aspect, the computer-readable medium according to the fifth aspect, and the computer program product according to the sixth aspect.

[0087] The seventh application has the following advantages: it can help users fine-tune generative AI models for specific industrial plants in an efficient, reliable and time-saving manner; it can also help users achieve zero-shot prediction or forecasting.

[0088] More specifically, trained generative AI models can be understood as foundational models of time series data with context in industrial applications. A key advantage of these foundational models is that they are base models (i.e., trained on large time series datasets), allowing for zero-shot predictions or forecasts. Therefore, foundational models require only a minimal amount of input data (i.e., a sequence of time series data with "context") to enable rapid startup, for example, for a newly installed industrial plant, and to predict the future of such a plant using well-performing and readily available multimodal models (e.g., time series and topology-based). By sampling multiple future trajectories, a probability forecast distribution and uncertainty interval can be obtained, thereby improving confidentiality. Based on a language model architecture, trained generative AI models promise to exhibit next-word prediction capabilities as impressive as language models themselves. Thanks to the integration of industrial context information disclosed in this paper, existing techniques for pure time-series forecasting are extended, and trained generative AI models can hopefully provide a solid foundation for delivering fairly good and rapid outputs, especially given the significant resources required for training (the cost of time, computation, and expertise needed to train specialized models, and the cost of data). As described, once a trained generative AI model is available, it can be used as a model to generate effective forecasts using only a very small amount of input data.

[0089] Furthermore, it benefits operators. Regarding new topologies and "new industrial plants," for example, using only one week's worth of recorded data, leveraging "base model knowledge" and zero-shot capability, a reasonably reasonable prediction of the operation of a "new industrial plant" can be obtained after fine-tuning a trained generative AI model (i.e., after fine-tuning a specific base model using this week's worth of recorded data), whereas no specific model could be trained earlier or better using only one week's worth of recorded data. Therefore, trained generative AI models are particularly valuable in scenarios requiring rapid time-series prediction models; for any use case, whether it's pattern analysis, anomaly detection, or operational pattern classification. Regarding topology modifications and "modified industrial plants," the reliability of a trained generative AI model may decrease compared to a model trained only on a specific "modified industrial plant," but its accuracy will still be higher than a model trained using only one week's worth of data.

[0090] Furthermore, this also benefits dashboard users. For dashboard users who work with algorithm-based dashboards that observe or monitor processes in operations based on time-series data, and if the dashboard user has an AI / ML model or algorithm to analyze this time-series data, the dashboard user will be able to set up dashboards more quickly using trained generative AI models.

[0091] In addition, trained generative AI models are very valuable when one or more operational processes in an industrial plant require urgent prediction, and / or when the industrial plant does not have a dedicated model available or cannot afford one. Reliability can be improved by sampling several inference runs and thus obtaining a probability forecast distribution.

[0092] Based on several examples of this disclosure, and more specifically, the trained generative AI model according to the first aspect can be used in several different application areas. Generally, as outlined above with reference to the method according to the first aspect, the trained generative AI model is used for at least one of the following: i) fine-tuning the trained generative AI model for a second industrial plant different from the first industrial plant; ii) predicting further progress of the operational processes of the second industrial plant. Additionally and / or more specifically, it should be noted that the trained generative AI model can be used for at least one of efficient anomaly detection, quality prediction, and data analysis in newly constructed plants.

[0093] Regarding efficient anomaly detection, for example, detecting faults such as leaks is more reliable using video or audio data (rather than time-series data), with sound and time-series data sometimes being more efficient for detecting pump faults, for instance. For example, a certain fault (like the pump fault mentioned) might cause at least a portion of an industrial plant to emit abnormal sounds or noises. Depending on the type of fault, different sounds or noises may occur. Therefore, based on the different sounds or noises and their corresponding time-series data, the type of fault can be more efficiently guessed, proposed, or determined. Similarly, a certain fault (like the pump fault mentioned) might cause at least a portion of an industrial plant to exhibit abnormal movement or vibration, or visible refrigerant or oil leaks. Depending on the type of fault, different movements, vibrations, or leaks may occur. Therefore, based on the different movements, vibrations, or leaks and their corresponding time-series data, the type of fault can be more efficiently guessed, proposed, or determined.

[0094] According to several examples of this disclosure, industrial context information may include at least one of video data, image data, and / or audio data. For example, video data, image data, and / or audio data may be associated with one or more plant components of a first industrial plant and / or (e.g., for fine-tuning) a second industrial plant. Additionally or alternatively, video data, image data, and / or audio data may be associated with one or more operating processes of the first industrial plant and / or (e.g., for fine-tuning) the second industrial plant. For example, video data, image data, and / or audio data may be associated with one or more starting materials processed by the first industrial plant and / or the second industrial plant, one or more intermediate (manufactured) products produced by the first industrial plant and / or the second industrial plant, and / or one or more final (manufactured) products produced by the first industrial plant and / or the second industrial plant.

[0095] Therefore, for example, context lexical units can represent time units appended with video data, image data, and / or audio data. Thus, a sequence of context lexical units can include a sequence of time units appended with corresponding video data, image data, and / or audio data.

[0096] The use of the trained generative AI model can then include detecting anomalies in the operation process and / or operation data of the second industrial plant, preferably based on a sequence of time units supplemented with corresponding video data, image data and / or audio data.

[0097] Regarding quality forecasts, for example, using videos or subsequent images of the product and timelines, it is possible to better predict the quality of products manufactured in the intermediate stages.

[0098] Therefore, according to several examples of this disclosure, forecasting the operation of a second industrial plant may include forecasting the quality of intermediate or final manufactured products to be manufactured at the second industrial plant.

[0099] Regarding enabling data analytics in new factories, it is very difficult to conduct data analytics or apply ML-based models from the outset. Therefore, having a trained generative AI model (optionally, with some initial fine-tuning) can help provide several uses or benefits.

[0100] Therefore, according to several examples of this disclosure, the second industrial plant may be a newly built plant.

[0101] With appropriate modifications, the optional features described in the first aspect can form part of any of the second to seventh aspects.

[0102] The computer-readable medium described in the fifth aspect can store the computer program product described in the sixth aspect.

[0103] As used herein, the term “acquire” can include, for example, receiving from another system, device, (AI / ML) model, or process; receiving via interaction with a user; loading or retrieving from a storage device or memory; measuring or capturing using a sensor or other data acquisition device; and receiving or acquiring results from one or more data processing steps.

[0104] The indefinite articles “one” or “a” do not exclude multiple. Furthermore, unless otherwise specified or clearly indicated from the context to be a single form, the articles “one” and “a” as used herein should generally be interpreted as meaning “one or more”.

[0105] Unless otherwise specified or as may be clearly understood from the context, the phrases “one or more of A, B, and C,” “at least one of A, B, and C,” and “A, B, and / or C” as used herein are intended to mean all possible combinations of one or more of the listed items. That is, the phrase “A and / or B” means (A), (B), or (A and B), while “A, B, and / or C” means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C).

[0106] The term "comprising" does not exclude other elements or steps. Furthermore, the terms "comprising," "including," "having," etc., may be used interchangeably herein.

[0107] This invention may include one or more aspects, examples, or features that can be implemented individually or in combination, whether specifically disclosed in combination or individually. Any optional feature or sub-aspect of one of the foregoing aspects may be applied to any other aspect as appropriate.

[0108] The aspects described above will become apparent and will be clarified through the specific implementations provided below. Attached Figure Description

[0109] The following detailed description will be given with reference to the accompanying drawings, using only examples: - Figure 1 Examples of lexicalization procedures according to several examples of this disclosure are illustrated; - Figure 2 The diagram illustrates the training and inference phases of a base model according to several examples of this disclosure; - Figure 3 The diagram illustrates flowcharts of methods according to several examples of this disclosure; - Figure 4 The illustration shows a block diagram schematically illustrating a data processing apparatus according to several examples of the present disclosure; and - Figure 5 The illustrations schematically depict how to train and use a base model according to several examples of this disclosure. Detailed Implementation

[0110] Based on several examples of this disclosure, a method (e.g., and data processing apparatus and data processing system) is disclosed for the functionality of a basic model for obtaining time-series data of an operational process with industrial context information for industrial applications.

[0111] It's important to note that the term "base model functionality" refers to the functionality of a ML-based model. More specifically, this means that the ML-based model possesses extensive knowledge and understanding of the subject of interest and is typically ready to use out of the box without any additional work (such as further dedicated, resource-intensive data-driven training, calibration, tuning, or optimization), exhibiting good applicability and usability. In other words, a ML-based model with base model functionality can, for example, possess extensive knowledge and understanding of the operational processes of starting and / or optimizing one or more specific industrial plants (i.e., the subject of interest). For example, these one or more industrial plants can be understood as representing one or more industrial plants used in a particular industrial sector (such as in the mining industry or the chemical industry). Therefore, if the industrial plant is newly incorporated or embedded in such a specific industrial sector, for example, a ML-based model with base model functionality is already quite suitable—that is, ready to use out of the box and well applied to the newly incorporated, embedded, or installed industrial plant.

[0112] Therefore, the base model requires very little input data (like an input time series data sequence) to enable rapid startup based on this data (e.g., for a newly installed industrial plant) and to enable various data-driven applications (such as anomaly detection, forecasting, and data labeling) for the newly installed industrial plant, which have well-performing and readily available multimodal (like time series and topology) models.

[0113] Specifically, according to several examples of this disclosure, a method and data processing apparatus for abstracting representations (i.e., encoding) of time-series data and its industrial context information are disclosed. Using the time-series data and its industrial context information, a base model can be pre-trained using historical operational data and / or future operational data. As one of its greatest advantages, the base model can then be used for (1) low-cost fine-tuning and / or (2) prompting with zero-sample forecasting and predictive capabilities.

[0114] This means that the advantage of having such a base model is that users require very little training or input data, whether for fine-tuning or direct hints. For example, a user might only need a short time-series data sequence and some industrial context information (i.e., fragments of industrial context information). This could include, for example, which plant component the short time-series data sequence originates from, what industry (or industrial sector) it relates to, or the current temperature. Users require very little training or input data to make fairly confident future predictions based on this data. Additionally or alternatively, various data-driven applications, such as anomaly detection, forecasting, and data annotation, can also be implemented. It should be noted that, according to several examples of this disclosure, an industrial plant may include multiple plant components, and the time-series data may include plant component time-series data of the plant component operation processes associated with the multiple plant components. It should also be noted that, according to several examples of this disclosure, the plant component operation processes may consist of fragments of the industrial plant's operation processes, and fragments of plant component time-series data may correspond to one or more corresponding plant components among the multiple plant components. Industrial context information, or fragments of industrial context information (as outlined in more detail below), may include industrial context information of the plant components that is associated with the plant component operation process.

[0115] "Relatively confident future predictions" means that industrial plant owners or operators do not need to train a completely new model from scratch (which could take weeks), but can quickly get started using a base model. For example, for a newly installed industrial plant, they can use a "not useless," fairly well-functioning model that has seen a large amount of time-series data incorporating industrial context information. It should be noted that this contrasts with a base model that only has time series data (without context), where one cannot rely on obtaining reasonably context-dependent outputs, because without additional context, the time series of plant component "X" might be very similar to, for example, the time series of physical vitality parameters or the financial stock valuation process. Therefore, because context is taken into account, and because the base model is trained on time series that include context, one can rely on the predictions made by the base model based on several examples of this disclosure (at least better than no predictions), since all these predictions are at least somewhat relevant to the industrial plant due to context.

[0116] Note that for “fairly confident future predictions,” the confidence can be further enhanced using standard methods used for word / sentence predictions in large language models (LLMs) based on standard text. This can be simple, for example, by having the base model make ten future predictions with the same input, and then averaging those predictions—that is, sampling the trajectories of multiple predictions to obtain a prediction distribution with a most probable trajectory and an interval of uncertainty. Alternatively, for example, by excluding certain behaviors of an industrial plant (or one or more industrial plant components) to which the base model is applied, the user can request additional information from the base model to reduce uncertainty.

[0117] Based on several examples of this disclosure and reference Figure 1 (It shows a high-level visualization of a lexicalization procedure according to several examples of this disclosure), the training of this base model can be performed as follows (i.e., implementation details will be given below):

[0118] The model underlying this base model is based on a language model architecture. In the lexicalization step (a preparation step for training the base model), for example, after optionally scaling, mean scaling 103, and / or quantizing 104 of the time-series data 101, the time-series data 101 is transformed into a series of “preliminary lexical units” 102 (e.g., as shown, these preliminary lexical units are segmented and grouped into preliminary lexical units 102a, 102b, 102c, 102d, and 102e). For example, a moving time window or similar method can be used. Each of these preliminary lexical units 102a through 102e is appended with an additional context vector or context label containing additional information such as industrial context information 105. For example, this additional information could be about one or more (industrial plants of interest) plant components from which time-series data has been recorded or obtained, and this time-series data is included in the time-series data 101. Additionally or alternatively, this additional information or industrial context information can be part of the overall plant topology (i.e., the plant topology of the industrial plant of interest), and / or may even indicate environmental conditions, such as weather or temperature information, timestamped video recordings, etc., so that they can be aligned with time-series data. This yields the "final context lexicon"106. Furthermore, the resulting embeddings can be placed together in a joint embedding space for ML-based model algorithms to learn on specific tasks.

[0119] It is worth noting that, according to Figure 1 The proposed lexicalization scheme or procedure illustrated surpasses existing technologies because it integrates additional contextual information along with or aligned with the corresponding time-series data. Therefore, a reasonable industrial time-series foundation model can be obtained based on this.

[0120] See now Figure 2 The illustration depicts a training phase 201 and an inference phase 202 of a base model according to several examples of this disclosure, using methods as described in... Figure 1 The illustrated lexicalization scheme or procedure can be used to obtain the lexical sequence 106 (e.g., Figure 1 The diagram illustrates training a (temporal) language model 203 on a lexical sequence 106. This training may include feeding the lexical sequence 106 into the language model 203; and training the language model 203 using, for example, masking techniques (i.e., masking the corresponding next lexical and training to predict this corresponding next lexical) and a suitable loss function. For example, the loss function could be a cross-entropy loss 204. The language model 203 can be an encoder-decoder model or a decoder-only model. After the language model 203 is trained, the language model 203' can be understood as a representation of the underlying model.

[0121] For pre-training of the base model (i.e., language model 203), there are transformer architectures for temporal data and methods for probabilistic prediction, which can be used to train a neural network on lexicalized samples 106. Furthermore, pre-training can be performed in a self-supervised manner using tasks such as multivariate prediction and / or introducing perturbations into the data. For example, introducing perturbations into the data can include, for instance, interpolating missing values ​​by randomly or strategically masking certain time points. Introducing perturbations into the data can also include outlier interpolation by identifying and removing or masking outliers. Additionally, multimodal temporal models can be made to work with fewer modalities; for example, when data from some modalities is missing, a set of zero values ​​can be strategically or randomly introduced into some modalities. For example, zero values ​​can be introduced into video data in some samples to indicate that data is missing from the video modality; while in other samples, both video and audio data can be set to zero values. This helps the model learn from a single or few modalities, enabling inference even when some input data is missing.

[0122] To fine-tune the base model (i.e., language model 203'), the pre-trained base model can be further fine-tuned on downstream tasks (such as prediction, multi-class classification, anomaly detection, etc.).

[0123] According to several examples of this disclosure, regarding implementation, the training of the base model can preferably employ at least "quantized-aware training," and the quantized and non-quantized states of the weights can be preserved, thereby improving performance. Furthermore, the forward propagation (for inference) of the ML / AI model can use this quantized version, while the backpropagation can use the non-quantized version to improve performance.

[0124] After training is complete, in the next stage, namely in inference phase 202 (i.e., the use or application phase of the base model), probabilistic predictions can be obtained by sampling multiple future trajectories given the historical context. More precisely, 205 words are autoregressively sampled from the base model and mapped back to numerical values ​​representing temporal sequence and context.

[0125] As mentioned above, multiple trajectories can be sampled to obtain a more reliable prediction distribution. It's important to note that a simple consistency check can be performed by comparing the (decoded) output context information 206 with the initial (unencoded) input context information. For example, simply put, if the temperature at a time point / time interval in the input time series 101 is 25°C, and only in the "predicted output" does it suddenly change to 10°C after decoding, then this is clearly an error.

[0126] It should be noted that, further in terms of training implementation and data requirements, a relatively large amount of data is needed to train or pre-train the language model 203 so that it can serve as a reasonable and comprehensive base model. This means, for example, data from various plant components in the plant topology (e.g., distillation columns, storage tanks, reactors, level transmitters, etc.) and under various environmental conditions (e.g., warm, cold, humid, dry, etc.) are needed, ideally covering almost all combinations.

[0127] The more comprehensive the understanding of the "world" of language model 203' (i.e., the more data language model 203 has about various conditions and their variations and all combinations, or the more data language model 203' has to train about various conditions and their variations and all combinations), the better language model 203' will be able to reason in future use cases, because language model 203' has learned how the "world" works.

[0128] In addition, it should be noted that large amounts of data can be collected or recorded from actual factories, or they can be obtained as synthetic data, for example, by using Gaussian processes.

[0129] Furthermore, according to several examples of this disclosure, distributed data collection and training are also possible. More specifically, since the pre-training task of the language model 203 can be completed in a self-supervised manner, the language model 203 can be pre-trained and weights merged in a distributed manner using data from different departments of the industrial plant of interest or from different industrial plants of interest from the same customer. The same approach is also applicable to fine-tuning, especially in the cases of forecasting and anomaly detection, where forecast values ​​are typically available after a certain period of time, and such samples can be easily collected online. For anomaly detection, the time series can be considered normal until an alarm or human intervention is detected, after which feedback from the operator can be requested for data annotation. These methods help overcome the problem of data availability.

[0130] The resulting base model (i.e., the trained language model 203') can be fine-tuned individually, for example, for the pulp and paper industry (e.g., compared to the mining industry), or for northern climates (e.g., compared to southern climates), or for chemical process industrial plants and specific plant components and alarms, etc. (e.g., compared to oil and gas plants with other conditions).

[0131] Similar to mainstream text-based LLMs, the potential and reliability of these basic models can be improved by fine-tuning them with domain-specific time-series data and domain-specific contextual information.

[0132] It is important to note that, under normal circumstances, the alignment of multimodal context information data with time-series data is most likely or most frequently accomplished by using appropriate timestamps.

[0133] Based on several examples of this disclosure, one or more of the following data sources may be used or considered for obtaining time-series data and / or industry context information: - Sensor data; - Tank level, reaction time, volume of raw materials and / or products per unit time, quantity of raw materials and / or products per unit time, etc.; - Alarms and / or events; - Manual logging; - (Operating) parameters and states or state changes, and combinations thereof, of industrial plants and / or plant components; - Weather data, such as temperature or humidity data; - etc.

[0134] According to several examples of this disclosure, industrial context information may include at least one of the following: - The topology of an industrial plant, which indicates which plant components are linked to which other plant components(s), and therefore, which plant components’ timing data declines or rises might affect which other plant components, etc. - Industrial information models and / or industrial sector information models, such as how time-series data relate to each other or how they relate to the overall operation of an industrial plant; - etc.

[0135] Based on several examples of this disclosure, the underlying model disclosed herein can be used, for example, out of the box, directly for any time-related application, such as for time-based anomaly detection, pattern analysis, or driving runtime pattern classification, or for maintaining due date calculations, etc.

[0136] It should be noted that, for ease of understanding, although explicit dedicated data analysis models can be used instead of the basic model disclosed herein, these dedicated models require a significant investment of effort and time to train specialized schemes, whereas the basic model method disclosed herein, based on several examples, does not require such effort.

[0137] Based on several examples in this disclosure, the basic model can function in different application areas.

[0138] For example, for efficient anomaly detection: video or audio data (not to mention time-series data) can be used to detect faults such as leaks more robustly, with pump fault detection sometimes being more efficient using both audio and time-series data.

[0139] For example, for quality forecasting: using video or subsequent images of the product along with the timeline can better predict the quality of products manufactured in the intermediate stages.

[0140] For example, when implementing data analytics in a new factory: it is very difficult to perform data analytics or ML models from scratch for a new factory, where pre-training a base model and making some initial fine-tuning can help provide several uses.

[0141] Therefore, many industrial use cases can be realized by using datasets more efficiently and comprehensively.

[0142] See now Figure 3 , Figure 3 A flowchart illustrating methods according to several examples of this disclosure is provided. This method is a foundational model for obtaining time-series data of an operational process with industrial context information for industrial applications. The time-series data can be as described in the reference... Figure 1 The illustrated time series data 101. Industrial context information can be as shown in the reference. Figure 1 The industrial context information 105 is illustrated.

[0143] This method starts from S300.

[0144] In S310, the method includes transforming time-series data of an operation process associated with at least a first industrial plant into a time-unit sequence associated with a corresponding segment of the operation process.

[0145] In S320, the method includes attaching industrial context information to time units in a time unit sequence, wherein each segment of the industrial context information is associated with a segment of the operation process, and wherein the time unit is attached with concurrent segments of industrial context information.

[0146] In S330, the method includes obtaining a sequence of context lexical units from appended information, where each context lexical unit represents a time unit appended with concurrent fragments of industrial context information. For example, the sequence of context lexical units can be obtained using a lexicalization method. The sequence of context lexical units can be a reference... Figure 1 and Figure 2 The illustrated context lexical sequence 106.

[0147] In S340, the method includes training a generative AI model on a sequence of context lexical units using a loss function. This generative AI model (either untrained or undergoing pre-training) can be a reference. Figure 2 The language model 203 shown in the figure.

[0148] In S350, the method includes using a trained generative AI model with basic model functionality to perform at least one of the following operations: i) fine-tuning the trained generative AI model for a second industrial plant different from the first industrial plant; and ii) predicting further developments in the operational processes of the second industrial plant. The trained generative AI model can be a reference... Figure 2 The illustrated example is a trained language model 203'.

[0149] This method ends at S360.

[0150] According to several examples of this disclosure, the transformation in S310 may include: obtaining a time unit sequence by mean scaling and / or quantizing the time series data, and transforming the time series data into a time unit sequence. Figure 1 Illustrative examples of mean scaling and / or quantization are shown in the figure.

[0151] According to several examples of this disclosure, the additions in S320 may include: an additional context vector for each time unit, the context vector being associated with a segment of industrial context information, each segment being associated with a corresponding time unit. (Reference) Figure 1 The context vector is explained in more detail above.

[0152] According to several examples of this disclosure, fragments of industrial context information may also indicate at least one of the following: - Environmental data associated with the operational processes of at least the first industrial plant; - Environmental data associated with the operation of multiple factory components; - An industrial domain in which at least one industrial plant is operated; - Alarm and event data associated with at least the first industrial plant; - Multimodal data regarding multiple factory components; and - Topological context, in which at least the first industrial plant is embedded.

[0153] According to several examples of this disclosure, training in S340 may include performing the training in a self-supervised manner using at least one of the following task weights: multivariate prediction for context lexical sequences, and introducing perturbations into context lexical sequences.

[0154] According to several examples of this disclosure, i) fine-tuning in S350 may include: fine-tuning a trained generative AI model using fine-tuning training data, wherein the fine-tuning training data is associated with the operation process of the second industrial plant; wherein ii) forecasting in S350 may include forecasting by using the fine-tuned trained generative AI model.

[0155] According to several examples of this disclosure, i) fine-tuning in S350 may include fine-tuning the trained generative AI model for at least one of the following: - Applications that are to be operated in specific types of industries in secondary industrial plants; - Applications that execute specific processes in a secondary industrial plant; and - Applications in a second industrial plant that need to operate in a specific climate zone.

[0156] According to several examples of this disclosure, fine-tuning training data may include temporal fine-tuning data and associated fine-tuning context information, wherein the temporal fine-tuning data and fine-tuning context information may be associated with the operation of a second industrial plant and / or with the operation of plant components of the second industrial plant. According to several examples of this disclosure, the temporal fine-tuning data may be specific to an industrial context domain associated with the second industrial plant. According to several examples of this disclosure, the context information may be specific to an industrial context domain.

[0157] According to several examples of this disclosure, time-series data and / or time-series fine-tuning data may include at least one of three types of training data: historical data from actual operational processes, synthetic data from synthetic operational processes, and synthetic data from historical data. Synthetic data from historical data may be synthesized via data manipulation (e.g., by adding noise).

[0158] According to several examples of this disclosure, the forecast in step S350 (ii) may include: obtaining a probabilistic forecast of the operational behavior of the second industrial plant by sampling multiple forecasted future trajectories.

[0159] According to several examples of this disclosure, a time unit may include a point in time and / or a period of time.

[0160] See now Figure 4 , Figure 4 A block diagram is shown, schematically illustrating a data processing apparatus 400 according to several examples of the present disclosure. Specifically, according to several examples of the present disclosure, a data processing apparatus 400 is provided for the functional basis model of obtaining time-series data of an operational process with industrial context information for an industrial application. The data processing apparatus 400 includes one or more processors 401 configured to perform operations according to... Figure 1 and Figure 2 The indicated methods and / or execution are as described above. Figure 3 The methods outlined herein.

[0161] According to several examples of this disclosure, the data processing apparatus 400 may include a component that acts as referenced above. Figure 2 The language model 203 and / or the trained language model 203' described herein.

[0162] More specifically, based on various examples, it is configured to execute Figure 3 The data processing apparatus 400 of the method may include a processing circuit system, processing functions, processing components, processing units, or a processor 401. These components enable the data processing apparatus 400 to participate in the basic model functionality of obtaining timing data of operational processes with industrial context information for industrial applications. The processor 401 may include one or more processing portions or functions, wherein the processing portions or functions may be provided as one or more physical or virtual entities. The data processing apparatus 400 may include one or more communication interfaces 402. The data processing apparatus 400 may also include a memory or storage unit 403 for storing data, programs, and / or instructions to be executed by the processor. The memory 403 may be internal to the data processing apparatus 400 or external to the data processing apparatus 400 (e.g., on a cloud server). The processor 401 may include one or more portions that enable the data processing apparatus 400 to perform, for example... Figure 3 The method. According to several examples of this disclosure, the transformation section 410 can be configured to according to Figure 3 The S310 is used to perform this transformation, and the additional part 420 can be configured to perform the transformation according to... Figure 3 The S320 is used to perform this additional operation, and part 430 can be configured to perform this operation according to... Figure 3 The S330 is used to perform this acquisition, and the training part 440 can be configured to perform it according to... Figure 3 The S340 is used to perform this training, and a portion of the 440 can be configured to... Figure 3 The S350 is used to perform this usage.

[0163] According to several examples of this disclosure, corresponding portions of the data processing apparatus 400 may also be understood as components for performing specific functions.

[0164] See now Figure 5 , Figure 5 The illustrations schematically depict how to train and use a base model (i.e., how to train a generative AI model to obtain the desired base model) based on several examples of this disclosure.

[0165] In light of this, it is important to note that, according to several examples in this disclosure, time-series data is considered together with industry context information. In other words, industry context information is not simply considered as textual or linguistic context information. Rather, all kinds of context information related to an industrial plant are also considered industry context information. See below for reference. Figure 5 This will be described in more detail.

[0166] All the contextual information related to industrial plants is explicitly not limited to text, such as industry-specific or plant-specific information or manuals. It also includes multimodal data, such as: - Topology information, such as: how the component or plant component of interest is embedded in the overall topology, what other components or plant components are before or after the corresponding component, what characteristics those components have, what alarms, adjustments, etc. The terms "before" and "after" will be understood, for example, when viewing the topology and determining which first component is connected to the second component and which third component is connected to the second component—the process can run from the first component to the second component, and then to the third component. For example, steam runs from the first tank (first component) to the second tank (third component) via a control valve (second component). - Environmental information, such as temperature, humidity or other time-series information (such as seasonality, date, etc.). - Information on other influencing factors of the relevant components of interest, such as, for a distillation column, the information could be a level controller, connected tanks, or quantities such as physical quantities, tank levels, sensor values, key performance indicators (KPIs), which can be represented in a knowledge graph or other graphical format.

[0167] For example, such as Figure 5 As illustrated in the system 500, timing data 501 and industrial context information 502 (e.g., including audio recordings 502a, video recordings 502b, and alarm and event data 502c) are stored in or accessible to one or more data storage devices 503.

[0168] At data preparation step 504, timing data 501 and industrial context information 502 can be prepared for one or more devices. For example, timing data for first device 505a, as well as recording and alarm / event data for first device 505a, can be prepared. Similar preparation can be performed for devices 505b and 505c. In particular, the timing data is aligned with respect to different assets and their corresponding topologies and other modalities.

[0169] As a result of data preparation step 504, joint embeddings of time series (TS), video, alarms and events (A&E), and topology can be obtained by step 507 (pre-training pipeline, e.g., contrastive learning with selective sparse context).

[0170] Based on several examples of this disclosure, the term “sparse” should be understood as having fewer “coupled” instances.

[0171] More specifically, in the pre-training pipeline of the base model, there can be, for example, a "contrastive learning" setting with selectively sparse context, where "sparse" means having fewer "coupled" examples. That is, for example, only a small number of time series 501 with topological context information serve as industrial context information 502, only a small number of time series 501 with temperature context information serve as industrial context information 502, only a small number of time series 501 with industrial segmentation information serve as industrial context information 502, and so on. Therefore, the data point space is very sparse overall, but there is still sufficient overlap across all "context" dimensions.

[0172] By training the base model on a sparse context (i.e., sparse industrial context information) according to several examples of this disclosure, such as through contrastive learning, it is also possible to input time series as well as sparse context. This permission can be understood as a detectable differentiator, where the sparse context can be understood as a flexible context. This means that since the base model is not trained with the complete context (complete industrial context information), for example, for interpretation purposes, the training data above has time series with topological structure and other time series with temperature information, but only a very small number of time series have both topological structure and temperature information, it is not necessary for the time series to include this complete information during the inference phase. Conversely, time series with sparse context can also be processed. This allows for great flexibility.

[0173] The obtained basic model 508 (e.g.) Figure 5 The context-informed timing model illustrated in the figure can represent the above reference. Figure 2 The trained language model 203' outlined above, i.e., the one mentioned above, refers to... Figure 3 The trained generative AI model outlined in steps S340 and S350. Optionally, as already based on... Figure 3 As indicated in step S350, the base model 508 can be fine-tuned to obtain a fine-tuned base model 509, which can be fine-tuned for different tasks (e.g., for anomaly detection or quality prediction). This fine-tuning can be performed for a specific plant 510, for example, as described above. Figure 3The second factory mentioned in (step S350). The second factory 510 can be a newly built factory or a renovated factory.

[0174] According to several examples of this disclosure, a data processing system is provided for the functionalities of a basic model for obtaining time-series data of an operational process with industrial context information for industrial applications. The data processing system includes, according to... Figure 4 The data processing device 400 and / or includes a means for performing according to Figure 3 The system can be a component of the method. Figure 5 The system 500 shown in the figure.

[0175] According to several examples of this disclosure, an industrial plant is provided, the industrial plant comprising, according to Figure 4 The data processing device 400 and / or the data processing system as outlined above. The industrial plant may be an industrial plant to which a trained generative AI model will be applied, i.e., the industrial plant may be as described in the reference above. Figure 3 The second industrial plant outlined herein.

[0176] According to several examples of this disclosure, a computer-readable medium including instructions that, when executed by a computing system, cause the computing system to perform as described above (referenced above) Figure 1 and Figure 2 The indicated methods and / or execution are as per the reference. Figure 3 The methods outlined herein. The computer-readable medium may be transient or non-transient, volatile or non-volatile.

[0177] According to several examples of this disclosure, a computer program product including instructions is provided that, when executed by a computing system, cause the computing system to perform as described above. Figure 1 and Figure 2 The indicated method, and / or the execution as per reference Figure 3 The methods outlined above. The computer program product may include a computer-readable medium that includes instructions for the computer program product. The computer-readable medium, as mentioned above, may store the computer program product thereon.

[0178] According to several examples of this disclosure, uses are provided for a data processing apparatus 400, a data processing system as outlined above, an industrial plant as outlined above, a computer-readable medium as outlined above, and / or a computer program product as outlined above. In particular, uses are provided as described in the references... Figure 3The outlined method is used to provide a foundational model for obtaining time-series data of operational processes with industrial context information for industrial applications. Furthermore, by fine-tuning the trained generative AI model on a second industrial plant, and / or by predicting the operational processes of the second industrial plant, the trained generative AI model (i.e., Figure 2 The use of the trained language model 203' shown in the figure.

[0179] After appropriate modifications, for reference Figure 3 The optional features of the methods outlined can form part of data processing apparatus 400, data processing system, industrial plant, computer-readable medium, computer program product, and their uses.

[0180] Any unit, module, circuit system, or method described herein may be implemented using hardware, software, and / or firmware configured to perform any of the operations described herein. Hardware may include one or more processor cores, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), etc. Software may be embodied as software packages, code, instructions, instruction sets, and / or data recorded on at least one transient or non-transitory computer-readable storage medium. Firmware may be embodied as code, instructions, or instruction sets and / or data hard-coded in a memory device (e.g., a non-volatile memory device).

[0181] If implemented in software, the function can be stored on or transmitted through a computer-readable medium as one or more instructions or code on that medium. Computer-readable media includes computer-readable storage media. A computer-readable storage medium can be any storage medium accessible to a computer. By way of example, and not limitation, such a computer-readable storage medium can include FLASH storage media, RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, disk storage devices or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and is accessible to a computer. As used herein, “disk” and “optical disc” include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs (BDs), where disks typically copy data magnetically and optical discs typically copy data optically using lasers. Furthermore, the propagation of signals can also be included within the scope of computer-readable storage media. Computer-readable media also includes communication media, which includes any medium that facilitates the transfer of a computer program from one place to another. For example, a connection can be a communication medium. For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared, radio, and microwave), then that coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave) is included in the definition of communication medium. Combinations of the above technologies should also be included within the scope of computer-readable media.

[0182] The applicant hereby separately discloses each individual feature described herein, as well as any combination of two or more such features, provided that such feature or combination can be implemented as a whole based on the description in accordance with the common sense of those skilled in the art, regardless of whether such feature or combination solves any problem disclosed herein, and without being limited by the scope of the claims. The applicant notes that aspects of the invention can consist of any such individual feature or combination of features.

[0183] It should be noted that embodiments of the present invention are described with reference to different categories. In particular, some examples are described with reference to methods, while other embodiments are described with reference to apparatus. However, those skilled in the art will understand from the description that, unless otherwise notified, any combination of features related to different categories, in addition to any combination of features belonging to one category, is also considered to be disclosed in this application. However, all features can be combined to provide a synergistic effect greater than the simple sum of the features.

[0184] Although the invention has been illustrated and described in detail in the accompanying drawings and foregoing description, such illustrations and descriptions should be considered exemplary rather than limiting. The invention is not limited to the disclosed embodiments. By studying the drawings, the disclosure, and the appended claims, those skilled in the art will understand and implement other variations of the disclosed embodiments.

[0185] The fact that certain measures are described in mutually different dependent claims does not mean that combinations of these measures cannot be used advantageously.

[0186] No reference numerals in the claims shall be construed as limiting the scope of the claims.

Claims

1. A computer-implemented method for obtaining a basic model functionality of timing data (101) of an operational process with industrial context information for industrial applications, the method comprising: The time-series data (101) of the operation process associated with at least the first industrial plant is transformed (S310) into a time unit sequence associated with the corresponding segment of the operation process; The industrial context information (105) is appended (S320) to the time units in the time unit sequence, wherein each segment of the industrial context information is associated with each segment of the operation process, and wherein the time unit is appended with concurrent segments of the industrial context information; A context lexical sequence (106) is obtained from the appended information using a lexicalization method (S330), wherein one context lexical represents a time unit, the time unit being appended with concurrent fragments of the industrial context information; Using a loss function, a generative artificial intelligence (AI) model (203) is trained (S340) on the context lexical sequence (106); and The trained generative AI model (203') with basic model functionality is used (S350) for at least one of the following operations: i) fine-tuning the trained generative AI model (203') for a second industrial plant different from the first industrial plant; and ii) predicting further progress of the operation process of the second industrial plant.

2. The method according to claim 1, wherein the transformation comprises: The time series data is transformed into the time unit sequence by means scaling and / or quantization to obtain the time unit sequence.

3. The method according to claim 1 or 2, wherein the additional component comprises: Each time unit in the time unit is attached with a context vector, which is associated with the segments of the industrial context information, which are associated with the corresponding time units.

4. The method according to any one of claims 1 to 3, The aforementioned at least the first industrial plant includes multiple plant components; The timing data includes factory component timing data of factory component operation processes associated with the plurality of factory components, wherein the factory component operation process consists of segments of the operation process, and wherein each segment of the factory component timing data corresponds to one or more corresponding factory components among the plurality of factory components; and The fragments of the industrial context information include industrial context information of the plant components associated with the operation process of the plant components.

5. The method according to any one of claims 1 to 4, wherein each fragment of the industrial context information further indicates at least one of the following: Environmental data associated with the operating process of the at least first industrial plant; Environmental data associated with the operation of the plant components of the plurality of plant components; An industrial domain, in which at least the first industrial plant is operated; Alarm and event data associated with at least the first industrial plant; Multimodal data regarding the multiple factory components; as well as A topological context, wherein at least the first industrial plant is embedded in the topological context.

6. The method according to any one of claims 1 to 5, wherein the training comprises: The training is performed in a self-supervised manner by using at least one of the following tasks: multivariate prediction for the context lexical sequence and introducing perturbations into the context lexical sequence.

7. The method according to any one of claims 1 to 6, wherein the fine-tuning comprises: The trained generative AI model is fine-tuned using fine-tuning training data, wherein the fine-tuning training data is associated with the operation process of the second industrial plant; The forecasts mentioned include forecasting using a fine-tuned, trained generative AI model.

8. The method according to any one of claims 1 to 7, wherein the fine-tuning comprises: The trained generative AI model is fine-tuned for at least one of the following: Applications that are to be operated in a specific type of industry at the second industrial plant; Applications that execute specific processes on the second industrial plant; as well as Applications that need to be operated in a specific climate zone at the second industrial plant.

9. The method according to claim 7 or 8, wherein the fine-tuning training data comprises: Timing fine-tuning data and fine-tuning context information associated with the timing fine-tuning data, wherein the timing fine-tuning data and the fine-tuning context information are associated with the operation process of the second industrial plant and / or with the operation process of the plant components of the second industrial plant.

10. The method according to any one of claims 1 to 9, wherein the time-series data and / or time-series fine-tuning data includes at least one of the following three types of training data: historical data from actual operation processes, synthetic data from synthetic operation processes, and synthetic data from historical data.

11. The method according to any one of claims 1 to 10, wherein the forecast comprises: By sampling multiple predicted future trajectories, a probabilistic prediction of the operational behavior of the second industrial plant is obtained.

12. The method according to any one of claims 1 to 11, The loss function mentioned therein is the cross-entropy loss function; and / or The time unit includes a point in time and / or a time period; and / or The industrial context information is sparse industrial context information, and the context lexical sequence is obtained by appending the sparse industrial context information to the time unit.

13. A data processing apparatus (400) comprising one or more processors configured to perform the method according to any one of claims 1 to 12.

14. A computer program product comprising instructions that, when executed by a computing system, cause the computing system to perform the method according to any one of claims 1 to 12, and / or enable the computing system to perform the method according to any one of claims 1 to 12.

15. A computer-readable medium having thereon stored a computer program product according to claim 14.