Key process variable prediction method based on InduDK-LLM model

By constructing the InduDK-LLM model and combining industrial technical documents and time series data of operational variables, the relationship between process variables and operational variables is established, which solves the problem of insufficient prediction accuracy in existing methods and achieves high-precision prediction of industrial process variables.

CN121093104AActive Publication Date: 2025-12-09JIANGNAN UNIV

Patent Information

Application Number
CN202511642862.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2025-12-09
Estimated Expiration
2045-11-11

AI Technical Summary

Technical Problem

Existing deep learning methods cannot effectively capture the complex logical relationships between process variables and operational variables in industrial time series forecasting, resulting in insufficient prediction accuracy. Existing large language models do not fully utilize industrial domain knowledge and cannot achieve high-precision industrial process variable forecasting.

Method used

We construct an InduDK-LLM-based model, which establishes the relationship between process variables and operational variables by combining industrial technical documents and time series data of operational variables with an inverted embedding module, a static wavelet decomposition module, a process joint block module, a domain knowledge and operational variable embedding module, an interaction module, and a prediction module. We then leverage the powerful reasoning capabilities of a pre-trained large language model for prediction.

Benefits of technology

It achieves high-precision prediction of key industrial process variables, improves the accuracy and efficiency of prediction models, effectively captures complex dependencies and potential logical connections between variables, and is applicable to a variety of industrial processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121093104A_ABST
    Figure CN121093104A_ABST
Patent Text Reader

Abstract

The invention discloses a key process variable prediction method based on an InduDK-LLM model, and the method specifically separates the modeling of an industrial process variable and an operation variable, and combines historical time sequence data and industrial field knowledge. Firstly, a pre-trained LLM is introduced to process time series data of domain knowledge and operation variables so as to generate industrial prior codes rich in semantics, and the pre-trained LLM reduces the requirement for manual prompt engineering. In addition, the process variables are decomposed through a multi-scale method, and meanwhile, a learnable joint mark is designed to promote interaction between the two types of variables. And finally, designing an interaction module to model an interdependency relationship and a potential logical relationship between the two types of variables. According to the method, accurate prediction of industrial key process variables is realized, and reference and guidance are provided for actual industrial production.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a key process variable prediction method based on an InduDK-LLM model and belongs to the field of industrial key process variable prediction. BACKGROUND

[0002] In industrial processes, accurate prediction of key process variables is crucial for ensuring product quality, process safety, and production efficiency in industrial production. Reliable prediction can often achieve proactive control, timely troubleshooting, and improve the consistency of the final product, thereby becoming the basis for process monitoring and control. Actual industrial time series data usually has the characteristics of high dimensionality, heterogeneity and multivariate. They can be conceptually divided into process variables (such as temperature, concentration) and operation variables (such as solvent addition rate, valve state, operation step). These two types of variables are often closely coupled and subject to time-varying logical relationships. Changes in operation variables often cause transient or long-term changes in process variables, and the strength of these coupling relationships may change with changes in operating conditions. Therefore, extracting interaction features that can describe the dependence relationship between the two types of variables is of great significance for accurate industrial process modeling and prediction.

[0003] Existing deep learning methods, such as convolutional neural networks (CNN), long short-term memory networks (LSTM), and gated recurrent units (GRU), have been widely used in industrial time series analysis and prediction. CNN can effectively extract representative local features in industrial time series by sliding one-dimensional convolution kernels in the time dimension; recurrent neural networks and their variants are excellent at processing sequence data, effectively capturing long-term patterns and periodic rules in time series. However, when faced with long sequence industrial time data, the limited receptive field of CNN gradually averages, smooths or dilutes long-term dependency information after multiple convolution layers; and the long time transmission chain of recurrent neural networks and their variants cannot be parallel computed; these lead to low accuracy and efficiency of the two methods, which cannot meet the requirements of actual industrial production. In order to solve this problem, models based on Transformer have received widespread attention. Transformer uses multi-head attention mechanism to efficiently model long-distance dependencies in sequence data; and through the expansion of parallel computing, it effectively improves the efficiency of the model. However, existing Transformer-based methods often ignore the complex logical relationship between process variables and operation variables in industrial data, which greatly affects the accuracy of model prediction.

[0004] In recent studies, large language models (LLMs) have revolutionized numerous fields. The potential knowledge acquired through pre-training on large corpora can be transferred to time series analysis, allowing the model to capture complex historical dependencies between data through carefully designed prompts, resulting in excellent performance and generalization ability of LLM-based prediction methods. Existing methods for applying LLMs to time series prediction can be broadly divided into two categories: feature engineering recoding and fine-tuning LLMs. Feature engineering recoding refers to extracting, transforming, and selecting useful information from raw data and converting it into a form that the model can handle more easily. This method is suitable for scenarios with small amounts of data, obvious features, and the need for rapid implementation. Through carefully designed features and prompts, good results can be achieved on limited data. Fine-tuning LLMs involves additional training on specific task data based on pre-trained models to optimize model performance on that task. This method is suitable for scenarios with large amounts of data and the need for models to automatically learn complex patterns. Through fine-tuning, the powerful capabilities of LLMs can be fully utilized to achieve better performance.

[0005] Current large language model frameworks in the industrial field often only input external text information (such as news reports, holiday events, etc.) or time series features, allowing the large language model to process in a single modality. While these methods can capture the correlation between external events and time series to some extent, they still have limited understanding of the internal mechanisms, structured knowledge, and industrial production processes of industrial systems. In particular, existing methods do not fully exploit and utilize the rich domain knowledge contained in industrial technical documents, device manuals, control logic explanations, and other texts. They also fail to allow the powerful reasoning capabilities of large language models to work synergistically on both technical document understanding and time series modeling. Therefore, neither feature engineering recoding-based methods nor fine-tuning LLMs can achieve high-precision prediction for complex industrial processes. SUMMARY

[0006] To achieve high-precision prediction of key process variables in industrial processes, the present invention proposes a key process variable prediction method based on the InduDK-LLM model, which comprises: Step one, for a specific industrial process, obtain its technical document data and collect historical data of industrial variables to construct an industrial time series dataset. Preprocess the industrial time series dataset to obtain a standard dataset, and divide the standard dataset into training, validation, and test sets according to the proportion; Step two, divide the industrial variables in the standard dataset obtained in step one into process variables and operation variables; Step three, construct a prediction model based on InduDK-LLM; Step four, using the data set classified by variables in step two to train the prediction model constructed in step three; Step five, using the prediction model trained in step four to predict the key process variables of the industry.

[0007] Optionally, the prediction model based on InduDK-LLM in step three includes an inverted embedding module, a static wavelet decomposition module, a process joint blocking module, a domain knowledge and operation variable embedding module, an interaction module, and a prediction module. The inverted embedding module is used to project the process variables and operation variables obtained by dividing in step two into high-dimensional embedding in the time dimension respectively, to obtain process variable inverted embedding sequences and operation variable inverted embedding sequences. The static wavelet decomposition module is used to decompose the process variable inverted embedding sequences to obtain corresponding coefficient sequences; the process joint blocking module is used to divide the coefficient sequences obtained by the static wavelet decomposition module into multiple patches according to channels. The domain knowledge and operation variable embedding module is used to convert the obtained technical document data into labeled numerical embedding, and generate industrial domain semantic enhanced industrial priori coding combined with the operation variable inverted embedding sequences. The interaction module is used to process the multiple patches of each component output by the process joint blocking module and the industrial priori coding output by the domain knowledge and operation variable embedding module, so as to establish the connection between the process variables and the operation variables, and output multi-scale process variable components. The prediction module includes an inverse static wavelet transform module and a head layer, the inverse static wavelet transform module is used to reconstruct the signal according to the multi-scale process variable components output by the interaction module, and the head layer obtains the prediction result through a Flatten layer and a linear layer.

[0008] Optionally, the static wavelet decomposition module includes padding operation and convolution operation; the input process variable inverted embedding sequence is first subjected to padding operation to ensure constant length, and then subjected to first layer decomposition to obtain low frequency part and high frequency part, which are respectively subjected to low pass filter and high pass filter for convolution, to extract approximation coefficients and detail coefficients , wherein the approximation coefficients are passed to the next layer for further decomposition and convolution operation to obtain higher level low frequency and high frequency features, until the preset decomposition number S is reached, and finally output the coefficient sequence composed of the highest layer approximation coefficient and each layer detail coefficient ; the hole rate of each layer convolution operation is expanded layer by layer to expand the receptive field.

[0009] Optionally, the process joint patch module performs a channel-independent patch operation on each channel of each coefficient in the coefficient sequence, and adds a learnable joint mark at the end of each channel patch to facilitate subsequent interaction between the process variables and the operation variables.

[0010] Optionally, the domain knowledge and operation variable embedding module includes a domain knowledge branch, an operation variable branch, and a pre-trained large model empowerment, wherein the domain knowledge branch converts the obtained technical document data into token numerical embedding, the operation variable branch uses a multi-head cross-attention mechanism to align the operation variable inverted embedding sequence with the pre-trained LLM discrete word embedding, and the pre-trained large model empowerment is used to parse the domain knowledge and operation variable inverted embedding sequence and output semantic-rich industrial prior encoding.

[0011] Optionally, the interaction module includes a multi-head self-attention layer, a multi-head cross-attention layer, a feedforward network layer, and a head layer, each sub-layer is connected by residual connection and layer normalization, and the interaction module takes the multiple patches of each component output by the process joint patch module and the industrial prior encoding output by the domain knowledge and operation variable embedding module as input to establish the relationship between the process variables and the operation variables.

[0012] Optionally, the specific industrial process includes industrial alcohol precipitation, industrial Chinese medicine extraction, industrial concentration, industrial sulfur recovery, industrial hydrocracking, and froth flotation process.

[0013] Optionally, the preprocessing in step one includes Z-Score normalization processing.

[0014] Optionally, the domain knowledge branch in the domain knowledge and operation variable embedding module converts the obtained technical document data into token numerical embedding, including: The text and symbol flow data in the technical document data are first subjected to a frozen pre-trained LLM tokenizer for tokenization, and then a pre-trained LLM embedder is used to convert the tokens into token numerical embedding.

[0015] The present application has the following advantages: The modeling of process variables and operating variables is explicitly separated, and domain knowledge and time series data are integrated to improve prediction accuracy. First, the domain knowledge and time series data of operating variables contained in the industrial technical documents are combined as the input of the pre-trained LLM, which reduces the workload of manual prompting engineering and generates industrial priori coding with semantic enhancement. Second, a learnable joint mark is introduced to establish the logical connection between process variables and operating variables. Finally, an interaction module is designed to model the interdependence and potential logical relationship between the two types of variables. The method realizes accurate prediction of key process variables in industry and provides a reference for actual industrial production. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0017] Figure 1 is a flowchart of the key process variable prediction method based on the InduDK-LLM model provided by the embodiments of the present application; Figure 2 is a prediction model structure diagram based on InduDK-LLM in the key process variable prediction method based on InduDK-LLM model provided by the embodiments of the present application; Figure 3 is a SWT module diagram of the prediction model based on InduDK-LLM in the key process variable prediction method based on InduDK-LLM model provided by the embodiments of the present application; Figure 4 is an ISWT module diagram of the prediction module in the prediction model based on InduDK-LLM in the key process variable prediction method based on InduDK-LLM model provided by the embodiments of the present application; Figure 5 is a domain knowledge and operating variable embedding module diagram of the prediction model based on InduDK-LLM in the key process variable prediction method based on InduDK-LLM model provided by the embodiments of the present application; Figure 6 is an interaction module diagram of the prediction model based on InduDK-LLM in the key process variable prediction method based on InduDK-LLM model provided by the embodiments of the present application; Figure 7is a schematic diagram of the prediction result of the temperature under the industrial alcohol precipitation tank by using the key process variable prediction method based on the InduDK-LLM model provided in the embodiment of the present application. Figure 8 is a schematic diagram of the prediction result of the temperature under the industrial alcohol precipitation tank by using the existing model DLinear. Figure 9 is a schematic diagram of the prediction result of the temperature under the industrial alcohol precipitation tank by using the existing model TimeMixer. Figure 10 is a schematic diagram of the prediction result of the temperature under the industrial alcohol precipitation tank by using the existing model PatchTST. Figure 11 is a schematic diagram of the prediction result of the temperature under the industrial alcohol precipitation tank by using the existing model Time-LLM. Figure 12 is a schematic diagram of the prediction result of the temperature under the industrial alcohol precipitation tank by using the existing model TimeCMA. Figure 13 is a schematic diagram of the error distribution violin of the prediction result of the temperature under the industrial alcohol precipitation tank by using the key process variable prediction method based on the InduDK-LLM model provided in the embodiment of the present application and the existing model. Figure 14 is a comparison diagram of the evaluation results of different prediction models. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical scheme and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0019] Embodiment one The embodiment provides a key process variable prediction method based on an InduDK-LLM model, and a flowchart of the method is shown in Figure 1 First, industrial production data is collected to construct a data set, and the data is cleaned and preprocessed, wherein the industrial production data refers to various variable data in an industrial process, which is collectively referred to as industrial variables; then the industrial variables are divided into process variables and operation variables; subsequently, a prediction model based on InduDK-LLM is constructed, and the prediction model is trained and evaluated according to the constructed data set; finally, the trained prediction model is used to predict the key process variables of the industry.

[0020] Specifically, it includes: Step one: collect the historical data of each industrial variable in a specific industrial process to construct an industrial time series data set, pre-process the industrial time series data set to obtain a standard data set, and divide the standard data set into a training set, a validation set and a test set according to a proportion.

[0021] Specific industrial processes include industrial alcohol precipitation, industrial Chinese medicine extraction, industrial concentration, industrial sulfur recovery, industrial hydrocracking, froth flotation process, etc. The corresponding industrial variables are different for different industrial processes; the preprocessing includes Z-Score normalization processing, subtracting the mean of each channel and dividing by the standard deviation, so that the measured values are standardized to the same range, and the specific expression is: ; wherein, and represent the data before and after normalization, and represent the mean and standard deviation of ; The ratio of training set, validation set and test set is 7:2:1.

[0022] Step two: divide the industrial variables in the standard data set obtained in step one into process variables and operation variables.

[0023] According to the properties of the variables, the variables in the industrial data are divided, and the process variables and operation variables are modeled by different ways, and an interaction module is designed to capture their mutual dependence and potential logical relationship; wherein, the process variables include temperature, concentration and other physical or chemical quantities that can be measured or observed, these variables reflect the state of the process, the operation variables include solvent addition rate, valve state, operation steps and other variables directly controlled and adjusted by the operator or control system.

[0024] Step three: build a prediction model based on InduDK-LLM.

[0025] As shown in Figure 2 , the prediction model based on InduDK-LLM includes an inverted embedding module, a static wavelet decomposition module (SWT), a process joint blocking module, a domain knowledge and operation variable embedding module, an interaction module and a prediction module.

[0026] The inverted embedding module includes a linear layer, which is used to process process variable sequences and operation variable sequences, specifically, the multivariate time series is regarded as a single label, and is projected into a high-dimensional embedding along the time dimension, thereby generating a time-aligned representation for each variable. This inverted embedding enables the model to more effectively capture complex cross-variable dynamics, which can be represented as: ; wherein, represents the number of process variables or operation variables, and represent the length of the original sequence before embedding and the hidden dimension of embedding, respectively.

[0027] As Figure 3 shown, the static wavelet decomposition module (SWT) includes padding operations and convolution operations; the input process variable inverted embedded sequence is first padded to ensure constant length, and then decomposed at the first layer to obtain low-frequency and high-frequency parts, which are respectively convolved through low-pass filter and high-pass filter to extract the approximation coefficient and the detail coefficient . In each layer of decomposition, the network does not perform downsampling like traditional discrete wavelets, but expands the receptive field through an increasing dilation rate (dilation = 1, 2, 4, …) to enable the convolution to capture longer time-domain features while maintaining time resolution.

[0028] The approximation coefficient obtained by decomposition is passed to the next layer for convolution operation, thereby obtaining higher-level low-frequency and high-frequency features, until the preset decomposition number is reached. The final output is formed by the highest-level approximation coefficient and the detail coefficients of each layer, forming a complete representation of the signal in multiple scales; its expression is: ; wherein, h and represent the learnable low-pass filter and high-pass filter, respectively, represents the number of decompositions. The process variable embedded sequence is decomposed into approximation coefficients and detail coefficients, wherein the approximation coefficients reflect long-term trends, and the detail coefficients capture short-term fluctuations or sudden changes; at the same time, the SWT does not downsample the coefficients obtained by decomposition, and its time shift invariance ensures that the decomposed coefficients are point-to-point aligned with the original time series on the time axis.

[0029] The process joint patching module is used for each coefficient sequence produced by the SWT; patch operations are performed on each channel of each coefficient, and transient events that are localized in time and causally related to the manipulated variable are amplified in their original local environment, making them more detectable for downstream local feature extractors and attention mechanisms; in addition, a learnable joint label is added at the end of each channel patch, providing a trainable bridge to facilitate subsequent interaction between process variables and manipulated variables; specifically, taking the th detail coefficient as an example, for each channel of the coefficient, the output after patching is wherein represents the number of channels, represents the length of each patch, represents the total number of patches, which is expressed as: wherein, represents the sliding step of the patching operation, represents the integer part of this number; a shape of is then introduced is learned jointly, the patch representation of each channel becomes is embedded into a higher dimension wherein is the hidden dimension.

[0030] As shown in Figure 5 , the domain knowledge and operation variable embedding module includes a domain knowledge branch, an operation variable branch, and a pre-trained LLM enabled three parts, which uses the powerful ability of the pre-trained LLM to generate semantic enhanced industrial domain semantic enhanced industrial priori coding, which helps the model to understand the industrial process and mine the internal logical relationship of the two kinds of industrial variables.

[0031] The domain knowledge branch converts the text data from the industrial technical document into a token numerical embedding, specifically, the text and symbolic flow data are first processed by the frozen pre-trained LLM tokenizer for word segmentation, and then the pre-trained LLM embedder is used to convert the token into a token numerical embedding , this branch includes the dataset summary, industrial equipment description, standard operation procedure and the explanation of related industrial variables into the model.

[0032] The operation variable branch uses a multi-head cross-attention mechanism to align the continuous value embedding of the operation variable with the pre-trained LLM discrete word embedding; in order to reduce the amount of calculation, avoid information redundancy and improve the efficiency and accuracy of the prediction task, the original pre-trained LLM word embedding is projected into an industrial domain related word embedding , which is specifically expressed as: ; wherein, and represent the size of and , respectively, is the input dimension of the LLM, and represent linear layer learnable parameters.

[0033] embed operation variables as queries, industrial-related word embeddings as keys and values, perform multi-head cross-attention mechanism between the two to achieve alignment purposes, the expression of the th single attention head is: ; wherein, , , respectively represent the query, key and value of the th single attention head, , , is a learnable weight matrix; then merge the results of all single attention heads to obtain: ; wherein, represents the number of attention heads; Finally, through a linear layer, project to the dimension of LLM, the expression is: ; wherein, and represent the learnable parameters of the linear layer.

[0034] As shown in Figure 6 , the interaction module includes a multi-head self-attention layer, a multi-head cross-attention layer, a feedforward network layer and a layer, each sub-layer is connected by residual connection and layer normalization, which is used to establish the relationship between process variables and operation variables, capture the internal complex time and causal relationship between the two, which includes a multi-head self-attention mechanism, a multi-head cross-attention mechanism, a feedforward network layer and a layer, each sub-layer is connected by residual connection and layer normalization, which stabilizes the training process and relieves the problem of gradient vanishing.

[0035] The input of the multi-head self-attention mechanism is the patch embedding result of the process variable , to model the dependency between the internal blocks of the process variable, and make the introduced learnable joint label obtain knowledge related to the process variable from other blocks; the output of a single attention head is: ; aggregate single attention heads, the expression of the output of this sub-layer is: ; in, Indicates the first The output of each attention head, Represents the number of attention heads, , This is the output weight matrix.

[0036] The output after introducing residual connections and layer normalization at this layer is: ; in, Representative level normalization, This represents a regularization technique to prevent model overfitting; The query for multi-head cross-attention mechanism is Learnable joint labeling of process variable embeddings Both the keys and values ​​are the outputs of the pre-trained LLM. This enables global dependency modeling of process and operational variables, aggregating a single attention head, and representing the output of this sublayer. The expression is: ; in, Indicates the first The output of a cross-attention head, Represents the number of attention heads, , This is the output weight matrix.

[0037] The output after introducing residual connections and layer normalization at this layer is: ; The feedforward network layer consists of two linear layers. A nonlinear transformation is introduced into these two linear layers through the GeLu activation function, and its output... The expression is: ; As head The layer takes input and the interactive module takes output. The expression is: ; in Indicates the first The output results of multiple scale components.

[0038] The prediction module includes inverse static wavelet transform ( ) and a layer; like Figure 4As shown, the inverse static wavelet decomposition module (ISWT) includes padding operations and convolution operations; the ISWT implements a process of gradually restoring the original scale signal by the multi-scale coefficients, which takes the approximation coefficients and the detail coefficients of the highest layer as the starting point, respectively fills them, and then obtains the reconstructed low-frequency and high-frequency parts through convolution by the reconstruction filters and . Adding and averaging the two parts obtains new approximation coefficients , and then the expansion rate is halved, and the result is input into the next convolution fusion together with the detail coefficients of the next layer; through such layer-by-layer iteration, the approximation and detail information are continuously fused, and finally the reconstructed signal is obtained; through inverse static wavelet transform, all multi-scale component features are reconstructed into original scale representation, where , , represents the decomposition times; the expression of the inverse transform is: ; wherein, and represent the learnable reconstruction low-pass filter and the reconstruction high-pass filter, respectively.

[0039] Then, through one layer, the normalized prediction result is obtained: ; wherein, represents stacking all multi-scale components; using the inverse normalization , the final prediction result is obtained.

[0040] Step four: using the data set classified by the variable in step two to train the prediction model constructed in step three.

[0041] By comparing the model prediction value with the true value, the prediction model parameters and weights are optimized, and when the mean square error loss of the validation set does not decrease for three consecutive iterations, the model stops training.

[0042] Step five: predicting the industrial key process variables according to the prediction model obtained in step four.

[0043] The prediction of the industrial key process variables is realized in the prediction model based on InduDK-LLM.

[0044] Embodiment two This embodiment provides a method for predicting key process variables based on the InduDK-LLM model, taking the industrial alcohol precipitation process as an example: using a preprocessed dataset, the prediction of key industrial process variables is achieved based on the InduDK-LLM prediction model.

[0045] This embodiment uses six industrial alcohol settling tank datasets to verify the performance of our proposed InduDK-LLM-based prediction model. These datasets were recorded in January 2024, sampled once per minute, for a total of approximately 44,480 time steps.

[0046] To verify the performance of the InduDK-LLM-based prediction model proposed in this invention, it is compared with DLinear, TimeMixer, PatchTST, Time-LLM, and TimeCMA in predicting the temperature in industrial alcohol precipitation tanks; among them... For DLinear, please refer to Zeng, A., Chen, M., Zhang, L.,&Xu, Q. (2023, June). Are transformers effective for time series forecasting?. In Proceedings of the AAAI conference on artificial intelligence (Vol. 37, No. 9, pp. 11121-11128); For TimeMixer, please refer to Wang, S., Wu, H., Shi, X., Hu, T., Luo, H., Ma, L., ... & Zhou, J. (2024). Timemixer: Decomposable multiscale mixing for timeseries forecasting. arXiv preprint arXiv:2405.14616 ; For PatchTST, please refer to Nie, Y. (2022). A Time Series is Worth 64Words: Long-term Forecasting with Transformers. arXiv preprint arXiv:2211.14730 ; Time-LLM can refer to Jin, M., Wang, S., Ma, L., Chu, Z., Zhang, J. Y., Shi, X.,... & Wen, Q. (2023). Time-llm: Time series forecasting by reprogramming large language models. arXiv preprint arXiv:2310.01728 ; TimeCMA can refer to Liu, C., Xu, Q., Miao, H., Yang, S., Zhang, L., Long, C.,... & Zhao, R. (2025, April). Timecma: Towards llm-empowered multivariate time series forecasting via cross-modality alignment. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (pp. 18780-18788). Proceedings of the AAAI Conference on Artificial Intelligence

[0047] The mean absolute error (MAE), mean squared error (MSE), and coefficient of determination (R2) are used as evaluation criteria; the smaller the values of MAE and MSE, the better the performance of the model; if the value of R2 is closer to 1, it indicates that the model has a better fitting effect and stronger ability to explain the observed values; the specific calculation formulas of these three indicators are as follows: ; ; ; ; wherein, , and represent the true value, the predicted value, and the average value of the true value, respectively; at the same time, denotes the sample size.

[0048] The results of the comparison are shown in Figure 14 , where the bolded results represent the best.

[0049] Overall, the prediction model based on InduDK-LLM proposed in the present application achieves the lowest MAE, the lowest MSE, and the highest R2 on each dataset. ​​, which indicates that its performance is always stable and has a significant advantage; for example, in dataset 1, the MAEs of other comparative models are 0.1664, 0.1684, 0.5206, 0.3539 and 0.1544 respectively, while the MAE of InduDK-LLM is 0.1056, which is reduced by 36.54%, 37.29%, 79.72%, 70.16% and 31.61% respectively; in dataset 2, the proposed model reduces the MAE value by 0.2825, 0.1674, 0.2099, 0.4579, 0.0365, and the average increase is ((56.76% + 43.75% + 49.38% + 68.03% + 14.50%) / 5) = 46.48%; at the same time, the proposed model obtains the highest values in datasets 5 and 6, which are 0.9901 and 0.9957 respectively, and from the perspective of , it also performs better. The proposed prediction model based on InduDK-LLM can effectively extract the hidden nonlinear dynamic relationship and logical relationship in these variables by explicitly dividing complex industrial alcohol precipitation variables, integrating domain knowledge and operating variables, and using the powerful reasoning ability of the pre-trained large language model to enhance the understanding of the industrial process, thereby significantly improving the prediction accuracy.

[0050] Figure 7 、 Figure 8 、 Figure 9 、 Figure 10 、 Figure 11 、 Figure 12 Figures are the prediction results of the industrial alcohol precipitation tank temperature of InduDK-LLM, DLinear, TimeMixer, PatchTST, Time-LLM and TimeCMA respectively; as can be seen from the figure, the prediction curve of the prediction model based on InduDK-LLM is highly consistent with the actual temperature curve, and its performance at temperature stable and turning point, as well as temperature peak and trough value is better than that of the other five comparative models; Figure 13 Fig. 6 is a violin plot of the absolute error distribution of the prediction model based on InduDK-LLM provided by the present application and the existing models on dataset 6, the prediction model based on InduDK-LLM presents a narrow and compact chord diagram, the median is closer to zero, and the interquartile range is shorter, which indicates that most of the predicted values are close to the true value, and the situation of large error is rare; it is proved that the prediction model based on InduDK-LLM proposed by the present application can realize accurate prediction.

[0051] Some steps in the embodiments of the present application can be realized by software, and the corresponding software program can be stored in a readable storage medium, such as an optical disc or a hard disk.

[0052] The above description is only the preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for predicting key process variables based on InduDK-LLM model, characterized in that, The method comprises: Step one, for a specific industrial process, obtain its technical document data, collect historical data of industrial variables to construct an industrial time series dataset, preprocess the industrial time series dataset to obtain a standard dataset, and divide the standard dataset into a training set, a validation set and a test set according to a proportion; Step two, divide the industrial variables in the standard dataset obtained in step one into process variables and operation variables; Step three, construct a prediction model based on InduDK-LLM; Step four, train the prediction model constructed in step three using the dataset classified by variables in step two; Step five, use the prediction model trained in step four to predict the key process variables of the industry; The prediction model based on InduDK-LLM in step three comprises an inverted embedding module, a static wavelet decomposition module, a process joint blocking module, a domain knowledge and operation variable embedding module, an interaction module and a prediction module; The inverted embedding module is used to project the process variables and operation variables divided in step two into high-dimensional embeddings in the time dimension respectively to obtain process variable inverted embedding sequences and operation variable inverted embedding sequences; The static wavelet decomposition module is used to decompose the process variable inverted embedding sequences to obtain corresponding coefficient sequences; the process joint blocking module is used to divide the coefficient sequences decomposed by the static wavelet decomposition module into multiple patches according to channels; The domain knowledge and operation variable embedding module is used to convert the obtained technical document data into labeled numerical embedding, and generate industrial domain semantic enhanced industrial priori coding in combination with the operation variable inverted embedding sequences; The interaction module is used to process the multiple patches of each component output by the process joint blocking module and the industrial priori coding output by the domain knowledge and operation variable embedding module, so as to establish the connection between the process variables and the operation variables, and output multi-scale process variable components; The prediction module comprises an inverse static wavelet transform module and a head layer, the inverse static wavelet transform module is used to reconstruct the signal according to the multi-scale process variable components output by the interaction module, and the head layer obtains the prediction result through a Flatten layer and a linear layer.

2. The method of claim 1, wherein, The static wavelet decomposition module comprises a padding operation and a convolution operation; the input process variable inverted embedded sequence is firstly subjected to the padding operation to ensure the length unchanged, and then is subjected to the first layer decomposition to obtain a low frequency part and a high frequency part , which are subjected to low pass filtering and high pass filtering , respectively, to perform convolution to extract an approximation coefficient and a detail coefficient, wherein the approximation coefficient is transmitted to the next layer to continue the decomposition and the convolution operation to obtain higher layer low frequency and high frequency features, until a preset decomposition number S is reached, and finally an output coefficient sequence composed of the highest layer approximation coefficient and each layer detail coefficient is obtained; the hole rate of each layer convolution operation is expanded layer by layer to expand the receptive field.

3. The method of claim 2, wherein, The process joint blocking module performs channel-independent patch operations on each channel of each coefficient in the coefficient sequence, and adds a learnable joint label at the end of each channel patch to facilitate subsequent interaction between the process variables and the operation variables.

4. The method of claim 3, wherein, The domain knowledge and operation variable embedding module comprises a domain knowledge branch, an operation variable branch and a pre-trained large model empowerment part, wherein the domain knowledge branch converts the obtained technical document data into labeled numerical embedding, the operation variable branch uses a multi-head cross attention mechanism to align the operation variable inverted embedding sequences with the pre-trained LLM discrete word embedding, and the pre-trained large model empowerment part is used to analyze the domain knowledge and operation variable inverted embedding sequences and output semantic-rich industrial priori coding.

5. The method of claim 4, wherein, The interaction module comprises a multi-head self-attention layer, a multi-head cross-attention layer, a feedforward network layer and a head layer, each sub-layer is connected through residual connection and layer normalization, and each component of the process joint block module and the industrial prior encoding output by the domain knowledge and operation variable embedding module are taken as inputs of the interaction module to establish the connection between the process variables and the operation variables.

6. The method of claim 5, wherein, The specific industrial processes include industrial alcohol precipitation, industrial Chinese medicine extraction, industrial concentration, industrial sulfur recovery, industrial hydrocracking and froth flotation process.

7. The method of claim 6, wherein, The preprocessing in step one comprises Z-Score normalization processing.

8. The method of claim 7, wherein, In the domain knowledge branch of the domain knowledge and operation variable embedding module, the obtained technical document data is converted into token numerical embedding, comprising: The text and symbol flow data in the technical document data are subjected to tokenization processing through a frozen pre-trained LLM tokenizer, and then the tokens are converted into token numerical embedding through a pre-trained LLM embedder.

Citation Information

Patent Citations

  • Building big data acquisition and intelligent settlement monitoring method and Internet of Things system

    CN117128924A

  • Prediction of process variables by simulation based on partially only measurable initial states

    CN117581169A

  • Interactive evolution graph intelligent design method based on GNN and LLM

    CN119337928A

  • Industrial fault diagnosis method and system based on intelligent causal correction

    CN120217262A

  • Time sequence prediction method and device based on wavelet transform and large language model

    CN120724152A

Cited By

  • Variable working condition industrial temperature prediction method based on KWNet model

    CN121682449A