Industrial quality variable prediction method based on target self-adaption-hierarchical fusion Transform

Through the target adaptive-level fusion Transformer model, dynamically adjusting feature weights and adapting to the delay characteristics of industrial processes, the problem of lack of target task guidance and feature loss in the existing methods is solved, and high-accurate industrial quality variable prediction is achieved.

CN120278323APending Publication Date: 2025-07-08HANGZHOU NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510357383.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing industrial quality variable prediction methods lack the guidance of feature variables related to the target task, resulting in low prediction accuracy, and increasing the number of model layers will lead to loss of feature information, affecting the prediction accuracy.

Method used

The target adaptive-level fusion Transformer (TAIF-TF) model is adopted, and the target adaptive attention layer and the hierarchical fusion attention layer are dynamically adjusted, and the key quality variables are predicted in combination with the linear regression layer to adapt to the time delay characteristics and data distribution changes of different industrial processes.

Benefits of technology

It improves the prediction accuracy and robustness of the model, can adapt to complex industrial processes, meets the efficient and real-time requirements of modern industrial processes for data processing, and significantly improves the prediction effect of key quality variables.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278323A_ABST
    Figure CN120278323A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial quality variable prediction method based on a target adaptive-hierarchical fusion Transform. In order to solve the problem that a key quality variable is difficult to detect in real time in real industrial production, the key quality variable is predicted by constructing a TAIF-TF model; collecting process data in real industrial production to train the TAIF-TF model; each layer of encoder in the model has a target adaptive attention layer, a multi-head attention layer and a mask multi-head attention layer. The output of the current encoder layer is used as the input of the next encoder layer, the output of each encoder layer passes through the hierarchical fusion attention layer, different hierarchical attention weights are obtained, and the final output is obtained through weighting; and finally, calculating a predicted key quality variable through a linear regression layer. According to the method, the complex relationship between the key quality variable and the auxiliary variable can be more accurately captured, the accuracy of the prediction result is remarkably improved, and the adaptability of the model to the complex industrial process is enhanced; the method is suitable for various industrial scenes; and large-scale or high-dimensional data can be efficiently processed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of industrial processes, and in particular, to an industrial quality variable prediction method based on Target Adaptive-Interlevel Fusion Attention Transformer (TAIF-TF). Background Art

[0002] In actual industrial manufacturing engineering, the control of key quality variables is very important, which is the core key of the industrial quality prediction task, because key quality variables often determine the success or failure of industrial production. To grasp the quality of industrial products, it is necessary to monitor various quality variables in the production process in real time. However, due to various reasons such as complex production processes, harsh industrial environments, difficult machine operations, dangerous production processes, expensive operating instruments, and time delays in industrial processes, it is difficult to directly obtain data through hardware sensors for monitoring some quality variables in the production process. Therefore, soft sensing technology is required to build a model to predict key quality variables.

[0003] Soft sensing is based on easily detectable variables, and uses the relationship between easily detectable variables and difficult-to-detect variables to make predictions through various mathematical relationships. Soft sensing modeling methods can be mainly classified into the following three types: mechanism-based modeling, data-driven modeling, and hybrid modeling. Deep learning in data-driven modeling methods is a powerful data-driven method for dealing with nonlinear problems. It can effectively mine useful information from massive process data and apply it to the field of industrial process soft sensing modeling.

[0004] In recent years, many data-driven soft sensing modeling methods have been applied to the field of industrial production. However, with the increasing complexity and difficulty of control in industrial manufacturing in recent years, many traditional methods still face many problems. For example, existing modeling methods often ignore that different industrial processes have different industrial requirements and target tasks, and the prediction of key quality variables is closely related to the target task. Therefore, existing modeling methods lack the guidance of feature variables related to the target task, resulting in low prediction accuracy. And in order to adapt to large-scale or high-dimensional industrial data, existing modeling methods often increase the number of model layers. However, increasing the number of model layers will cause another problem, that is, the features extracted by the model will be lost during transmission, resulting in a decrease in the prediction accuracy of the model. Summary of the Invention

[0005] Aiming at the problems of lack of guidance for key quality variables related to the target task and loss of features extracted during transmission in the prior art, the present invention proposes an industrial quality variable prediction method based on Target Adaptive-Interlevel Fusion Transformer.

[0006] The present invention specifically includes the following steps: Step 1: Data collection. Key quality variables y and auxiliary variables x in the previous working process are collected. According to the collection time sequence, a training set, a validation set, and a test set are separated. The data pairs collected earlier are selected as the training set, the data pairs collected in the middle are used as the validation set, and the data collected later are used as the test set. The key quality variables y and auxiliary variables x in the data set are normalized so that the result values are mapped between 0 and 1.

[0007] Step 2: Construct a target adaptive - hierarchical fusion Transformer industrial quality variable prediction model TAIF - TF. The TAIF - TF model includes k layers of encoders, a hierarchical fusion layer, and a linear regression layer, where k is a set value.

[0008] Each layer of the encoder has a target adaptive attention layer, a multi - head attention layer, and a masked multi - head attention layer.

[0009] The auxiliary variables input to the encoder are respectively subjected to feature extraction through the target adaptive attention layer and the multi - head attention layer. The target adaptive attention layer calculates the target adaptive attention value, assigns a weight to each feature variable in the auxiliary variables input to the encoder, and identifies the feature variables with a relatively large relationship with the guiding quality variable of the input encoder. The auxiliary variables after feature extraction by the target adaptive attention layer are input to the masked multi - head attention layer for feature extraction. The auxiliary variables after feature extraction by the target adaptive attention layer, the multi - head attention layer, and the masked multi - head attention layer are subjected to the first residual connection and the first layer normalization; then, through the non - linear transformation and feature enhancement processing of the feed - forward network layer; the outputs of the first layer normalization and the feed - forward network layer are subjected to the second residual connection and the second layer normalization to obtain the output of the current encoder layer .

[0010] The output of the current encoder layer is used as the input of the next layer of the encoder, and the above operations are repeated for each encoder.

[0011] The output of each layer of the encoder passes through the hierarchical fusion attention layer to obtain different hierarchical attention weights, and the weighted result is the final output.

[0012] Finally, based on the final output, the predicted key quality variable is calculated through the linear regression layer.

[0013] Step 3: According to the time - delay characteristics of the industrial process, the historical key quality variable y detected a time stamps ago is used as the guiding quality variable for the current prediction time stamp . The length a of the sliding window is dynamically adjusted according to the time - delay characteristics of the industrial process.

[0014] To optimize the training model, the key quality variable y at the current timestamp in the training set is used as the guiding quality variable .

[0015] Set the model learning rate, the number of model training iterations, and the batch size. According to the set batch size, perform batch processing on the guiding quality variable and the auxiliary variable x in the training set. Use the auxiliary variable x in the training set and the guiding quality variable after dimension expansion to train the model. In each forward propagation process, calculate the error between the predicted value and the true value through the loss function. The loss value reflects the accuracy of the current prediction of the model. After the loss value is calculated, calculate the gradient of the loss function with respect to the model parameters through the backpropagation algorithm

[0016] The validation set is used to evaluate the generalization ability of the model, monitor the training process, and perform hyperparameter tuning. Training and validation are alternated. After each training cycle ends, the model will perform a forward propagation using the validation set and calculate the loss value on the validation set to determine whether the model is overfitting or underfitting

[0017] Step 4: Input the auxiliary variable of the test set and the guiding feature variable corresponding to this auxiliary variable into the trained TAIF-TF model to obtain the predicted key quality variable value. Combine the predicted key quality variable and the true key quality variable, and use regression evaluation metrics to evaluate the model

[0018] Step 5: Input any auxiliary variable in the industrial production process and the corresponding guiding quality variable into the TAIF-TF model evaluated in Step 4, and output the predicted key quality variable at the current moment

[0019] Further, the processing process of each group of guiding quality variables and the corresponding auxiliary variable x in Step 3 is as follows Perform dimension expansion on the guiding quality variable in the training set to make the dimension of the guiding quality variable consistent with the dimension of the corresponding auxiliary variable x. Perform dense layer embedding and positional encoding processing on the guiding quality variable with consistent dimensions and the auxiliary variable to obtain the auxiliary variable entering the encoder and the guiding quality variable .

[0020] The target adaptive attention layer calculates the target adaptive attention value , assigns a weight to each feature variable in the auxiliary variable , and identifies the relationship with the guiding quality variable The feature variables with greater relevance are calculated as follows: ; In the formula is the target adaptive attention value of the d-th feature variable, indicating the correlation between the d-th feature variable and the guidance quality variable ; represents the parameter matrix to be learned and updated; GELU represents the Gaussian error linear unit function in the activation function; is an intermediate variable, , ; In the formula, represents the output of the previous layer encoder. When i = 1, is the auxiliary variable entering the encoder ; represents the d-th feature variable in the output feature vector of the previous layer encoder. , , are all parameter matrices to be learned and updated. During the iterative training of the model, the parameter matrices are updated by calculating the gradients through the backpropagation algorithm.

[0021] Then all the obtained target adaptive attention values are normalized, and the normalized target adaptive attention value ; In the formula, is the normalized target adaptive attention value.

[0022] Finally, attention weighting is performed to obtain the feature vector after feature weight sharing by the target adaptive attention layer : .

[0023] The auxiliary variable entering the encoder is input to the multi-head attention layer for feature extraction while the feature weight sharing is performed in the input target adaptive attention layer. The multi-head attention layer outputs the feature vector ; The feature vector output by the target adaptive attention layer is input to the masked multi-head attention layer for feature extraction, and the masked multi-head attention layer outputs the feature vector ; The feature vector output by the multi-head attention layer, the feature vector output by the target adaptive attention layer, and the feature vector output by the masked multi-head attention layer are subjected to residual connection and layer normalization to obtain the intermediate variable : , where LN represents layer normalization processing.

[0024] Then through the feedforward network layer and layer normalization, the output of the current encoder layer is obtained : ; Where FFN represents feed-forward network layer processing.

[0025] The output of each encoder layer passes through the hierarchical fusion attention layer to obtain different hierarchical attention weights, and the weighted feature vector is obtained after feature decomposition through the hierarchical fusion attention layer. : ; In the formula, represents the normalized level attention value, ; is the layer-level attention value, which indicates the correlation between the output of the i-th layer in the k-layer encoder and the key quality variable. ; Represents the parameter matrix to be learned and updated; GELU represents the Gaussian error linear unit function in the activation function; and represents the intermediate variable, ; ; Represents the output of the previous encoder. When i=1, is the feature vector output by the target adaptive attention layer in the first encoder layer; is the output of the current encoder layer. and They are all parameter matrices to be learned and updated. During the iterative training of the model, the gradient is calculated through the back propagation algorithm to update the parameter matrix.

[0026] After extracting feature information from multiple layers of key attention layers, the feature vector is obtained , the feature vector Input the linear regression layer to calculate the predicted value of the key quality variable , .

[0027] Preferably, a broadcast mechanism is used to expand the dimension of the guidance quality variable g in the training set so that the dimension of the guidance quality variable g is consistent with the dimension of the corresponding auxiliary variable x.

[0028] Preferably, in the training process described in step three, the model is trained using the Adam gradient descent optimization algorithm and the MSE loss function.

[0029] The beneficial effects of the present invention are: The TAIF-TF model constructed by the present invention uses the key quality variables of the task objective through the target adaptive attention mechanism to guide the model to extract features in the auxiliary variables, obtain the correlation between them, so as to ensure that the input features with higher relevance to the target features can be retained, thereby improving the prediction accuracy of the model; through the hierarchical fusion attention mechanism, an inter-level attention weight is assigned to the output of each encoder layer, which can retain the feature information extracted by the model when increasing the number of layers of the model under the requirements of processing large-scale data or high-dimensional data in the model, so as to improve the prediction accuracy of the key quality variables in actual industrial production.

[0030] It can dynamically adjust the weights of input features according to the task objective, avoid the interference of irrelevant features on the prediction results, and also solve the problem of feature information loss in deep learning models. It can capture the complex relationship between key quality variables and auxiliary variables more accurately, significantly improve the accuracy of prediction results, and enhance the adaptability of the model to complex industrial processes; and it can adapt to the time-delay characteristics and data distribution changes of different industrial processes, has strong robustness and generalization ability, and is applicable to a variety of industrial scenarios; it can efficiently process large-scale or high-dimensional data, meeting the requirements of high efficiency and real-time of data processing in modern industrial processes. Brief Description of the Drawings

[0031] Figure 1 is a schematic structural diagram of the TAIF-TF model of the present invention; Figure 2 is a schematic structural diagram of the target adaptive attention mechanism in the TAIF-TF model of the present invention; Figure 3 is a schematic structural diagram of the hierarchical fusion attention mechanism in the TAIF-TF model of the present invention; Figure 4 is a flow chart of model training, verification and testing of the present invention; Figure 5 is a working flow chart of a reforming furnace in an application example of the present invention.

[0032] Figure 6 is a scatter plot of the predicted values and true values of each model in the application example of the present invention. Detailed Embodiments

[0033] The present invention will be described in detail below according to the drawings and preferred embodiments. The purpose and effects of the present invention will become more clear. The present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0034] An industrial quality variable prediction method based on target adaptive-hierarchical fusion Transformer specifically includes the following steps: Step 1: Data collection. Collect the key quality variable y and the auxiliary variable x during the previous work process. According to the collection time series, divide them into a training set, a validation set, and a test set. Select the data pairs collected earlier as the training set, the data pairs collected in the middle as the validation set, and the data collected later as the test set. In this embodiment, the first 60% of the data is used as the training set, the middle 20% of the data is used as the validation set, and the last 20% of the data is used as the test set.

[0035] Normalize the key quality variable y and the auxiliary variable x in the dataset so that the result values are mapped between 0 and 1.

[0036] Step 2: Construct a Target Adaptive-Hierarchical Fusion Transformer (TAIF-TF) industrial quality variable prediction model. As Figure 1 shown, the TAIF-TF model includes k layers of encoders, a hierarchical fusion layer, and a linear regression layer, where k is a set value.

[0037] Each layer of the encoder has a target adaptive attention layer, a multi-head attention layer, and a masked multi-head attention layer.

[0038] The auxiliary variables input to the encoder are respectively subjected to feature extraction through the target adaptive attention layer and the multi-head attention layer. As Figure 2 shown, the target adaptive attention layer calculates the target adaptive attention value, assigns a weight to each feature variable in the auxiliary variables input to the encoder, and identifies the feature variables that have a greater relationship with the guiding quality variable of the input encoder. The auxiliary variables after feature extraction by the target adaptive attention layer are input to the masked multi-head attention layer for feature extraction. The auxiliary variables after feature extraction by the target adaptive attention layer, the multi-head attention layer, and the masked multi-head attention layer are subjected to the first residual connection and the first layer normalization; then, through the non-linear transformation and feature enhancement processing of the feed-forward network layer; the outputs of the first layer normalization and the feed-forward network layer are subjected to the second residual connection and the second layer normalization to obtain the output of the current encoder layer. .

[0039] Use the output of the current encoder layer as the input to the next layer of the encoder, and repeat the above operations for each encoder.

[0040] The multi-head attention layer operates multiple independent attention heads in parallel to overcome the problem of insufficient feature vector representation caused by a single attention head; and it also expands the model's ability to focus on different subspaces at different positions, thus diversifying its feature expression ability. The masked multi-head attention layer introduces a masking mechanism on the basis of the multi-head attention layer, enabling the model to rely only on known information when predicting key quality variables and avoiding leaking future information to the model. According to the attention calculation and weight assignment of multiple key mechanisms, the effectiveness of feature extraction is enhanced.

[0041] As Figure 3 shown, the output of each layer of the encoder passes through a hierarchical fusion attention layer to obtain different hierarchical attention weights and is weighted to obtain the final output.

[0042] Finally, based on the final output, the predicted key quality variables are calculated through a linear regression layer.

[0043] Step 3: According to the time-delay characteristics of the industrial process, use the historical key quality variable y detected a time stamps ago as the guiding quality variable for the current prediction time stamp . The length a of the sliding window is dynamically adjusted according to the time-delay characteristics of the industrial process.

[0044] To optimize the training of the model, use the key quality variable y at the current time stamp in the training set as the guiding quality variable .

[0045] As Figure 4 shown, use the auxiliary variable x in the training set and the guiding quality variable after dimensional expansion to train the model. During training, use the Adam gradient descent optimization algorithm and use Mean Squared Error (MSE) as the loss function. In each forward propagation process, calculate the error between the predicted value and the true value through the loss function. The loss value reflects the accuracy of the current prediction of the model. After the loss value is calculated, calculate the gradient of the loss function with respect to the model parameters through the backpropagation algorithm. The gradient value propagates layer by layer from the output layer. Through the chain rule, the gradient of the loss with respect to the parameters of each layer is calculated layer by layer. Through the gradient descent algorithm, the model parameters are updated according to the calculated gradient.

[0046] Loss function ; where N represents the number of samples in the training set, represents the key quality variable of the i-th training sample, represents the predicted value of the i-th training sample under the TAIF-TF model.

[0047] The lower the MSE value, the smaller the deviation between the predicted value of the model and the true value, and the better the prediction performance of the model.

[0048] The validation set is used to evaluate the generalization ability of the model, monitor the training process, and perform hyperparameter tuning. Training and validation are carried out alternately. After each training epoch, the model performs a forward pass using the validation set to calculate the loss value on the validation set, which is used to determine whether the model is overfitting or underfitting.

[0049] Set the model learning rate, the number of model training epochs, and the batch size. According to the set batch processing, for the guidance quality variable in the training set and the auxiliary variable x are processed in batches. Each group of guidance quality variables and the corresponding auxiliary variable x are processed as follows: In this embodiment, the broadcast mechanism is used to expand the dimension of the guidance quality variable in the training set so that the dimension of the guidance quality variable is consistent with the dimension of the corresponding auxiliary variable x. For the guidance quality variable with consistent dimensions and the auxiliary variable perform dense layer embedding and position encoding processing to obtain the auxiliary variable entering the encoder and the guidance quality variable .

[0050] As Figure 2 shown, the target adaptive attention layer calculates the target adaptive attention value to assign a weight to each feature variable in the auxiliary variable and identify the feature variables that have a greater relationship with the guidance quality variable . The calculation process is as follows: ; In the formula is the target adaptive attention value of the d-th feature variable, indicating the correlation between the d-th feature variable and the guidance quality variable ; represents the parameter matrix to be learned and updated; GELU represents the Gaussian error linear unit function in the activation function; is an intermediate variable, , ; In the formula, represents the output of the previous layer of the encoder. When i = 1, is the auxiliary variable entering the encoder; represents the d-th feature variable in the output feature vector of the previous layer of the encoder. , , are all parameter matrices to be learned and updated. During the iterative training of the model, the parameter matrices are updated by calculating the gradients through the backpropagation algorithm.

[0051] Then, all the obtained target adaptive attention values are normalized, and the normalized target adaptive attention values ; in the formula, is the normalized target adaptive attention value, and this formula ensures that the sum of all target adaptive attention values is 1.

[0052] Finally, attention weighting is performed to obtain the feature vector after feature weight sharing by the target adaptive attention layer : .

[0053] Auxiliary variable input to the encoder While performing feature weight sharing in the input target adaptive attention layer, it is also input to the multi-head attention layer for feature extraction, and the multi-head attention layer outputs the feature vector ; the feature vector output by the target adaptive attention layer is input to the masked multi-head attention layer for feature extraction, and the masked multi-head attention layer outputs the feature vector ; the feature vector output by the multi-head attention layer, the feature vector output by the target adaptive attention layer, and the feature vector output by the masked multi-head attention layer are subjected to residual connection and layer normalization to obtain the intermediate variable , where LN represents layer normalization processing.

[0054] Then, through the feed-forward network layer and layer normalization, the output of the current encoder layer is obtained: ; in the formula, FFN represents feed-forward network layer processing.

[0055] As Figure 3 shown, the output of each layer of the encoder passes through the hierarchical fusion attention layer to obtain different hierarchical attention weights, and after weighting, the feature vector obtained after feature weight sharing by the hierarchical fusion attention layer is obtained: ; in the formula, represents the normalized hierarchical attention value, ; is the hierarchical attention value, representing the correlation between the output of the i-th layer in the k-th layer of the encoder and the key quality variable, and the sum of all hierarchical attention values is 1. ; represents the parameter matrix to be learned and updated; GELU represents the Gaussian error linear unit function in the activation function; and represents an intermediate variable, ; ; represents the output of the previous layer of the encoder. When i = 1, is the feature vector output by the target adaptive attention layer in the first layer of the encoder; is the output of the current encoder layer. In the formula, and are both parameter matrices to be learned and updated. During the iterative training of the model, the parameter matrices are updated by calculating gradients through the backpropagation algorithm.

[0056] After extracting the feature information through multiple layers of key attention layers, the feature vector is obtained. The feature vector is input into the linear regression layer, and the predicted value of the key quality variable is calculated, .

[0057] Step Four: As shown in Figure 4 , the auxiliary variable of the test set and the corresponding guiding feature variable are input into the trained TAIF-TF model to obtain the predicted value of the key quality variable. Combining the predicted key quality variable and the true key quality variable, the model is evaluated using regression evaluation metrics.

[0058] Step Five: Any auxiliary variable in the industrial production process and the corresponding guiding quality variable are input into the TAIF-TF model evaluated in Step Four, and the predicted key quality variable at the current moment is output.

[0059] Application Example Taking the prediction of industrial quality variables in the hydrogen production unit during ammonia synthesis as an example, the performance of the industrial quality variable prediction method described in the present invention is detected.

[0060] The dataset of this application example is extracted from the hydrogen production unit during ammonia synthesis. NH3 produced during ammonia synthesis is usually the main material in the urea synthesis process. In the ammonia synthesis process, hydrogen is one of the important raw materials and is usually converted from the raw material methane. The foregoing process of the ammonia synthesis process is the methane conversion unit, which includes a pre-reformer, a primary reformer, and a secondary reformer. According to the process flow design, the reforming reaction mainly takes place in the primary reformer (as shown in Figure 5 ). Therefore, it is particularly important to optimize the process control of this device, which can effectively improve the production and purity of hydrogen.

[0061] The reaction temperature is a key factor in ensuring hydrogen production in the primary reformer. The combustion situation should be monitored in a timely manner to keep the temperature stable at a certain level. To stabilize the combustion conditions, one of the key measures is to control the oxygen content in the furnace within a set range. In the actual process, the oxygen content is measured by an expensive mass spectrometer, but due to the fast combustion speed, the oxygen content in the furnace changes frequently. Therefore, in this application example, the oxygen content at the furnace top is selected as the key quality variable y to be predicted.

[0062] To build a soft-sensing model, 13 auxiliary variables that affect the oxygen content (see Table 1), including temperature, pressure, and flow rate, etc., are selected. The process instruments are Figure 5 marked with gray boxes in the figure. The light gray blocks represent the auxiliary variables, while the dark gray block represents the key variable, the oxygen content.

[0063] Table 1 Auxiliary variables in the primary reformer process and the oxygen content at the furnace top Label Description FR03001.PV Natural gas fuel flowing into 03B001 FR03002.PV Exhaust gas fuel flowing into 03B001 PC03002.PV Exhaust gas fuel pressure at the outlet of 03E005 PC03007.PV Furnace flue gas pressure at the outlet of 03B001 TI03001.PV Exhaust gas fuel temperature at the outlet of 03E005 TI03009.PV Natural gas fuel temperature at the outlet of 03B002E06 TI03013.PV Furnace flue gas temperature above the left of 03B001 TI03014.PV Furnace flue gas temperature above the right of 03B001 TR03012.PV Gas feedstock temperature at the inlet of 03B001 TR03015.PV Mixed furnace flue gas temperature at the top of 03B001 TR03016.PV Reformed gas temperature at the left outlet of 03B001 TR03017.PV Reformed gas temperature at the right outlet of 3B001 TR03020.PV Reformed gas temperature at the outlet of 03B001 AR03001.PV Oxygen content at the furnace top A total of 2500 samples are collected from the actual primary reformer process. These data are divided into a training set containing the first 60% of the entire dataset, a validation set containing the middle 20% of the dataset, and a test set containing the last 20% according to the time sequence. In this experiment, the currently detected oxygen content at the furnace top (according to the time-delay characteristics of the ammonia synthesis process, the currently detected is the oxygen content 3 time stamps ago, i.e., a = 3) is used as the guiding feature variable, and historical data is dynamically used for prediction.

[0064] The learning rate of the TAIF-TF model is set to 0.0005, the number of model training iterations (times) is 50, and the batch size is 32.

[0065] Meanwhile, to demonstrate the performance of the TAIF-TF model described in the present invention, this model is compared with five network models, namely the Long- and Short-term Time-series network (hereinafter referred to as LSTnet), the Long Short-Term Memory network (hereinafter referred to as LSTM), the Stacked Autoencoder (hereinafter referred to as SAE), the Support Vector Regression (hereinafter referred to as SVR), and the Principal Component Regression (hereinafter referred to as PCR), for quality prediction. In PCR, the number of features after dimensionality reduction is the same as the number of features in the dataset. The radial basis function (RBF) is used as the kernel function of SVR. The hidden layer sizes of LSTM and LSTnet are both set to 64, and the number of training iterations and the batch size are the same as those of TAIF-TF.

[0066] The Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), and Coefficient of Determination (R-Square, R 2 ) are used for evaluation. The smaller the values of MAE and RMSE, the higher the prediction accuracy of the model; the value of R 2 ranges from , and the closer it is to 1, the higher the prediction accuracy of the model.

[0067] After training the above models with the same training set and validation set, the above models are evaluated using the same test set (the evaluation results are shown in Table 2), and at the same time, a scatter plot of the predicted values and true values of each model as shown in Figure 6 is obtained.

[0068] Table 2 Evaluation Results of Six Models

[0069] As can be seen from Table 2, the TAIF-TF model of the present invention has the best prediction effect, with MAE being 0.042, RMSE being 0.056, and R 2 being 0.899, all reaching the optimal. It can also be seen from Figure 6 that the data points of the TAIF-TF model are closest to the diagonal line, indicating that the predicted values of the TAIF-TF model are closest to the true values and the prediction accuracy is the highest.

[0070] Those of ordinary skill in the art can understand that the above are only preferred examples of the invention and are not used to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, for those skilled in the art, they can still modify the technical solutions described in the foregoing examples, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, etc. made within the spirit and principle of the invention shall be included within the protection scope of the invention.

Claims

1. An industrial quality variable prediction method based on target adaptive - hierarchical fusion Transformer, characterized in that: Specifically, it includes the following steps: Step 1: Data collection. Collect the key quality variable y and the auxiliary variable x during the previous work process. According to the collection time series, divide them into a training set, a validation set, and a test set. Select the data pairs collected earlier as the training set, the data pairs collected in the middle as the validation set, and the data collected later as the test set. Normalize the key quality variable y and the auxiliary variable x in the data set so that the result values are mapped between 0 and 1; Step 2: Construct the target adaptive - hierarchical fusion Transformer industrial quality variable prediction model TAIF - TF. The TAIF - TF model includes k layers of encoders, a hierarchical fusion layer, and a linear regression layer, where k is a set value; Each layer of the encoder has a target adaptive attention layer, a multi - head attention layer, and a masked multi - head attention layer; The auxiliary variables input to the encoder are respectively subjected to feature extraction through the target adaptive attention layer and the multi - head attention layer; The target adaptive attention layer calculates the target adaptive attention value, assigns a weight to each feature variable in the auxiliary variables input to the encoder, and identifies the feature variables with a relatively large relationship with the guiding quality variable of the input encoder; Input the auxiliary variables after feature extraction by the target adaptive attention layer into the masked multi - head attention layer for feature extraction; Perform the first residual connection and the first layer normalization on the auxiliary variables after feature extraction by the target adaptive attention layer, the multi - head attention layer, and the masked multi - head attention layer; Then perform non - linear transformation and feature enhancement processing through the feed - forward network layer; Perform a second residual connection and a second layer normalization on the output of the first layer normalization and the feed-forward network layer to obtain the output of the current encoder layer ; Take the output of the current encoder layer as the input to the next encoder layer, and repeat the above operation for each encoder; The output of each layer of the encoder passes through the hierarchical fusion attention layer to obtain different hierarchical attention weights, and the weighted result is the final output; Finally, based on the final output, calculate the predicted key quality variable through the linear regression layer; Step 3: According to the time-delay characteristics of the industrial process, use the historical key quality variable y detected a timestamps ago as the guiding quality variable for the current prediction timestamp. The length a of the sliding window is dynamically adjusted according to the time-delay characteristics of the industrial process. To optimize the training model, the key quality variable y at the current timestamp in the training set is used as the guiding quality variable ; Set the model learning rate, the number of model training iterations, and the batch size; according to the set batch processing, for the guidance quality variable in the training set Perform batch processing with the auxiliary variable x; use the auxiliary variable x in the training set and the guidance quality variable after dimensional expansion Train the model. In each forward propagation process, calculate the error between the predicted value and the true value through the loss function. The loss value reflects the current prediction accuracy of the model; After the loss value is calculated, calculate the gradient of the loss function with respect to the model parameters through the backpropagation algorithm; The validation set is used to evaluate the generalization ability of the model, monitor the training process, and perform hyperparameter tuning; Training and validation are carried out alternately. After each training cycle ends, the model will perform a forward propagation using the validation set to calculate the loss value on the validation set, which is used to determine whether the model is overfitting or underfitting; Step 4: Input the auxiliary variables of the test set and the corresponding guiding feature variables into the trained TAIF - TF model to obtain the predicted key quality variable values; Combine the predicted key quality variables and the true key quality variables, and evaluate the model using regression evaluation indicators; Step 5: Input any auxiliary variable and the corresponding guiding quality variable in the industrial production process into the TAIF - TF model evaluated in Step 4, and output the predicted key quality variable at the current moment.

2. The industrial quality variable prediction method based on target adaptive-hierarchical fusion Transformer according to claim 1, wherein: The quality variable of each group of guidance in step three The processing process with the corresponding auxiliary variable x is as follows: Perform dimensionality expansion on the guidance quality variable in the training set so that the dimension of the guidance quality variable is consistent with the dimension of the corresponding auxiliary variable x; perform dense layer embedding and positional encoding processing on the guidance quality variable with consistent dimensions and the auxiliary variable to obtain the auxiliary variable entering the encoder and the guidance quality variable ; The target adaptive attention layer calculates the target adaptive attention value , assigns a weight to each feature variable in the auxiliary variable , identifies the feature variables that have a greater relationship with the guidance quality variable , and the calculation process is as follows: ; In the formula is the target adaptive attention value of the d-th feature variable, indicating the correlation between the d-th feature variable and the guidance quality variable ; represents the parameter matrix to be learned and updated; GELU represents the Gaussian error linear unit function in the activation function; is an intermediate variable, , ; In the formula, represents the output of the previous layer encoder. When i = 1, is the auxiliary variable entering the encoder ; represents the d-th feature variable in the output feature vector of the previous layer encoder; , , are all parameter matrices to be learned and updated. During the iterative training of the model, the parameter matrices are updated by calculating the gradients through the backpropagation algorithm; Then, all the obtained target adaptive attention values are normalized, and the normalized target adaptive attention values ; where is the normalized target adaptive attention value; Finally, attention weighting is performed to obtain the feature vector after feature weight sharing by the target adaptive attention layer : ; Auxiliary variables of the input encoder While performing feature decentralization in the target adaptive attention layer of the input, it is also input into the multi-head attention layer for feature extraction, and the multi-head attention layer outputs a feature vector ; The feature vector output by the target adaptive attention layer is input into the masked multi-head attention layer for feature extraction, and the masked multi-head attention layer outputs a feature vector ; The feature vector output by the multi-head attention layer , the feature vector output by the target adaptive attention layer , and the feature vector output by the masked multi-head attention layer are subjected to residual connection and layer normalization to obtain an intermediate variable : , where LN represents layer normalization processing; Then, through the feed-forward network layer and layer normalization, the output of the current encoder layer is obtained : ; where FFN represents the processing of the feed-forward network layer; The output of each encoder layer passes through the hierarchical fusion attention layer to obtain different hierarchical attention weights, and the weighted feature vector is obtained after feature decomposition through the hierarchical fusion attention layer. : ; In the formula, represents the normalized hierarchical attention value, ; is the hierarchical attention value, representing the correlation between the output of the i-th layer in the k-th layer encoder and the key quality variable; ; represents the parameter matrix to be learned and updated; GELU represents the Gaussian error linear unit function in the activation function; and represent intermediate variables, ; ; represents the output of the previous layer encoder. When i = 1, is the feature vector output by the target adaptive attention layer in the first layer encoder; is the output of the current encoder layer; In the formula, and are both parameter matrices to be learned and updated. During the iterative training of the model, the parameter matrices are updated by calculating the gradient through the backpropagation algorithm. After extracting the feature information through multiple layers of key attention layers, a feature vector is obtained , the feature vector is input into the linear regression layer, and the predicted value of the key quality variable is calculated , .

3. The industrial quality variable prediction method based on target adaptive-hierarchical fusion Transformer according to claim 1 or 2, characterized in that: Use the broadcast mechanism to expand the dimension of the guiding quality variable g in the training set so that the dimension of the guiding quality variable g is consistent with the dimension of the corresponding auxiliary variable x.

4. The industrial quality variable prediction method based on target adaptive-hierarchical fusion Transformer according to claim 1, wherein: During the training process in Step 3, use the Adam gradient descent optimization algorithm and the MSE loss function to train the model.