Industrial key quality indicator prediction method based on multi-level spatio-temporal correlation features

By dividing the entire industrial production process into sub-processes and constructing a graph adjacency matrix, and using multi-level spatiotemporal dependency sensing blocks to extract global and local spatiotemporal features, the model is optimized to learn the topological relationships between process variables. This solves the problem of poor prediction accuracy in existing technologies and achieves more accurate prediction of key industrial quality indicators.

CN119168454BActive Publication Date: 2025-12-19CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411191442.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2025-12-19
Estimated Expiration
2044-08-28

AI Technical Summary

Technical Problem

Existing methods for predicting key industrial quality indicators neglect local information and complex nonlinear dynamic behavior of each sub-process when dealing with complex industrial processes, resulting in poor prediction accuracy.

Method used

The entire industrial production process is divided into multiple sub-processes. A graph adjacency matrix is ​​constructed for the entire process and each sub-process. Global and local spatiotemporal feature representations are extracted using multi-level spatiotemporal dependency-aware blocks. The adjacency matrix model, spatiotemporal dependency-aware feature extraction model, and prediction model are combined for optimization to learn the topological relationships and time dependencies between process variables.

Benefits of technology

It improves the accuracy and reliability of forecasting key industrial quality indicators by mining the spatial dependencies between process variables and the temporal dependencies in time series data, thereby enhancing the accuracy of forecasts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119168454B_ABST
    Figure CN119168454B_ABST
Patent Text Reader

Abstract

The application is suitable for the technical field of data processing, and provides an industrial key quality index prediction method based on multi-level space-time correlation characteristics, comprising: determining process variables related to the key quality index, and obtaining training samples; regarding the whole process and each sub-process as a target process respectively, and performing the following for each target process: screening target process variables corresponding to the target process from all the determined process variables, and obtaining a graph adjacency matrix of the target process by using an adjacency matrix model and the target process variable data; obtaining space-time feature representation corresponding to the training samples of the target process by using a space-time dependence perception feature extraction model; optimizing the adjacency matrix model, the space-time dependence perception feature extraction model and a prediction model based on the obtained space-time feature representation; and predicting the industrial key quality index at a prediction time by using the model after parameter optimization. The application can improve the accuracy and reliability of the industrial key quality index prediction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of data processing, and particularly relates to an industrial key quality index prediction method based on multi-level space-time correlation characteristics. BACKGROUND

[0002] With the continuous progress of modern industrial technology, industrial production processes are increasingly large-scale and complex. These processes involve a series of complex material energy conversion and physical and chemical reactions, aiming to efficiently convert raw materials into products. In this context, real-time optimization and control of the entire production process are crucial for improving production efficiency and reducing energy consumption. Generally, real-time detection of key quality indicators is the most effective reflection of the state of industrial production and the basis for realizing real-time optimization and control. As a key tool for modeling complex industrial processes and driving the transformation of process industries to intelligent manufacturing, it plays an important role in industrial process control.

[0003] However, due to the limitations of measurement technology and industrial environment, these quality indicators cannot be measured in time and need to be determined through laboratory tests, which leads to serious time lag in subsequent industrial process optimization and control, greatly reducing the timeliness and effectiveness of industrial process control. Due to the continuous accumulation of industrial data in distributed control systems, data-driven soft sensor modeling methods have become more effective. By building mathematical models between difficult-to-measure key quality indicators and easily-measured process variables such as temperature, liquid level, flow, and current, indirect sensing of key quality indicators can be achieved. In recent years, due to the excellent ability of deep learning in high-dimensional data processing and complex function representation, it has been widely applied to data-driven soft measurement modeling of industrial processes.

[0004] Although the prediction of key quality indicators in industrial processes has been extensively studied, there are still several key issues to be addressed: 1) Industrial production processes are usually composed of multiple groups of different equipment units, and existing data-driven models often only use global information of the entire industrial production process to guide the feature extraction process, ignoring the local information of each group of sub-processes. In fact, there is a mutual coupling relationship between the production equipment units of each group of sub-processes, and they are affected by uncertain factors such as changes in operating conditions and fluctuations in production demand, resulting in differences in working conditions of each group of sub-processes and exhibiting different local characteristics. 2) Spatially, there is a complex nonlinear coupling relationship between process variables in industrial processes due to the interaction of material flow and energy flow and the influence of process connections; temporally, due to the complexity of physical and chemical reaction mechanisms, feedback control, and changes in operating conditions, industrial processes have inherently complex nonlinear dynamic behavior. Therefore, the current prediction of industrial key quality indicators has the problem of poor accuracy SUMMARY

[0005] The embodiment of the present application provides a method for predicting an industrial key quality index based on multi-level space-time correlation characteristics, which can solve the problem of poor prediction accuracy of the industrial key quality index.

[0006] The embodiment of the present application provides a method for predicting an industrial key quality index based on multi-level space-time correlation characteristics, which can solve the problem of poor prediction accuracy of the industrial key quality index.

[0007] Determine process variables related to the key quality index of industrial production, and obtain a plurality of training samples; the training samples include process variable data at a historical time and corresponding key quality index values; the industrial production whole process includes a plurality of sub-processes;

[0008] The whole process and each sub-process are respectively taken as a target process, and the following is performed for each target process: for each training sample, target process variables corresponding to the target process are selected from all the determined process variables, and a graph adjacency matrix corresponding to the training sample of the target process is obtained according to an adjacency matrix model and target process variable data corresponding to the target process variables in the training sample; a space-time feature representation corresponding to the training sample of the target process is obtained by using a space-time dependence perception feature extraction model; the space-time dependence perception feature extraction model includes a skip connection module and L space-time dependence perception blocks connected in turn, the output ends of the L space-time dependence perception blocks are connected with the input end of the skip connection module, the space-time dependence perception block includes a spatial topology feature extraction module, a time dynamic feature extraction module and a space-time feature fusion module, the output end of the spatial topology feature extraction module and the output end of the time dynamic feature extraction module are connected with the input end of the space-time feature fusion module, the output end of the space-time feature fusion module is the output end of the space-time dependence perception block, the input end of the spatial topology feature extraction module and the input end of the time dynamic feature extraction module are the input end of the space-time dependence perception block, the input of the spatial topology feature extraction module of the first space-time dependence perception block in the L space-time dependence perception blocks is the graph adjacency matrix corresponding to the training sample of the target process and the target process variable data, and the input of the time dynamic feature extraction module of the first space-time dependence perception block is the target process variable data corresponding to the training sample of the target process;

[0009] For each training sample, a key quality index prediction value corresponding to all target processes of the training sample is obtained according to a prediction model and the space-time feature representations corresponding to all target processes of the training sample, and parameters in the adjacency matrix model, the space-time dependence perception feature extraction model and the prediction model are optimized by using the key quality index prediction values and the key quality index values corresponding to all training samples.

[0010] The adjacency matrix model, the space-time dependence perception feature extraction model and the prediction model after parameter optimization are used to predict the industrial key quality index at a prediction time.

[0011] Optionally, the adjacency matrix model comprises an adjacency weight formula and a matrix fusion formula; according to the adjacency matrix model and target process variable data corresponding to the target process variables in the training samples, a graph adjacency matrix corresponding to the training samples of the target process is obtained, comprising:

[0012] respectively for each target process variable in all target process variables: generating a node embedding based on the corresponding target process variable data of the target process variable; respectively for each other target process variable in all target process variables except the target process variable, generating a node embedding based on the corresponding target process variable data of the other target process variable, and calculating the adjacency weight between the target process variable and the other target process variable according to the generated node embedding and the adjacency weight formula;

[0013] Based on all adjacency weights, an asymmetric adjacency matrix is constructed, and the asymmetric adjacency matrix is sparsified to obtain a sparsified asymmetric adjacency matrix; each element in the asymmetric adjacency matrix is the adjacency weight of two target process variables in all target process variables, and each element corresponds to different adjacency weights;

[0014] From the plurality of training samples, target process variable data corresponding to the target process variables is selected, and the correlation between each two target process variables is calculated according to the target process variable data corresponding to each two target process variables in all target process variables, and a correlation adjacency matrix is constructed based on all calculated correlations; each element in the correlation adjacency matrix is the correlation between two target process variables in all target process variables, and each element corresponds to different correlations;

[0015] The sparsified asymmetric adjacency matrix and the correlation adjacency matrix are fused by the matrix fusion formula to obtain the graph adjacency matrix corresponding to the training samples of the target process.

[0016] Optionally, the adjacency weight formula is:

[0017] M 1i =tanh(αE 1i Γ1)

[0018] M 2j =tanh(αE 2j Γ2)

[0019]

[0020] Wherein, E 1i represents the node embedding based on the i-th target process variable, E 2j represents the node embedding based on the j-th target process variable, M represents the number of target process variables, d represents the node embedding dimension, Γ1 and Γ2 are both learnable parameters, α is a hyper-parameter controlling the saturation rate of the activation function, tanh(·), ReLU(·) represent the activation function, A learnij denotes the adjacency weight between the i-th target process variable and the j-th target process variable, i≠j.

[0021] Optionally, the matrix fusion formula is:

[0022] A=α′A corr +(1-α)A learn

[0023] wherein A denotes a graph adjacency matrix, α' denotes a learnable trade-off parameter, A corr denotes a correlation adjacency matrix, A learn denotes a sparse asymmetric adjacency matrix.

[0024] Optionally, the spatial topology feature extraction module comprises a spatial position attention sub-module and a mixed jump graph convolution sub-module connected in sequence, an input end of the spatial position attention sub-module is an input end of the spatial topology feature extraction module, and an output end of the mixed jump graph convolution sub-module is an output end of the spatial topology feature extraction module.

[0025] Optionally, the time dynamic feature extraction module comprises a time position attention sub-module and a multi-scale dilated causal convolution sub-module connected in sequence, an input end of the time position attention sub-module is an input end of the time dynamic feature extraction module, and an output end of the multi-scale dilated causal convolution sub-module is an output end of the time dynamic feature extraction module.

[0026] Optionally, the spatio-temporal feature fusion module is configured to fuse the spatial topology feature output by the spatial topology feature extraction module and the time dynamic feature output by the time dynamic feature extraction module through a feature fusion formula to obtain a spatio-temporal feature fusion result.

[0027] The feature fusion formula is: Z l denotes a spatio-temporal feature fusion result output by a spatio-temporal feature fusion module of the l-th spatio-temporal dependency perception block, s l denotes a spatial attention vector of the l-th spatio-temporal dependency perception block, denotes a spatial topology feature output by a spatial topology feature extraction module of the l-th spatio-temporal dependency perception block, t l denotes a time attention vector of the l-th spatio-temporal dependency perception block, denotes a time dynamic feature output by a time dynamic feature extraction module of the l-th spatio-temporal dependency perception block.

[0028] Optionally,

[0029]

[0030] wherein, α1, α2, β1, β2 are all 1x1 convolutional layers, and σ represents a sigmoid activation function.

[0031] Optionally, according to the prediction model and the spatio-temporal feature representation of all target processes corresponding to the training sample, a predicted value of the key quality indicator corresponding to the training sample of all target processes is obtained, including:

[0032] For the whole process, the spatio-temporal feature representation of all target processes corresponding to the training sample is fused, and the fused spatio-temporal feature representation is input into the prediction model for processing to obtain a predicted value of the key quality indicator corresponding to the training sample of the whole process.

[0033] For each sub-process, the spatio-temporal feature representation of the sub-process corresponding to the training sample is input into the prediction model for processing to obtain a predicted value of the key quality indicator corresponding to the training sample of the sub-process.

[0034] Optionally, the parameters in the adjacency matrix model, the spatio-temporal dependence perception feature extraction model and the prediction model are optimized using the predicted value of the key quality indicator corresponding to all training samples and the value of the key quality indicator, including:

[0035] The prediction loss value is calculated based on the loss function, and the parameters in the adjacency matrix model, the spatio-temporal dependence perception feature extraction model and the prediction model are optimized using the prediction loss value.

[0036] The loss function is:

[0037]

[0038] wherein, J(Θ) represents the prediction loss value, Θ represents the learnable parameters in the adjacency matrix model, the spatio-temporal dependence perception feature extraction model and the prediction model, N represents the number of training samples, groupN represents the number of sub-processes, represents a predicted value of the key quality indicator corresponding to the i-th training sample of the whole process, y i represents a value of the key quality indicator in the i-th training sample, represents a predicted value of the key quality indicator corresponding to the i-th training sample of the j-th sub-process.

[0039] The above-mentioned scheme of the present application has the following beneficial effects:

[0040] In the embodiments of the present application, by dividing the whole industrial production process into multiple sub-processes, constructing graph adjacency matrices of the whole process and each sub-process, on the basis of each graph adjacency matrix, using a multi-level spatio-temporal dependency perception block to extract global spatio-temporal feature representation of the whole process and local spatio-temporal feature representation of each sub-process, and then combining the global spatio-temporal feature representation and each local spatio-temporal feature representation to optimize the adjacency matrix model, the spatio-temporal dependency perception feature extraction model and the prediction model, the adjacency matrix model, the spatio-temporal dependency perception feature extraction model and the prediction model can learn the underlying topological structure relationship between the process variables related to the key quality indicators of industrial production, mine the spatial dependency between the process variables and the temporal dependency in the time series data, and further enable the prediction of the key quality indicators of industry to greatly improve the accuracy and reliability of the prediction of the key quality indicators of industry when using the optimized adjacency matrix model, the spatio-temporal dependency perception feature extraction model and the prediction model.

[0041] Other benefits of the present application will be described in detail in the subsequent specific embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0043] Figure 1 A flowchart of an industrial key quality indicator prediction method based on multi-level spatio-temporal correlation features provided by an embodiment of the present application;

[0044] Figure 2 A structural schematic diagram of a spatio-temporal dependency perception feature extraction model provided by an embodiment of the present application;

[0045] Figure 3 A simplified flowchart of a potassium salt flotation process in the related art;

[0046] Figure 4 A comparison curve diagram of predicted values and true values of concentrate potassium ion grade in a potassium salt flotation process by an LSTNet model in an example;

[0047] Figure 5 A comparison curve diagram of predicted values and true values of concentrate potassium ion grade in a potassium salt flotation process by a VW-SAE model in an example;

[0048] Figure 6 A comparison curve diagram of predicted values and true values of concentrate potassium ion grade in a potassium salt flotation process by a LogTrans model in an example;

[0049] Figure 7 Figure 6 is a graph showing the comparison between the predicted value and the true value of the potassium ion grade of the concentrate in the potassium salt flotation process according to the prediction method of the present application. DETAILED DESCRIPTION

[0050] In the following description, for purposes of explanation and not limitation, specific details are set forth such as particular architectures, techniques, etc. in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known methods, devices, circuits, and

[0051] It should be understood that the term "comprises" when used in this specification and the appended claims indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0052] It should also be understood that the term "and / or" when used in this specification and the appended claims indicates that the associated listed items can be present one or more of the associated listed items, and that the items are not limited to only one of the associated listed items.

[0053] As used in this specification and the appended claims, the term "if' can be construed to mean "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [a described condition or event] is detected" can be construed to mean "upon determining" or "in response to determining" or "upon detecting [the described condition or event]" or "in response to detecting [the described condition or event]," depending on the context.

[0054] In addition, the terms "first," "second," "third," etc. are used herein only to distinguish one element from another, and do not imply a relative importance or a given order.

[0055] Reference within the specification of this application to "one embodiment" or "some embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrase "in one embodiment" or "in some embodiments" in various places within specifications are not necessarily all referring to the same embodiment, however, are meant to signify that "one or more, but not all embodiments" of the application so described are contemplated to develop a full and enabling disclosure. The terms "including," "comprising," "having," and variations thereof, are meant to encompass the items listed thereafter and equivalents thereof as well as additional items.

[0056] In view of the poor prediction accuracy of the current industrial key quality indicators, the embodiments of the present application provide an industrial key quality indicator prediction method based on multi-level space-time correlation features. The full process of industrial production is divided into multiple sub-processes, and the graph adjacency matrix of the full process and each sub-process is constructed. On the basis of each graph adjacency matrix, the multi-level space-time dependence perception block is used to extract the global space-time feature representation of the full process and the local space-time feature representation of each sub-process. Then, the global space-time feature representation and each local space-time feature representation are combined to optimize the adjacency matrix model, the space-time dependence perception feature extraction model, and the prediction model, so that the adjacency matrix model, the space-time dependence perception feature extraction model, and the prediction model can learn the underlying topological structure relationship between the process variables related to the key quality indicators of industrial production, and mine the spatial dependence between the process variables and the temporal dependence in the time series data. Thus, when the optimized adjacency matrix model, space-time dependence perception feature extraction model, and prediction model are used for industrial key quality indicator prediction, the accuracy and reliability of industrial key quality indicator prediction can be greatly improved.

[0057] The industrial key quality indicator prediction method based on multi-level space-time correlation features provided by the embodiments of the present application will be exemplarily described below in combination with specific embodiments.

[0058] As shown in Figure 1 The industrial key quality indicator prediction method based on multi-level space-time correlation features provided by the embodiments of the present application includes the following steps:

[0059] Step 11, determining the process variables related to the key quality indicators of industrial production, and obtaining a plurality of training samples.

[0060] The training samples include process variable data at a historical time and corresponding key quality indicator values, and the full process of industrial production includes multiple sub-processes. Specifically, the full process can be divided into multiple sub-processes based on the equipment involved in industrial production, that is, the process flow corresponding to each (or multiple) equipment in industrial production is regarded as a sub-process.

[0061] For example, in the industrial scene of potash flotation, the concentrate potassium ion grade can be determined as the key quality index of the potash flotation process through process mechanism, and a plurality of process variables (such as temperature, liquid level, flow, current, etc.) having high correlation with the concentrate potassium ion grade can be selected from the potash flotation process data as auxiliary variables for modeling. Correspondingly, the training samples described above include process variable data (i.e. temperature value, liquid level value, flow value, current value, etc.) and corresponding key quality index value (i.e. concentrate potassium ion grade value) at historical time, and the historical time corresponding to each training sample is different. In real-time cases, 2117 time industrial process data sample points can be obtained, each sample point including 60-dimensional features.

[0062] For subsequent data processing, after obtaining the historical industrial process data (i.e. process variable data at historical time and corresponding key quality index value), the historical industrial process data can be preprocessed, the preprocessed data can be segmented according to time window using sliding window technology, and the industrial process data features can be mapped to a specified dimension using an input convolutional layer to obtain each training sample.

[0063] In specific implementation, first, the historical industrial process data is preprocessed by missing value completion, outlier removal, and data filtering, and then the collected key quality index data (i.e. key quality index value) and its auxiliary variables for modeling (i.e. process variables) are normalized, assuming is the time series data set after preprocessing, where X and Y represent the auxiliary variables for modeling and the key quality index respectively, N and M represent the total number of samples and the number of auxiliary variables respectively, and the normalization calculation formula of χ is as follows:

[0064] The auxiliary variables X for modeling are normalized by the following formula:

[0065] X = (X - X min,i ) / (X max,i -X min,i )

[0066] The key quality index Y is normalized by the following formula:

[0067] Y = (Y - Y min ) / (Y max -Y min )

[0068] where X min,i and X max,i represent the minimum and maximum values of the i-th auxiliary variable, and Y min and Y max represent the minimum and maximum values of the key quality index (i.e. concentrate potassium ion grade).

[0069] In order to facilitate the subsequent model extraction time dynamic characteristics, further adopt sliding window technology to industrial process data set according to time window segmentation, obtain time window data as original input data. Wherein, R is the length of the historical sample set lookback window required by the model extraction time dynamic characteristics, P is the length of the key sample index prediction window, the step length of the window sliding is 1, then the time window data obtained can be expressed as the following formula:

[0070]

[0071] Wherein, X w and Y w respectively represent the auxiliary variable and the key quality index in the time window data, Slidingwindow R+P (·) represents that the data set is segmented according to the time window with the window length of R+P and the step length of 1.

[0072] Then, in order to be able to more efficiently extract and process the key space-time characteristics subsequently, the data characteristics are mapped to a high-dimensional latent space by using a 1x1 input convolutional layer, and the original input data characteristics of the subsequent model are obtained:

[0073]

[0074] Startconv(·) represents the input convolution operation, is the original input data feature of the subsequent space-time dependent feature extraction module, and c is the data feature dimension.

[0075] The first 90% of the time window data is used as the training set, and the last 10% of the time window data is used as the test set to test the generalization performance of the trained model subsequently. Wherein, the auxiliary variable and the key quality index in each time window data are a training sample.

[0076] Step 12, the whole process, each sub-process is respectively regarded as a target process, and each target process is executed respectively: for each training sample, the target process variable corresponding to the target process is selected from all the process variables determined, and the graph adjacency matrix corresponding to the training sample of the target process is obtained according to the adjacency matrix model and the target process variable data corresponding to the target process variable in the training sample; the space-time feature representation corresponding to the training sample of the target process is obtained by using the space-time dependent feature extraction model.

[0077] In some embodiments of the present application, in order to facilitate the training of the model and enable the model to learn the underlying topological structure relationship between the process variables, the spatio-temporal feature representation of each training sample in the whole process and each sub-process needs to be obtained. Since the data processing process is the same in each process, for the sake of description, the whole process and each sub-process are regarded as a target process, and the acquisition process of the spatio-temporal feature representation is described taking the target process as an example.

[0078] In a target process, the spatio-temporal feature representation of each training sample needs to be obtained. Specifically, for each training sample, first, the target process variable corresponding to the current target process needs to be selected from all the process variables determined in step 11, then the process variable data corresponding to the target process variable (i.e., the target process variable data) is determined from the process variable data of the training sample, and then the graph adjacency matrix corresponding to the training sample of the target process is obtained according to the adjacency matrix model and the target process variable data corresponding to the target process variable in the training sample.

[0079] The adjacency matrix model includes an adjacency weight formula and a matrix fusion formula. The specific implementation of obtaining the graph adjacency matrix corresponding to the training sample of the target process according to the adjacency matrix model and the target process variable data corresponding to the target process variable in the training sample includes the following steps:

[0080] Step 12.1, for each target process variable in all target process variables, respectively, perform: generating a node embedding based on the corresponding target process variable data of the target process variable; for each other target process variable in all target process variables except the target process variable, respectively, generate a node embedding based on the corresponding target process variable data of the other target process variable, and calculate the adjacency weight between the target process variable and the other target process variable according to the generated node embedding and the adjacency weight formula.

[0081] The adjacency weight formula is:

[0082] M 1i = tanh(aE 1i Γ1)

[0083] M 2j = tanh(aE 2j Γ2)

[0084]

[0085] wherein, E 1i represents the node embedding based on the i-th target process variable, E 2j represents the node embedding based on the j-th target process variable, M represents the number of target process variables, d represents the dimension of the node embedding, Γ1and Γ2are both learnable parameters during training, and a is a hyperparameter that controls the saturation rate of the activation function, which is calculated by The asymmetry of the adaptive adjacency matrix is realized, and tanh(·) and ReLU(·) each represent an activation function, A learnij represents the adjacency weight of the ith target process variable and the jth target process variable, i≠j, j=1,…,T, i=1,…,T, and T represents the number of target process variables. For example, when generating node embeddings, the target process variable data can be input into an embedding layer for processing to obtain corresponding node embeddings.

[0086] Step 12.2, based on all adjacency weights, an asymmetric adjacency matrix is constructed, and the asymmetric adjacency matrix is sparsified to obtain a sparsified asymmetric adjacency matrix. Each element in the asymmetric adjacency matrix is the adjacency weight of two target process variables in all target process variables, and each element corresponds to a different adjacency weight, i.e., each element corresponds to a different target process variable.

[0087] In some embodiments of the present application, the topk algorithm can be used to select the top k nearest nodes (i.e., other target process variables) of each node (i.e., target process variable) as its neighbor, while preserving the weights of connected nodes and setting the weights of non-connected nodes to zero, as follows:

[0088] for i=1,2,…,M,M+1

[0089] idx=argtopk(A learn [i,:])

[0090] A learn [i:-idx]=0

[0091] where argtopk(·) represents returning the top k most relevant indexes of a vector, i represents the index of the ith node, and M represents the number of target process variables. The above expression means that for a certain target process variable, the adjacency weights between it and other target process variables are sorted in descending order, and the top k adjacency weights are preserved, and the other adjacency weights are set to zero.

[0092] Step 12.3, target process variable data corresponding to the target process variables are screened out from a plurality of training samples, and the correlation between each two target process variables is calculated according to the target process variable data corresponding to each two target process variables in all target process variables, and a correlation adjacency matrix is constructed based on all the calculated correlations. Each element in the correlation adjacency matrix is the correlation between two target process variables in all target process variables, and each element corresponds to a different correlation, i.e., each element corresponds to a different target process variable.

[0093] In some embodiments of the present application, the correlation between the target process variables can be calculated by using the Pearson correlation coefficient, i.e., the Pearson correlation coefficient between the target process variables is taken as the correlation between the two. It should be noted that when calculating the correlation between the target process variables, the target process variable data in all training samples is used for calculation. Taking the temperature and liquid level as two process variables, the temperature values in all training samples are taken as a sequence, and the liquid level values in all training samples are taken as a sequence, and then the Pearson correlation coefficient of the two sequences is calculated.

[0094] wherein the calculation formula of the Pearson correlation coefficient is: Pearson(X i ,X j ) represents the Pearson correlation coefficient between the sequence X i and the sequence X j , cov(·) represents the covariance between the sequence X i and the sequence X j , var(·) represents the variance of the sequence, the sequence X i represents the sequence corresponding to the i-th target process variable, and the sequence X j represents the sequence corresponding to the j-th target process variable.

[0095] After the correlation is calculated, the correlation adjacency matrix can be obtained by the threshold parameter k. Specifically, the element A corr,ij in the correlation adjacency matrix can be determined by the following formula:

[0096]

[0097] A corr,ij represents the correlation weight between the i-th target process variable and the j-th target process variable, i≠j, j=1,…,T, i=1,…,T, T represents the number of target process variables. It can be understood that after the correlation weights between the target process variables are determined, the correlation adjacency matrix is constructed based on all the correlation weights, and each element in the correlation adjacency matrix is the correlation (i.e., the correlation weight) between two target process variables among all target process variables, and each element corresponds to a different correlation.

[0098] Step 12.4, fuse the sparse asymmetric adjacency matrix and the correlation adjacency matrix by the matrix fusion formula to obtain the graph adjacency matrix corresponding to the training sample of the target process.

[0099] The above matrix fusion formula is:

[0100] A=α′A corr +(1-α)Alearn

[0101] Where A represents the graph adjacency matrix, and α′ represents a learnable tradeoff parameter used to adjust the relevance adjacency matrix A. corr Asymmetric adjacency matrix A learn The importance of A in order to better capture the potentially complex relationships in industrial processes. corr Let A represent the relevance adjacency matrix. learn This represents the sparsified asymmetric adjacency matrix.

[0102] In some embodiments of this application, in a target process, it is necessary to obtain the spatiotemporal feature representation of each training sample. Specifically, the graph adjacency matrix of the training sample and the target process variable data in the training sample can be input into a spatiotemporal dependency-aware feature extraction model for processing to obtain the spatiotemporal feature representation of the training sample.

[0103] like Figure 2 As shown, the above spatiotemporal dependency-aware feature extraction model includes a skip connection module (such as...). Figure 2 (with circles containing +) and L sequentially connected spatiotemporal dependent perceptual blocks (such as ... Figure 2 The model consists of spatiotemporal dependency sensing blocks 1 to L. The outputs of all L spatiotemporal dependency sensing blocks are connected to the input of the jump connection module. The output of the jump connection module is the output of the spatiotemporal dependency sensing feature extraction model, used to output spatiotemporal feature representations. It can be understood that when the data input to the spatiotemporal dependency sensing feature extraction model corresponds to the entire process, the jump connection module outputs a global spatiotemporal feature representation of industrial production; when the data input to the spatiotemporal dependency sensing feature extraction model corresponds to data corresponding to a sub-process, the jump connection module outputs a local spatiotemporal feature representation of industrial production.

[0104] The aforementioned spatiotemporal dependency-aware block includes a spatial topology feature extraction module (such as...). Figure 2 The graph convolution spatial feature extraction module and the temporal dynamic feature extraction module (such as...) Figure 2The system consists of a temporal convolutional temporal feature extraction module and a spatiotemporal feature fusion module. The outputs of the spatial topology feature extraction module and the temporal dynamic feature extraction module are both connected to the input of the spatiotemporal feature fusion module. The output of the spatiotemporal feature fusion module is the output of the spatiotemporal dependency sensing block. The inputs of the spatial topology feature extraction module and the temporal dynamic feature extraction module are both inputs of the spatiotemporal dependency sensing block. The input of the spatial topology feature extraction module of the first spatiotemporal dependency sensing block is the graph adjacency matrix of the target process corresponding to the training samples and the target process variable data. The input of the temporal dynamic feature extraction module of the first spatiotemporal dependency sensing block is the target process variable data of the target process corresponding to the training samples. For other spatiotemporal dependency sensing blocks besides the first one, their input is the output of the previous spatiotemporal dependency sensing block.

[0105] The spatial topology feature extraction module is used to extract the spatial topology features of its input data. Specifically, the spatial topology feature extraction module includes sequentially connected spatial location attention sub-modules (such as...). Figure 2 Spatial location attention in (and hybrid jump graph convolutional submodules such as) and Figure 2 In the hybrid jump graph convolution (HPP), the input of the spatial location attention submodule is the input of the spatial topology feature extraction module, and the output of the hybrid jump graph convolution submodule is the output of the spatial topology feature extraction module.

[0106] The temporal dynamic feature extraction module is used to extract the temporal dynamic features of its input data. Specifically, the temporal dynamic feature extraction module includes sequentially connected temporal position attention sub-modules (such as...). Figure 2 Temporal positional attention) and multi-scale dilated causal convolutional submodules (such as Figure 2 In the multi-scale dilated causal temporal convolution, the input of the temporal position attention submodule is the input of the temporal dynamic feature extraction module, and the output of the multi-scale dilated causal convolution submodule is the output of the temporal dynamic feature extraction module.

[0107] The spatiotemporal feature fusion module uses an attention mechanism to deeply fuse the extracted spatial topological features and temporal dynamic features, uncovering rich spatiotemporal interaction patterns in the input data. Specifically, the module fuses the spatial topological features output by the spatial topological feature extraction module and the temporal dynamic features output by the temporal dynamic feature extraction module using a feature fusion formula, obtaining a spatiotemporal feature fusion result. This enables a more comprehensive and accurate description of the spatiotemporal dependencies of the input sequence.

[0108] The feature fusion formula is: Z l s represents the spatiotemporal feature fusion result output by the spatiotemporal feature fusion module of the l-th spatiotemporal dependent sensing block.l a spatial attention vector of the lth spatiotemporal dependent perception block, a spatial topology feature output by the spatial topology feature extraction module of the lth spatiotemporal dependent perception block, t l a temporal attention vector of the lth spatiotemporal dependent perception block, a temporal dynamic feature output by the temporal dynamic feature extraction module of the lth spatiotemporal dependent perception block, l = 1, …, L, L being the number of spatiotemporal dependent perception blocks.

[0109]

[0110] wherein a1, a2, b1, b2 are all 1x1 convolution layers, and s represents a sigmoid activation function.

[0111] The skip connection module is configured to add the spatiotemporal feature fusion results output by each spatiotemporal dependent perception block to obtain a spatiotemporal feature representation. As an optional example, the skip connection module can be a commonly used skip connection module.

[0112] The processing procedures of the spatial topology feature extraction module and the temporal dynamic feature extraction module on data are exemplarily described below.

[0113] The spatial topology feature extraction module mainly adopts a spatial position attention mechanism to locate key process variable nodes, and uses a hybrid skip graph convolution layer to fuse the feature information of each process variable and its neighbor process variable based on the constructed graph adjacency matrix, to obtain the spatial topology features between the process variables. As an optional example, the spatial position attention submodule in the spatial topology feature extraction module can be a commonly used position attention module, and the hybrid skip graph convolution submodule can be a commonly used hybrid skip graph convolution network. Of course, it can also be a spatial topology feature obtained by performing the following steps.

[0114] Specifically, the spatial position attention submodule in the spatial topology feature extraction module is configured to perform the following steps:

[0115] Step 12.5, a one-dimensional pooling kernel (1, W) in the time direction is used to aggregate each dimension feature along the variable (i.e., process variable) direction, to obtain a variable direction perceived feature vector by encoding the spatial dependency relationship between variables, the output of the cth channel and the hth variable which can be represented as follows:

[0116]

[0117] wherein W represents the dimension of the input data feature Z (i.e., the target process variable data) of the spatial position attention submodule in the time direction, z c(h, i) represents the input data feature of the cth channel hth variable, i represents the index of the data feature in the time direction.

[0118] Step 12.6, aggregate all feature vectors according to the generated variable direction y S , pass it through two 1x1 convolution transformations in sequence to encode the spatial dependence between input data variables, more accurately locate the feature representation of key variables, which can be represented as follows:

[0119] f S = δ (F S1 (y S ))

[0120] g S = δ (F S2 (f S ))

[0121] Where F S1 and F S2 represent two 1x1 convolution transformations respectively, δ is a nonlinear activation function, y S is the generated variable direction aggregation feature vector, f S is the intermediate feature mapping of variable direction encoding spatial information, and g S is the encoded variable direction attention weight.

[0122] Step 12.7, combine the encoded variable direction attention weight g S with the original input data feature Z to obtain the feature representation X S embedding the spatial dependence between variables, achieving more accurate representation of key variable features, which can be represented as:

[0123] X S = g S ·Z

[0124] Where X S is the input feature of the subsequent mixed jump graph convolution submodule.

[0125] The mixed jump graph convolution submodule in the spatial topology feature extraction module is used to perform the following steps:

[0126] Step 12.8, use the mixed jump graph convolution layer to process the information flow of each node (i.e. process variable) in the graph based on the constructed graph adjacency matrix. By propagating information horizontally, a part of the original state of the node is preserved during the propagation process, while preserving the local features of the original node, while exploring its deep neighborhood, which can be represented by the following formula:

[0127]

[0128] where H (k) denotes the hidden state of the node after the k-th propagation, H in denotes the input hidden features of the previous layer, and β is a hyperparameter that controls the proportion of the original node features to be reserved, is the normalized graph adjacency matrix, H (k-1) denotes the hidden state of the node before the k-th propagation, and k denotes the current propagation layer, is the diagonal matrix of the graph adjacency matrix A.

[0129] Step 12.9, an information selection step is introduced to filter out redundant information generated by each-hop graph convolution and retain important information, which can be represented as:

[0130]

[0131] where H out denotes the final selected node feature representation after multi-step propagation, K is the depth of the graph convolution propagation, and W (k) serves as a feature selector to select node feature information.

[0132] Step 12.10, the graph convolution module is composed of two mixed-hop graph convolution layers, which process the incoming and outgoing information through each node respectively, and the outputs of the two mixed-hop graph convolution layers are added to obtain the net incoming information of each node:

[0133] Z S = MixhopGCN(X S , A) + MixhopGCN(X S , A T )

[0134] where MixhopGCN(·) denotes the process of extracting spatial topological features between variables by the mixed-hop graph convolution layer, and Z S is the spatial topological feature extracted by the spatial topological feature extraction module.

[0135] In specific implementation, first, the spatial position attention perceives the spatial information of variables, and a convolutional transformation is used to encode the spatial dependence between variables to obtain an attention weight matrix in the variable direction, so as to emphasize the spatial feature representation of key variable nodes. Then, based on the constructed graph adjacency matrix, the node information is propagated horizontally and selected vertically, and the information flow of each node in the graph is processed by the graph convolution module to gradually extract the spatial topological features between variables.

[0136] The time dynamic feature extraction module mainly adopts a time position attention mechanism to locate key time points, and uses a multi-scale dilated causal convolution to efficiently capture the sequential patterns and dynamic changes of different scales of each variable, to obtain the time dynamic features between sequences. As an optional example, the time position attention submodule in the time dynamic feature extraction module can be a commonly used position attention, and the multi-scale dilated causal convolution submodule can be a commonly used multi-scale dilated causal convolution kernel. Of course, it can also be a time dynamic feature obtained by performing the following steps.

[0137] Among them, the time position attention submodule in the time dynamic feature extraction module is mainly used to perform the following steps:

[0138] Step 12.11, a one-dimensional pooling kernel (H, 1) in the spatial direction is used to aggregate each dimension of the feature along the time direction, and a feature vector perceived in the time direction is obtained by encoding the long-range dependency relationship of different time steps. The output of the cth channel and the wth time step is which can be represented as follows:

[0139]

[0140] Among them, H represents the dimension of the input data feature Z (i.e. the target process variable data) of the time position attention submodule in the spatial direction, and z c (j, w) represents the input data feature of the cth channel and the wth time step, and j represents the index of the data feature in the spatial direction.

[0141] Step 12.12, all feature vectors are aggregated according to the generated time direction to obtain y T , which is then passed through two 1x1 convolution transformations to encode the time dependency between different time steps of the input data and more accurately locate the feature representation of the key time step. This process can be represented as follows:

[0142] f T = δ (F T1 (y T ))

[0143] g T = δ (F T2 (f T ))

[0144] Among them, F T1 and F T2 represent two 1x1 convolution transformations, δ is a nonlinear activation function, y T is the generated time direction aggregated feature vector, f T is the intermediate feature mapping of the time direction encoding the time information, and g TThe encoded time direction attention weight g is obtained.

[0145] Step 12.13, the encoded time direction attention weight g is obtained. T Combined with the original input feature Z, a feature representation X embedding the time dependence between different time steps is obtained. T , to achieve a more accurate representation of the key time step feature, the process is represented as:

[0146] X T = g T · Z

[0147] Where X T is the input feature of the subsequent multi-scale dilated causal convolution sub-module.

[0148] The multi-scale dilated causal convolution sub-module in the time dynamic feature extraction module is mainly used to perform the following steps:

[0149] Step 12.14, a multi-scale dilated causal convolution is designed to capture the local time pattern of each variable (i.e., process variable) in the industrial time series data, and different layers of dilated convolution are used to expand the receptive field and increase the distance of information propagation, so that the model is more sensitive to the local dependence in the sequence data through multi-scale information extraction. The specific expression of the multi-scale dilated causal convolution process is:

[0150]

[0151] Where, represents a two-dimensional convolution filter, d represents a convolution operation with a dilated factor, d is the dilated factor, k i represents the size of the i-th convolution filter.

[0152] With the stacking of convolution layers, the receptive field is constantly expanded by the dilated factor to extract long-term information. Assuming that the initial dilated factor is 1, the receptive field R of an m-layer time convolution network with a convolution kernel size of c is:

[0153]

[0154] Where q is the exponential growth rate of the dilated factor of each time convolution layer.

[0155] Step 12.15, two multi-scale dilated causal convolution modules are used to realize deep extraction of time dynamic features, one of which is followed by a tanh activation function as a filter to capture key information, and the other is followed by a sigmoid activation function as a gate to control the outflow of information, to obtain the output time dynamic feature Z T as shown in the following formula:

[0156] Z T= tanh(X T * F1) x sigma(X T * F2)

[0157] wherein sigma(·) represents a sigmoid activation function, F1 represents a convolution filter used in combination with a tanh activation function, F2 represents a convolution filter used in combination with a sigmoid activation function, Z T is a time dynamic feature extracted by the time dynamic feature extraction module.

[0158] In specific implementation, firstly, time position attention is used to perceive time information, and a convolution transformation is used to encode dynamic dependency between different time steps of input data to obtain an attention weight matrix in the time direction, so as to highlight dynamic feature representation of key time steps. Then, a multi-scale dilated causal convolution and a gating mechanism are used to realize deep extraction of dynamic features of time series data.

[0159] It can be seen that by taking global information of the whole industrial production process and local information of each group of device unit sub-processes as input, and respectively passing through multiple space-time dependency perception blocks composed of a graph convolution spatial feature extraction module, a time convolution time feature extraction module and a space-time feature fusion module, effective space-time feature representation can be obtained.

[0160] That is, for an actual industrial production process, considering that the whole industrial production process is usually composed of multiple groups of different device unit sub-processes, in order to extract global space-time features of the whole industrial production process and local space-time features of each group of device unit sub-processes, the global information of the whole industrial production process and the local information of each group of device unit sub-processes are taken as input of the space-time dependency perception feature extraction model, and are processed through multiple stacked space-time dependency perception blocks, so that deep space-time features of the whole industrial production process and each group of sub-processes can be more fully extracted. Further, the outputs of each space-time dependency perception block are connected in a skip connection manner to obtain final global and local space-time feature representation of the industrial production.

[0161] Step 13, for each training sample, according to the prediction model and the space-time feature representation of all target processes corresponding to the training sample, the key quality indicator prediction value of all target processes corresponding to the training sample is obtained, and the parameters in the adjacency matrix model, the space-time dependency perception feature extraction model and the prediction model are optimized by using the key quality indicator prediction value and the key quality indicator value corresponding to all training samples.

[0162] In some embodiments of the present application, for each training sample, the key quality indicator prediction value of each target process corresponding to the training sample (i.e. the key quality indicator prediction value of the training sample under each target process) can be obtained in the following manner:

[0163] Step 13.1, for the whole process, the spatio-temporal feature representation corresponding to the training sample of all target processes is fused, and the fused spatio-temporal feature representation is input into the prediction model for processing to obtain the key quality indicator prediction value corresponding to the training sample of the whole process.

[0164] That is, for the whole process, when predicting the key quality indicator, the data input into the prediction model is the fusion result of the spatio-temporal feature representation corresponding to the training sample in all target processes (i.e., the whole process and each sub-process). The fusion result can be obtained by splicing the spatio-temporal feature representation corresponding to the training sample in all target processes. The fused spatio-temporal feature representation Z all is:

[0165] Z all = concat(Z,Z group1 ,...,Z groupN )

[0166] Wherein, Z represents the spatio-temporal feature representation corresponding to the training sample in the whole process, Z group1 represents the spatio-temporal feature representation corresponding to the training sample in the first sub-process, Z groupN represents the spatio-temporal feature representation corresponding to the training sample in the Nth sub-process, and groupN is the number of sub-processes.

[0167] Specifically, the fused spatio-temporal feature representation Z all is input into a 1x1 output convolution layer (which is an implementation of the prediction model) for processing to obtain the key quality indicator prediction value corresponding to the training sample of the whole process.

[0168]

[0169] Wherein, Endconv(·) represents an output convolution operation.

[0170] Step 13.2, for each sub-process respectively, the spatio-temporal feature representation corresponding to the training sample of the sub-process is input into the prediction model for processing to obtain the key quality indicator prediction value corresponding to the training sample of the sub-process.

[0171] That is, for the sub-process, when predicting the key quality indicator, the data input into the prediction model is the spatio-temporal feature representation corresponding to the training sample in the sub-process. Specifically, the spatio-temporal feature representation corresponding to the training sample in the sub-process is input into a 1x1 output convolution layer for processing to obtain the key quality indicator prediction value corresponding to the training sample of the sub-process.

[0172] In specific implementation, in order to comprehensively consider the global information of the whole process of industrial production and the local information of each group of device unit sub-processes, the deep spatio-temporal features extracted from the industrial production whole process and each group of sub-processes by fusing the stacked multiple spatio-temporal dependence perception blocks are utilized to obtain more effective multi-level spatio-temporal feature representation, so as to improve the prediction performance of the model. Then, the multi-level spatio-temporal features are utilized by the output convolution layer to predict the key quality indicators, and the predicted values of the key quality indicators are output. In this way, the comprehensive analysis and modeling of the complex spatio-temporal dependence relationship in the industrial production process are realized.

[0173] After obtaining the predicted values of the key quality indicators corresponding to all the training samples in all the target processes, the predicted loss value can be calculated based on the loss function, and the parameters in the adjacency matrix model, the spatio-temporal dependence perception feature extraction model and the prediction model can be optimized by using the predicted loss value.

[0174] The above loss function is:

[0175]

[0176] wherein, J(Θ) represents the predicted loss value, Θ represents the learnable parameters in the adjacency matrix model, the spatio-temporal dependence perception feature extraction model and the prediction model, N represents the number of training samples, groupN represents the number of sub-processes, represents the predicted value of the key quality indicator of the whole process corresponding to the i-th training sample, y i represents the value of the key quality indicator in the i-th training sample, represents the predicted value of the key quality indicator of the j-th sub-process corresponding to the i-th training sample.

[0177] In the training process, a suitable optimizer, learning rate and iteration number can be selected, and then the weight parameters (i.e. the learnable parameters) of the whole key indicator prediction model (i.e. the adjacency matrix model, the spatio-temporal dependence perception feature extraction model and the prediction model) are supervised and updated in reverse through the obtained loss function value (i.e. the above predicted loss value), until the model converges, and the trained key indicator prediction model is obtained, i.e. the optimized adjacency matrix model, the spatio-temporal dependence perception feature extraction model and the prediction model.

[0178] It can be understood that, in some embodiments of the present application, in order to improve the prediction accuracy, the trained key indicator prediction model is first verified for prediction accuracy and generalization performance by using the collected test set samples before being used for real-time prediction of the key quality indicators. After verification, the real-time collected modeling auxiliary variables are used to predict the future key quality indicator values, so as to realize real-time online prediction of the key quality indicators.

[0179] Step 14, using the parameter-optimized adjacency matrix model, the spatio-temporal dependence-aware feature extraction model and the prediction model to predict the industrial key quality indicators at the prediction time.

[0180] In some embodiments of the present application, when real-time prediction is performed using the parameter-optimized adjacency matrix model, the spatio-temporal dependence-aware feature extraction model and the prediction model, current process variable data can be collected, and then the current process variable data is input into the parameter-optimized adjacency matrix model, the spatio-temporal dependence-aware feature extraction model and the prediction model for processing to obtain the predicted value of the industrial key quality indicators at the prediction time, thereby realizing real-time prediction of the industrial key quality indicators.

[0181] It can be understood that in the real-time prediction process, the current process variable data includes the data of the current process variables (such as the current temperature value, liquid level value, flow value, current value, etc.), and then the graph adjacency matrix and the spatio-temporal feature representation of the current process variable data under the whole process and each sub-process are obtained according to the processing process of the foregoing training sample, the spatio-temporal feature representations corresponding to the whole process and all sub-processes are fused, and finally the fused spatio-temporal feature representation is input into the prediction model for processing, so that the predicted value of the industrial key quality indicators at the prediction time (i.e., the current prediction of the industrial key quality indicators) is obtained, thereby realizing real-time prediction of the industrial key quality indicators.

[0182] The method will be further described below in combination with a specific embodiment. As shown in Figure 3 Fig. 1, the multi-level spatio-temporal correlation feature spatio-temporal dependence-aware industrial key quality indicator prediction method will be further described below taking the potassium salt flotation process as an example.

[0183] The coefficient of determination (R 2 ), root mean square error (RMSE), mean absolute error (MAE) and mean absolute percentage error (MAPE) are used to quantitatively evaluate the prediction performance of the obtained key indicator prediction model. The LSTNet model, the VW-SAE model and the LogTrans model are respectively used for comparison with the prediction method of the present application to verify the effectiveness of the method, and the prediction results are shown in Figs. Figure 4 、 5 , 6, 7, and the comparison results of the prediction performance indicators of various methods are shown in Table 1.

[0184] Model [R 2 ]]> RMSE MAE MAPE LSTNet 0.7426 0.0877 0.0689 0.1908 VW-SAE 0.6722 0.0990 0.0800 0.2298 LogTrans 0.7629 0.0842 0.0643 0.1783 The prediction method of the present application 0.7875 0.0797 0.0601 0.1652

[0185] Table 1

[0186] It can be seen that the industrial key quality indicator prediction method based on multi-level space-time correlation features provided in the application has the best prediction effect on various evaluation indexes compared with traditional deep learning prediction methods such as LSTNet model, VW-SAE model and LogTrans model, which verifies the effectiveness of the prediction method provided in the application.

[0187] In summary, the industrial key quality indicator prediction method provided in the embodiments of the application uses technologies such as graph adjacency matrix, spatial topological feature and time dynamic feature extraction, and multi-level full-process global feature and each group of sub-process local feature fusion, can accurately obtain the underlying topological structure relationship between industrial process variables, on the basis of which the spatial dependency between industrial process variables and the time dependency in time series data are fully mined and utilized, while the global information of the whole industrial production process and the local information of each group of sub-processes are taken into account to obtain more effective multi-level space-time feature representation, thereby improving the accuracy and reliability of modeling and prediction, and further enabling the accuracy and reliability of industrial key quality indicator prediction to be greatly improved when the constructed model is used for industrial key quality indicator prediction.

[0188] The above is the preferred embodiment of the application. It should be noted that for those skilled in the art, without departing from the principles described in the application, several improvements and refinements can be made, which should also be considered as the protection scope of the application.

Claims

1. A method for predicting an industrial key quality indicator based on multi-level spatio-temporal correlation features, characterized in that, The method comprises the following steps: determining process variables related to a key quality indicator of industrial production, and obtaining a plurality of training samples; the training samples include process variable data and corresponding key quality indicator values at historical time points; the industrial production whole process includes a plurality of sub-processes; respectively taking the whole process and each of the sub-processes as a target process, and respectively performing the following steps for each target process: for each training sample, screening target process variables corresponding to the target process from all the determined process variables, and obtaining a graph adjacency matrix corresponding to the training sample for the target process according to an adjacency matrix model and target process variable data corresponding to the target process variables in the training sample; obtaining a spatio-temporal feature representation corresponding to the training sample for the target process using a spatio-temporal dependency perception feature extraction model; the spatio-temporal dependency perception feature extraction model comprises a skip connection module and L spatio-temporal dependency perception blocks connected in sequence, the output ends of the L spatio-temporal dependency perception blocks are connected to the input end of the skip connection module, the spatio-temporal dependency perception block comprises a spatial topology feature extraction module, a temporal dynamic feature extraction module and a spatio-temporal feature fusion module, the output ends of the spatial topology feature extraction module and the temporal dynamic feature extraction module are connected to the input end of the spatio-temporal feature fusion module, the output end of the spatio-temporal feature fusion module is the output end of the spatio-temporal dependency perception block, the input ends of the spatial topology feature extraction module and the temporal dynamic feature extraction module are the input ends of the spatio-temporal dependency perception block, the input of the spatial topology feature extraction module of the first spatio-temporal dependency perception block in the L spatio-temporal dependency perception blocks is the graph adjacency matrix corresponding to the training sample for the target process and the target process variable data, and the input of the temporal dynamic feature extraction module of the first spatio-temporal dependency perception block is the target process variable data corresponding to the training sample for the target process; for each training sample, obtaining key quality indicator prediction values corresponding to all target processes for the training sample according to a prediction model and spatio-temporal feature representations corresponding to all target processes for the training sample, and optimizing parameters in the adjacency matrix model, the spatio-temporal dependency perception feature extraction model and the prediction model using key quality indicator prediction values and key quality indicator values corresponding to all training samples; using the adjacency matrix model, the spatio-temporal dependency perception feature extraction model and the prediction model after parameter optimization to predict an industrial key quality indicator at a prediction time point.

2. The prediction method of claim 1, wherein, The adjacency matrix model comprises an adjacency weight formula and a matrix fusion formula; and the graph adjacency matrix corresponding to the training sample for the target process is obtained according to the adjacency matrix model and the target process variable data corresponding to the target process variables in the training sample, which comprises: respectively for each of all target process variables: generating a node embedding based on corresponding target process variable data of the target process variable; respectively for each of all target process variables other than the target process variable, generating a node embedding based on corresponding target process variable data of the other target process variable, and calculating an adjacency weight between the target process variable and the other target process variable according to the generated node embedding and the adjacency weight formula; constructing an asymmetric adjacency matrix based on all adjacency weights, and performing sparse processing on the asymmetric adjacency matrix to obtain a sparse asymmetric adjacency matrix; each element in the asymmetric adjacency matrix is an adjacency weight between two target process variables of all target process variables, and each element corresponds to a different adjacency weight; filtering out target process variable data corresponding to the target process variable from the plurality of training samples, and calculating a correlation between each two target process variables according to target process variable data corresponding to each two target process variables of all target process variables, and constructing a correlation adjacency matrix based on all calculated correlations; each element in the correlation adjacency matrix is a correlation between two target process variables of all target process variables, and each element corresponds to a different correlation; fusing the sparse asymmetric adjacency matrix and the correlation adjacency matrix through a matrix fusion formula to obtain a graph adjacency matrix corresponding to the target process and the training sample.

3. The prediction method of claim 2, wherein, The adjacency weight formula is: M 1i = tanh(aE 1i Γ1) M 2j = tanh(aE 2j Γ2) wherein, E 1i denotes the node embedding corresponding to the ith target process variable, E 2j denotes the node embedding corresponding to the jth target process variable, M denotes the number of target process variables, and d denotes the dimension of the node embedding, Γ1and Γ2are both learnable parameters, α is a hyperparameter that controls the saturation rate of the activation function, tanh(·) and ReLU(·) both denote an activation function, A learnij denotes the adjacency weight between the ith target process variable and the jth target process variable, i≠j.

4. The prediction method of claim 2, wherein, The matrix fusion formula is: A = a' A corr + (1 - a) A learn where A denotes a graph adjacency matrix, a' denotes a learnable trade-off parameter, A corr denotes a correlation adjacency matrix, A learn denotes a sparse asymmetric adjacency matrix.

5. The prediction method of claim 1, wherein, The spatial topology feature extraction module comprises a spatial position attention submodule and a mixed jump graph convolution submodule connected in sequence, an input end of the spatial position attention submodule is an input end of the spatial topology feature extraction module, and an output end of the mixed jump graph convolution submodule is an output end of the spatial topology feature extraction module.

6. The prediction method of claim 1, wherein, The time dynamic feature extraction module comprises a time position attention submodule and a multi-scale expansion causal convolution submodule connected in sequence, an input end of the time position attention submodule is an input end of the time dynamic feature extraction module, and an output end of the multi-scale expansion causal convolution submodule is an output end of the time dynamic feature extraction module.

7. The prediction method of claim 1, wherein, The spatio-temporal feature fusion module is configured to fuse the spatial topology feature output by the spatial topology feature extraction module and the time dynamic feature output by the time dynamic feature extraction module through a feature fusion formula to obtain a spatio-temporal feature fusion result. The feature fusion formula is: Z l s represents a spatio-temporal feature fusion result output by a spatio-temporal feature fusion module of an lth spatio-temporal dependent perception block, l s represents a spatial attention vector of the lth spatio-temporal dependent perception block, t represents a spatial topological feature output by a spatial topological feature extraction module of the lth spatio-temporal dependent perception block, l s represents a time attention vector of the lth spatio-temporal dependent perception block, s represents a time dynamic feature output by a time dynamic feature extraction module of the lth spatio-temporal dependent perception block.

8. The prediction method of claim 7, wherein, wherein α1, α2, β1, β2 are all 1x1 convolution layers, and σ represents a sigmoid activation function.

9. The prediction method of claim 1, wherein, According to the prediction model and the spatio-temporal feature representation of all target processes corresponding to the training sample, the key quality indicator prediction value of all target processes corresponding to the training sample is obtained, comprising: For the whole process, the spatio-temporal feature representations corresponding to the training sample of all target processes are fused, and the fused spatio-temporal feature representations are input into a prediction model for processing to obtain the predicted value of the key quality indicator corresponding to the training sample of the whole process; For each sub-process, the spatio-temporal feature representation corresponding to the training sample of the sub-process is input into a prediction model for processing to obtain the predicted value of the key quality indicator corresponding to the training sample of the sub-process.

10. The prediction method of claim 1, wherein, The parameters in the adjacent matrix model, the spatio-temporal dependence perception feature extraction model and the prediction model are optimized by using the predicted value and the value of the key quality indicator corresponding to all training samples, including: The loss function is used to calculate the prediction loss value, and the parameters in the adjacent matrix model, the spatio-temporal dependence perception feature extraction model and the prediction model are optimized by using the prediction loss value; The loss function is: wherein J(Θ) represents a prediction loss value, Θ represents learnable parameters in the adjacency matrix model, the spatio-temporal dependency-aware feature extraction model and the prediction model, N represents the number of training samples, groupN represents the number of sub-processes, represents a key quality indicator prediction value of the whole process corresponding to the i-th training sample, y i represents a key quality indicator value in the i-th training sample, represents a key quality indicator prediction value of the j-th sub-process corresponding to the i-th training sample.

Citation Information

Patent Citations

  • Federal learning-based space-time prediction algorithm on industrial internet-of-things edge device

    CN114265913A

  • Multi-dimensional time sequence prediction method based on self-attention mechanism and graph convolutional network

    CN114818515A