An industrial soft measurement method and system based on a time prior guided attention mechanism

CN122434371BActive Publication Date: 2026-09-22HUNAN NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610894008.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-09-22
Estimated Expiration
2046-06-22

AI Technical Summary

Technical Problem

基于此,本发明提供了一种基于时间先验引导注意力机制的工业软测量方法及系统,以解决背景技术中所提到现有的工业质量指标测量方法在处理不规则采样数据时存在的动态特征提取不足、长距离时间依赖容易丢失,以及缺乏预测不确定性量化的问题

Benefits of technology

由上述技术方案可知,本发明提出的一种基于时间先验引导注意力机制的工业软测量方法及系统,其有益效果在于:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122434371B_ABST
    Figure CN122434371B_ABST
Patent Text Reader

Abstract

The application provides an industrial soft measurement method and system based on a time prior guide attention mechanism, which comprises the following steps: acquiring time series data sets of industrial process input variables and quality variables, and constructing a time interval matrix; generating survival scores and risk scores by using a Weibull distribution, and constructing a survival-risk dual time representation; after the dual time representation and value embedding and position embedding are fused, inputting the dual time representation into a TPGA encoder, generating a variational distribution of attention weights through a posterior inference network and a prior inference network, and dynamically adjusting a prior confidence degree by using the risk score; in the model training, minimizing a prediction error and a KL divergence loss; in the test stage, obtaining a mean value and a standard deviation of a prediction value by using Monte Carlo random sampling, and constructing a dynamic confidence interval. The application can accurately capture complex dynamic evolution under non-uniform sampling, retain key long-distance dependence, and provide prediction reliability evaluation, and is suitable for soft measurement of complex industrial processes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of soft measurement technology, and in particular to an industrial soft measurement method and system based on a time-prior guided attention mechanism. Background Technology

[0002] In the control of complex industrial processes, real-time monitoring of key quality variables is crucial for ensuring production. Limited by the physical bottlenecks of high maintenance costs and large response delays of hardware sensors, soft measurement techniques that utilize easily measurable variables to predict difficult-to-measure indicators have emerged. However, early methods based on traditional linear models such as principal component analysis (PCA) struggled to accurately characterize the inherent high-dimensional nonlinear features of industrial processes.

[0003] Thanks to advancements in deep learning theory, soft measurement methods based on deep neural networks (such as LSTM and Transformer) have made significant progress in capturing long-range dependencies. However, most existing methods are built on the ideal assumption of "strictly uniform data sampling." In real-world industrial scenarios, due to factors such as sensor malfunctions or transmission anomalies, the actual sampling interval is often non-uniform and dynamically changing. This irregular temporal fluctuation itself contains important physical information about the dynamic evolution of the process, but existing methods often treat it as noise or missing values, causing the model to fail to capture the true dynamic characteristics of the system.

[0004] To address the problem of irregular sampling, existing research mainly employs two approaches: resampling methods, which force irregular sequences into regular ones, inevitably leading to the loss of crucial dynamic details; and time-aware methods, which often rely on fixed mathematical mapping functions (such as exponential decay or Gaussian kernels) to introduce time interval decay weights, making it difficult to characterize complex non-monotonic dynamic evolution and easily resulting in excessive weakening of long-distance dependencies between samples. Furthermore, existing methods generally lack a quantitative assessment mechanism for the uncertainty of prediction results, outputting only a single point prediction, which cannot provide reliable decision-making references for high-risk industrial operations. Therefore, there is an urgent need to develop a novel industrial soft measurement framework that can adaptively handle non-uniform sampling, accurately balance long and short-distance temporal correlations, and possess prediction reliability assessment capabilities. Summary of the Invention

[0005] (a) Technical problems to be solved Based on this, the present invention provides an industrial soft measurement method and system based on a time prior guided attention mechanism to solve the problems mentioned in the background art of insufficient dynamic feature extraction, easy loss of long-distance time dependence, and lack of prediction uncertainty quantification in existing industrial quality indicator measurement methods when processing irregular sampled data.

[0006] (II) Technical Solution To achieve the above objectives, this invention provides an industrial soft measurement method based on a time-prior-guided attention mechanism, comprising: Step S1: Through mechanistic analysis and expert knowledge, select several key process variables that affect quality variables from the industrial process as input variables. After several consecutive irregular samplings of the input variables and corresponding quality variables, the time series dataset of the input variables and quality variables is obtained, denoted as . ;in, , , The number of times the sample was collected. Normalization preprocessing is performed on each variable in the dataset to eliminate the influence of units, resulting in a new dataset denoted as . And use the new dataset as the training set; Step S2: Construct a soft measurement prediction model, which includes, in sequence: a time prior module, a TPGA encoder, a fully connected layer, and an output layer; and its corresponding timestamp sequence As input to the time prior module, the output of the time prior module is obtained. ;Will Input the TPGA encoder to get ;Will Input the fully connected layer to generate predicted values ​​for quality variables. ;in, Describe the normalized time series data of the input variables; Step S3: Train the model using the training set and construct the loss function; Step S4: In the testing phase, the time series data of the input variable to be predicted is input into the trained model, and several independent Monte Carlo random sampling predictions are performed. The mean of the predicted values ​​of the quality variable is calculated as the final predicted value. A 95% dynamic confidence interval is constructed based on statistics to evaluate the reliability of the prediction.

[0007] Specifically, in step S2, when describing the processing flow of the time prior module, firstly, based on the timestamp sequence, calculate... The absolute sampling time difference between each sample is used to construct a time interval matrix. Then, based on the time interval matrix, the Weibull distribution is used to calculate the survival score reflecting the correlation strength and the risk score reflecting information uncertainty. The survival score and risk score are then concatenated to construct a dual survival-risk time representation. Finally, for the sequence The sample data is embedded with value features and standard locations, and then combined with a dual-time representation after linear projection dimensionality reduction. Element-wise addition and fusion are performed to generate features with irregular temporal awareness. .

[0008] Specifically, in step S2, based on the timestamp sequence, calculate The absolute sampling time difference between each sample is used to construct a time interval matrix; specifically, it includes: Step a1: Obtain sample sequences from the normalized preprocessed dataset. And the corresponding timestamp sequence composed of the actual physical sampling times. ;in, Indicates the first One sample; This indicates the sampling time corresponding to the sample, i.e., the timestamp; Step a2: Based on timestamp sequence Calculate any two samples and absolute sampling time difference between ;in, , Indicates sample timestamp, Indicates sample timestamp; Step a3: Based on the absolute sampling time difference among all samples, construct a system of size [size missing]. Time interval matrix Its expression is: .

[0009] Specifically, in step S2, based on the time interval matrix, the survival score reflecting the correlation strength and the risk score reflecting the information uncertainty are calculated using the Weibull distribution, and the survival score and risk score are concatenated to construct a dual time representation of survival and risk. Specifically, this includes: Step b1: Construct a system containing learnable time-scale parameters and shape parameters The Weibull distribution; in order to effectively capture long-distance time dependencies under non-uniform sampling intervals, constraints are used during model parameter initialization and optimization. This ensures that the Weibull distribution satisfies the heavy-tailed property, enabling the model to retain reasonable long-distance correlations. Step b2: Based on the elements in the time interval matrix Dynamically calculate samples using the Weibull survival function. and Survival score between This adaptively maps the time interval to the strength of correlation retention between samples, as shown in the following formula:

[0010] Step b3: Quantify the information uncertainty accumulated over time using the Weibull risk function, and nonlinearly normalize the theoretically unbounded risk rate using the sigmoid activation function to generate a risk score. The calculation formula is:

[0011] Step b4: Set the survival score Risk Score Tensor concatenation is performed along the feature channel dimension to generate tensors for time sampling points. Survival-risk dual temporal characterization features This achieves dual encoding of irregular sampling time information in industry, as shown in the following formula:

[0012] in, Indicates a splicing operation; Thus, a dual time representation is obtained. .

[0013] Specifically, in step S2, the sequence is... The sample data is embedded with value features and standard locations, and then combined with a dual-time representation after linear projection dimensionality reduction. Element-wise addition and fusion are performed to generate features with irregular temporal awareness. Specifically, this includes: Step c1: Input a linear transformation layer and project it onto the preset hidden layer dimension of the model, that is, by processing the samples... Perform a linear mapping to generate the corresponding value embedding feature vector. The expression is:

[0014] in, and These represent the learnable weight matrix and bias vector of the linear transformation layer, respectively. Step c2: Calculate the time sampling points based on the timestamp sequence. Standard position encoding generates position embedding feature vectors with dimensions consistent with the model's hidden layers. The expression is:

[0015] in, This represents the standard position coding function; Step c3: Analyze the survival-risk dual temporal characteristics of each sample. The data is fed into a linear transformation layer to obtain a dual time representation after linear transformation. The expression is:

[0016] in, Indicates a linear transformation layer; Steps c1 and c3 essentially perform linear mapping; their operation mechanisms are the same, the only difference being the objects they process. Step c4: Embed the values ​​into the feature vector Location embedding feature vector and dual-time representation after linear transformation Element-wise addition and fusion are performed to generate features that include physical priors with irregular time intervals. As shown below:

[0017] Then obtained .

[0018] Specifically, in step S2, the TPGA encoder includes TPGA attention encoding blocks, of which The TPGA attention encoding block consists of, in sequence: variational attention mechanism, residual connection and normalization, feedforward neural network, residual connection and normalization; among which, the feedforward neural network contains 2 hidden layers. When describing the processing flow of the TPGA attention encoding block, let its input be... First, a variational attention mechanism is introduced: for input features... Perform linear transformations to generate query matrices respectively. Key matrix AND-value matrix ; by calculating the first query vectors With the Key vectors The dot product generates unnormalized base attention weights. The expression is as follows:

[0019]

[0020]

[0021]

[0022] in, , , Each represents a trainable parameter matrix; , , Both represent bias terms; Indicates the dimension of the key vector; Basic attention weights In the input posterior inference network, the mean of the posterior distribution of the generative attention weights is calculated. with standard deviation As shown below:

[0023]

[0024] The posterior inference network comprises three multilayer perceptrons, each using... , , express; Indicates the activation function; Represents the natural exponential function; Represent the natural logarithm function; The first Key vectors The input is fed into the prior inference network to calculate the mean of the prior distribution of the attention weights. Compared with the baseline standard deviation As shown below:

[0025]

[0026] The prior inference network comprises three multilayer perceptrons, each using... , , express; Based on risk score Building dynamic confidence gating Adaptive adjustment of the basic prior standard deviation Generate the final time prior standard deviation As shown below:

[0027] During model training, based on the prior distribution parameters of the attention weights , With attention weight posterior distribution parameters , Calculate the KL divergence loss :

[0028] Based on the reparameterization technique, dynamic nonnormalized attention weights are generated by sampling from the posterior distribution of attention weights. As shown below:

[0029] in, This is standard normally distributed noise; Nonnormalized attention weights Normalize the sequence by performing a softmax activation function along the sequence dimension, and then compare it with the sequence's first (i.e., the first) line. Value vectors Perform a weighted summation to obtain the TPGA attention encoding block for the first... Variational attention features at each time step As shown below:

[0030] Secondly, variational attention features Compared with the original input features Residual stacking is performed, followed by layer normalization, to obtain As shown below:

[0031] Then, The input is fed into the first hidden layer of the feedforward neural network, where it undergoes linear upscaling projection to generate features for the intermediate hidden layers. And perform a nonlinear activation operation on it to obtain As shown below:

[0032] in, , These represent the weight matrix and bias vector of the first hidden layer, respectively. Will In the second hidden layer of the input feedforward neural network, linear dimensionality reduction projection is performed to generate the feedforward output features. :

[0033] in, , These represent the weight matrix and bias vector of the second hidden layer, respectively. Finally, the feedforward output features Input features of feedforward neural networks Residual stacking and layer normalization are performed to generate the output features of the TPGA attention coding block. As shown below:

[0034] TPGA encoders include A series of stacked TPGA attention encoding blocks; the input of the first TPGA attention encoding block is... The output is The input to the second TPGA attention encoding block is The output is And so on, the first The input to each TPGA attention encoding block is The output is ,in The final output of the TPGA encoder is: ,Right now ; Will Input the fully connected layer to generate predicted values ​​for quality variables. .

[0035] Specifically, in step S3, firstly, the prediction error loss function within the training batch is calculated, as shown below:

[0036] in, Indicates the training batch size; This represents the true value of the quality variable; Represents the predicted value of a quality variable; The total KL divergence loss function is calculated as follows:

[0037] in, This indicates the number of TPGA attention coding blocks in the TPGA encoder; Indicates the first KL divergence loss for each TPGA attention encoding block; Then the total loss function for:

[0038] in, This represents the hyperparameter that controls the strength of KL regularization; Then, the Adam optimizer is used to iteratively update all learnable parameters of the entire model based on the total loss function until the model converges.

[0039] Specifically, in step S4, During the model testing phase, obtain the input variable sample sequence. And keep the random sampling mechanism of the variational posterior distribution in the model after training convergence active; normalize the Input model execution Subindependent Monte Carlo stochastic forward propagation predictions, obtaining Predicted values ​​of each quality variable ,in ; through calculation The mean of the predicted values ​​is used as the final predicted value, by calculating... The uncertainty of the model is assessed by using the standard deviation of each predicted value. calculate The mean of the predicted values ​​of each quality variable is used to eliminate random fluctuations in a single prediction, resulting in the final predicted values ​​of the quality variables, as shown below:

[0040] based on Predicted values ​​and mean values ​​of each quality variable Calculate the standard deviation of the predicted values. The cognitive standard deviation of the quantification model under non-uniform sampling and complex working conditions is shown below:

[0041] Based on the mean of the predicted values with standard deviation Construct a 95% dynamic confidence interval for the final prediction result. This provides a basis for reliability assessment for operational decisions and risk warnings in industrial production systems, as shown below: .

[0042] On the other hand, the present invention also discloses an industrial soft measurement system based on a time-prior guided attention mechanism, comprising: at least one processor; and at least one memory communicatively connected to the processor, wherein: the memory stores program instructions executable by the processor, and the processor can execute the above-described method by calling the program instructions.

[0043] (III) Beneficial Effects As can be seen from the above technical solution, the industrial soft measurement method and system based on a time-prior-guided attention mechanism proposed in this invention have the following beneficial effects: 1. A survival-risk dual time representation mechanism based on the Weibull distribution is introduced, directly mapping irregular time intervals into learnable, physically meaningful high-dimensional prior features. Compared with existing methods that use fixed function decay or simple missing value imputation, this invention can more accurately capture the complex dynamic evolution patterns under non-uniform sampling, significantly improving the model's adaptability to irregular data and prediction accuracy.

[0044] 2. By deeply integrating variational inference with the attention mechanism, the confidence level of the prior distribution is dynamically adjusted using time risk scores. Over long time intervals, the model can adaptively use KL divergence constraints to allow the posterior distribution to "safely fall back" to a more conservative prior distribution. This effectively filters out long-distance noise interference while preserving key long-term dependencies, solving the problem of local spurious correlations or loss of long-distance dependencies that traditional attention mechanisms are prone to when dealing with irregular data.

[0045] 3. Breaking through the black box limitation of traditional soft measurement models that only output a single "point prediction", by combining multiple Monte Carlo random samplings during the testing phase, the model can dynamically calculate the cognitive standard deviation of the prediction results and generate confidence intervals. This not only provides operators with a reliable range of prediction values, but also provides valuable quantitative basis for industrial decision-making, risk warning and process safety assessment under complex and fluctuating operating conditions. Attached Figure Description

[0046] The features and advantages of the invention will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the invention in any way. In the drawings: Figure 1 This is a flowchart of the industrial soft measurement method based on the time prior guided attention mechanism of the present invention; Figure 2 This is a schematic diagram of the soft measurement prediction model of the present invention; Figure 3 This is a schematic diagram of the time prior module of the present invention; Figure 4 This is a schematic diagram illustrating the process of generating the prior and posterior distributions of attention weights in this invention. Figure 5 This is a curve comparing the predicted and actual C4 concentration values ​​of the butane removal tower in an embodiment of the present invention. Figure 6 This is a curve comparing the predicted and actual SO2 concentration values ​​of the sulfur recovery unit in an embodiment of the present invention. Figure 7 This is a confidence graph showing the predicted C4 concentration of the butane removal tower in an embodiment of the present invention. Figure 8 This is a confidence graph showing the predicted SO2 concentration of the sulfur recovery unit in an embodiment of the present invention. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0048] like Figure 1 As shown, this invention provides an industrial soft measurement method based on a time-prior-guided attention mechanism, comprising: Step S1: Through mechanistic analysis and expert knowledge, select several key process variables that affect quality variables from the industrial process as input variables. After several consecutive irregular samplings of the input variables and corresponding quality variables, the time series dataset of the input variables and quality variables is obtained, denoted as . ;in, , , The number of times the sample was collected. Normalization preprocessing is performed on each variable in the dataset to eliminate the influence of units, resulting in a new dataset denoted as . And use the new dataset as the training set; Step S2: Construct a soft measurement prediction model, such as Figure 2 As shown, the model consists of: a temporal prior module, a TPGA encoder (Temporal Prior Guided Attention Encoder), a fully connected layer, and an output layer; and its corresponding timestamp sequence As input to the time prior module, the output of the time prior module is obtained. ;Will Input the TPGA encoder to get ;Will Input the fully connected layer to generate predicted values ​​for quality variables. ;in, Describe the normalized time series data of the input variables; (1) When describing the processing flow of the time prior module, such as Figure 3 As shown, firstly, based on the timestamp sequence, calculate... The absolute sampling time difference between each sample is used to construct a time interval matrix; specifically, it includes: Step a1: Obtain sample sequences from the normalized preprocessed dataset. And the corresponding timestamp sequence composed of the actual physical sampling times. ;in, Indicates the first One sample; This indicates the sampling time (i.e., timestamp) corresponding to the sample. Step a2: Based on timestamp sequence Calculate any two samples and absolute sampling time difference between ;in, , Indicates sample timestamp, Indicates sample timestamp; Step a3: Based on the absolute sampling time difference among all samples, construct a system of size [size missing]. Time interval matrix Its expression is: (1) Secondly, based on the time interval matrix, the survival score reflecting the strength of the correlation and the risk score reflecting the uncertainty of information are calculated using the Weibull distribution. The two scores (survival score and risk score) are then concatenated to construct a dual time representation of survival and risk. Specifically, this includes: Step b1: Construct a system containing learnable time-scale parameters and shape parameters The Weibull distribution; in order to effectively capture long-distance time dependencies under non-uniform sampling intervals, constraints are used during model parameter initialization and optimization. This ensures that the Weibull distribution satisfies the heavy-tailed property, enabling the model to retain reasonable long-distance correlations. The Weibull distribution is used to characterize the dynamic evolution of industrial data (i.e., variables) over time. As a mapping mechanism, it uses its survival function and risk function to process the time difference in the time interval matrix, transforming it into a dual time representation feature that reflects the strength of correlation retention between samples (survival score) and information uncertainty (risk score).

[0049] Step b2: Based on the elements in the time interval matrix Dynamically calculate samples using the Weibull survival function. and Survival score between This adaptively maps the time interval to the strength of correlation retention between samples, as shown in the following formula: (2) Step b3: Quantify the information uncertainty accumulated over time using the Weibull risk function, and nonlinearly normalize the theoretically unbounded risk rate using the sigmoid activation function to generate a risk score. The calculation formula is: (3) Step b4: Set the survival score Risk Score Tensor concatenation is performed along the feature channel dimension to generate tensors for time sampling points. Survival-risk dual temporal characterization features This achieves dual encoding of irregular sampling time information in industry, as shown in the following formula: (4) in, Indicates a splicing operation; Thus, a dual time representation is obtained. .

[0050] For example, if Then the survival score matrix Risk score matrix , , , ; where matrix subscripts (e.g. subscript () represents the dimension of the matrix.

[0051] Processing the time interval matrix using a learnable Weibull distribution transforms the raw time difference values ​​into features that reflect both the "information relevance strength" and the "information uncertainty." and This allows the model to "understand" the temporal patterns of irregular sampling.

[0052] Finally, for the sequence The sample data is embedded with value features and standard locations, and then combined with a dual-time representation after linear projection dimensionality reduction. Element-wise addition and fusion are performed to generate features with irregular temporal awareness. Specifically, this includes: Step c1: Input a linear transformation layer and project it onto the preset hidden layer dimension of the model, that is, by processing the samples... Perform a linear mapping to generate the corresponding value embedding feature vector. The expression is: (5) in, and These represent the learnable weight matrix and bias vector of the linear transformation layer, respectively.

[0053] Step c2: Calculate the time sampling points based on the timestamp sequence. Standard position encoding generates position embedding feature vectors with dimensions consistent with the model's hidden layers. The expression is: (6) in, This represents the standard positional encoding function.

[0054] Step c3: Analyze the survival-risk dual temporal characteristics of each sample. The data is fed into a linear transformation layer to obtain a dual time representation after linear transformation. The expression is: (7) in, This represents a linear transformation layer.

[0055] It is worth noting that steps c1 and c3 are essentially performing linear mappings, and their operation mechanisms are the same; the only difference is the object they process.

[0056] Step c4: Embed the values ​​into the feature vector Location embedding feature vector and dual-time representation after linear transformation Element-wise addition and fusion are performed to generate features that include physical priors with irregular time intervals. As shown below: (8) Then obtained .

[0057] (2) The TPGA encoder includes The TPGA attention encoding block consists of: variational attention mechanism (TPGA Attention), residual connection and normalization (Add&Norm), feedforward neural network, and residual connection and normalization; the feedforward neural network contains two hidden layers.

[0058] When describing the processing flow of the TPGA attention encoding block, let its input be... First, as Figure 4 As shown, a variational attention mechanism is introduced: for input features Perform linear transformations to generate query matrices respectively. Key matrix AND-value matrix ; by calculating the first query vectors With the Key vectors The dot product generates unnormalized base attention weights. The expression is as follows: (9) (10) (11) (12) in, , , Each represents a trainable parameter matrix; , , Both represent bias terms; This represents the dimension of the key vector.

[0059] Basic attention weights In the input posterior inference network, the mean of the posterior distribution of the generative attention weights is calculated. with standard deviation As shown below: (13) (14) The posterior inference network comprises three multilayer perceptrons, each using... , , express; Indicates the activation function; Represents the natural exponential function; This represents the natural logarithm function.

[0060] The first Key vectors The input is fed into the prior inference network to calculate the mean of the prior distribution of the attention weights. Compared with the baseline standard deviation As shown below: (15) (16) The prior inference network comprises three multilayer perceptrons, each using... , , express.

[0061] Based on risk score Building dynamic confidence gating Adaptive adjustment of the basic prior standard deviation Generate the final time prior standard deviation As shown below: (17) When the sampling time interval is long and the risk score is high, the gate value increases and the prior standard deviation decreases, making the prior distribution more concentrated and conservative.

[0062] During model training, based on the prior distribution parameters of the attention weights , With attention weight posterior distribution parameters , Calculate the KL divergence loss : (18) Based on the reparameterization technique, dynamic nonnormalized attention weights are generated by sampling from the posterior distribution of attention weights. As shown below: (19) in, This is standard normally distributed noise.

[0063] Nonnormalized attention weights Normalize the sequence by performing a softmax activation function along the sequence dimension, and then compare it with the sequence's first (i.e., the first) line. Value vectors Perform a weighted summation to obtain the TPGA attention encoding block for the first... Variational attention features at each time step As shown below: (20) Secondly, variational attention features Compared with the original input features Residual stacking is performed, followed by layer normalization, to obtain As shown below: (twenty one) Then, The input is fed into the first hidden layer of the feedforward neural network, where it undergoes linear upscaling projection to generate features for the intermediate hidden layers. And perform a nonlinear activation operation on it to obtain As shown below: (twenty two) in, , These represent the weight matrix and bias vector of the first hidden layer, respectively.

[0064] Will In the second hidden layer of the input feedforward neural network, linear dimensionality reduction projection is performed to generate the feedforward output features. : (twenty three) in, , These represent the weight matrix and bias vector of the second hidden layer, respectively.

[0065] Finally, the feedforward output features Input features of feedforward neural networks Residual stacking and layer normalization are performed to generate the output features of the TPGA attention coding block. As shown below: (twenty four) TPGA encoders include A series of stacked TPGA attention encoding blocks; the input of the first TPGA attention encoding block is... The output is The input to the second TPGA attention encoding block is The output is And so on, the first The input to each TPGA attention encoding block is The output is The final output of the TPGA encoder is: (Right now ).

[0066] By using stacked TPGA attention encoding blocks, deeper temporal dependency features are extracted layer by layer. In this embodiment, .

[0067] (3) Input the fully connected layer to generate predicted values ​​for quality variables. .

[0068] Step S3: Train the model using the training set and construct the loss function; First, calculate the prediction error loss function within the training batch, as shown below: (25) in, Indicates the training batch size; This represents the true value of the quality variable; This represents the predicted value of a quality variable.

[0069] The total KL divergence loss function is calculated as follows: (26) in, This indicates the number of TPGA attention coding blocks in the TPGA encoder; Indicates the first KL divergence loss for each TPGA attention encoding block.

[0070] Then the total loss function for: (27) in, This represents the hyperparameter that controls the strength of KL regularization.

[0071] Then, the Adam optimizer is used to iteratively update all learnable parameters of the entire model based on the total loss function until the model converges.

[0072] Step S4: In the testing phase, the time-series data of the input variable to be predicted is input into the trained model, and several independent Monte Carlo random sampling predictions are performed. The mean of the predicted values ​​of the quality variable is calculated as the final predicted value. A 95% dynamic confidence interval is constructed based on statistics to assess the reliability of the prediction; specifically including: During the model testing phase, obtain the input variable sample sequence. And keep the random sampling mechanism of the variational posterior distribution in the model after training convergence active; normalize the Input model execution Subindependent Monte Carlo stochastic forward propagation predictions, obtaining Predicted values ​​of each quality variable ,in .

[0073] Existing deep learning models typically produce deterministic outputs during testing (i.e., inputting data always results in the same fixed predicted value). Therefore, existing models disable dropout or random sampling mechanisms during testing. However, the soft measurement model for high-risk industrial environments in this invention provides both predicted values ​​and confidence intervals. Therefore, during testing, the random sampling mechanism of the variational posterior distribution in the model after training convergence remains active (i.e., random sampling of the variational distribution is not disabled, preserving the standard normal distribution noise in formula (19)). Thus, for the same test sample, each forward propagation of the model will output a slightly different result due to the interference of random noise. The model is then run independently. The Monte Carlo random sampling (i.e., sampling twice) is calculated... The mean of each different outcome is used as the final predicted value, and the uncertainty of the model is assessed by calculating the standard deviation of each outcome.

[0074] calculate The mean of the predicted values ​​of each quality variable is used to eliminate random fluctuations in a single prediction, resulting in the final predicted values ​​of the quality variables, as shown below: (28) based on Predicted values ​​and mean values ​​of each quality variable Calculate the standard deviation of the predicted values. The cognitive standard deviation of the quantification model under non-uniform sampling and complex working conditions is shown below: (29) Based on the mean of the predicted values with standard deviation Construct a 95% dynamic confidence interval for the final prediction result. This provides a basis for reliability assessment for operational decisions and risk warnings in industrial production systems, as shown below: (30) This embodiment was validated on two different industrial process simulation datasets: a butanizer dataset (predicting C4 concentration) and a sulfur recovery unit dataset (predicting SO2 concentration). The experimental results are as follows: Figure 5-8 As shown.

[0075] Figure 5 The model outputs a curve comparing the predicted and actual C4 concentration values ​​of the bottom product in a butanizer column after inputting irregularly sampled operating variables (such as top temperature, top pressure, reflux flow rate, and bottom temperature). The curves show the fitting of the predicted and actual values ​​of the bottom product C4 concentration. The horizontal axis represents the input sample number, the vertical axis represents the C4 concentration value, the blue curve represents the actual value, and the red curve represents the predicted value. Figure 6 The figures show the fitting curves (horizontal axis: input sample number; vertical axis: SO2 concentration value) of the predicted sulfur dioxide concentration in the tail gas output by the model to the actual value after inputting irregular sampling operation variables (feed gas flow rate, first air pipe flow rate, second air pipe flow rate, sulfur-containing zone gas flow rate, and sulfur-containing zone air flow rate) in the sulfur recovery unit. Both figures visually demonstrate the high-precision prediction capability of the model of this invention when dealing with non-uniform sampling and missing data.

[0076] Figure 7 and Figure 8 This further demonstrates the 95% dynamic confidence interval generated by the model (95%). The shaded areas in the figure precisely cover the dynamic fluctuation range of the true values. These confidence intervals indicate that the model can adaptively output the cognitive standard deviation for risk warning when facing complex and fluctuating operating conditions, demonstrating extremely strong on-site reliability and completely breaking the limitation of traditional black-box models that can only output a single predicted value.

[0077] The present invention also discloses an industrial soft measurement system based on a time prior guided attention mechanism, comprising: at least one processor; and at least one memory communicatively connected to the processor, wherein: the memory stores program instructions executable by the processor, and the processor can execute the above-described method by calling the program instructions.

[0078] Finally, it should be noted that the above methods can be converted into software program instructions, which can be implemented using a control system including a processor and memory, or by computer instructions stored in a non-transitory computer-readable storage medium. The integrated unit implemented as a software functional unit can be stored in a computer-readable storage medium. This software functional unit, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: a USB flash drive, a portable hard drive, a read-only memory (ROM). ROM, Random Access Memory (RAM), magnetic disks, optical disks, and other media that can store program code.

[0079] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. An industrial soft measurement method based on a time-prior-guided attention mechanism, characterized in that, include: Step S1: Through mechanistic analysis and expert knowledge, select several key process variables that affect quality variables from the industrial process as input variables. After several consecutive irregular samplings of the input variables and corresponding quality variables, the time series dataset of the input variables and quality variables is obtained, denoted as... ;in, , , The number of times the sample was collected. Normalization preprocessing is performed on each variable in the dataset to eliminate the influence of units, resulting in a new dataset denoted as . And use the new dataset as the training set; Step S2: Construct a soft measurement prediction model, which includes, in sequence: a time prior module, a TPGA encoder, a fully connected layer, and an output layer; and its corresponding timestamp sequence As input to the time prior module, the output of the time prior module is obtained. ;Will Input the TPGA encoder to get ;Will Input the fully connected layer to generate predicted values ​​for quality variables. ;in, Describe the normalized time series data of the input variables; When describing the processing flow of the time prior module, firstly, based on the timestamp sequence, calculate... The absolute sampling time difference between each sample is used to construct a time interval matrix. Then, based on the time interval matrix, the Weibull distribution is used to calculate the survival score reflecting the correlation strength and the risk score reflecting information uncertainty. The survival score and risk score are then concatenated to construct a dual survival-risk time representation. Finally, for the sequence The sample data is embedded with value features and standard locations, and then combined with a dual-time representation after linear projection dimensionality reduction. Element-wise addition and fusion are performed to generate features with irregular temporal awareness. ; Step S3: Train the model using the training set and construct the loss function; Step S4: In the testing phase, the time series data of the input variable to be predicted is input into the trained model, and several independent Monte Carlo random sampling predictions are performed. The mean of the predicted values ​​of the quality variable is calculated as the final predicted value. A 95% dynamic confidence interval is constructed based on statistics to evaluate the reliability of the prediction.

2. The method according to claim 1, characterized in that, In step S2, based on the timestamp sequence, calculate The absolute sampling time difference between each sample is used to construct a time interval matrix; Specifically, it includes: Step a1: Obtain sample sequences from the normalized preprocessed dataset. And the corresponding timestamp sequence composed of the actual physical sampling times. ;in, Indicates the first One sample; This indicates the sampling time corresponding to the sample, i.e., the timestamp; Step a2: Based on timestamp sequence Calculate any two samples and absolute sampling time difference between ;in, , Indicates sample timestamp, Indicates sample timestamp; Step a3: Based on the absolute sampling time difference among all samples, construct a system of size [size missing]. Time interval matrix Its expression is: 。 3. The method according to claim 2, characterized in that, In step S2, based on the time interval matrix, the survival score reflecting the correlation strength and the risk score reflecting the information uncertainty are calculated using the Weibull distribution. The survival score and the risk score are then concatenated to construct a dual time representation of survival and risk. Specifically, it includes: Step b1: Construct a system containing learnable time-scale parameters and shape parameters The Weibull distribution; in order to effectively capture long-distance time dependencies under non-uniform sampling intervals, constraints are used during model parameter initialization and optimization. This ensures that the Weibull distribution satisfies the heavy-tailed property, enabling the model to retain reasonable long-distance correlations. Step b2: Based on the elements in the time interval matrix Dynamically calculate samples using the Weibull survival function. and Survival score between This adaptively maps the time interval to the strength of correlation retention between samples, as shown in the following formula: Step b3: Quantify the information uncertainty accumulated over time using the Weibull risk function, and nonlinearly normalize the theoretically unbounded risk rate using the sigmoid activation function to generate a risk score. The calculation formula is: Step b4: Set survival score Risk Score Tensor concatenation is performed along the feature channel dimension to generate tensors for time sampling points. Survival-risk dual temporal characterization features This achieves dual encoding of irregular sampling time information in industry, as shown in the following formula: in, Indicates a splicing operation; Thus, a dual time representation is obtained. .

4. The method according to claim 3, characterized in that, In step S2, the sequence The sample data is embedded with value features and standard locations, and then combined with a dual-time representation after linear projection dimensionality reduction. Element-wise addition and fusion are performed to generate features with irregular temporal awareness. Specifically, it includes: Step c1: Input a linear transformation layer and project it onto the preset hidden layer dimension of the model, that is, by processing the samples... Perform a linear mapping to generate the corresponding value embedding feature vector. The expression is: in, and These represent the learnable weight matrix and bias vector of the linear transformation layer, respectively. Step c2: Calculate the time sampling points based on the timestamp sequence. Standard position encoding generates position embedding feature vectors with dimensions consistent with the model's hidden layers. The expression is: in, This represents the standard position coding function; Step c3: Analyze the survival-risk dual temporal characteristics of each sample. The data is fed into a linear transformation layer to obtain a dual time representation after linear transformation. The expression is: in, Indicates a linear transformation layer; Steps c1 and c3 essentially perform linear mapping; their operation mechanisms are the same, the only difference being the objects they process. Step c4: Embed the values ​​into the feature vector Location embedding feature vector and dual-time representation after linear transformation Element-wise addition and fusion are performed to generate features that include physical priors with irregular time intervals. As shown below: Then obtained .

5. The method according to claim 4, characterized in that, In step S2, the TPGA encoder includes TPGA attention encoding blocks, of which ; The TPGA attention encoding block consists of, in sequence: variational attention mechanism, residual connection and normalization, feedforward neural network, residual connection and normalization; wherein, the feedforward neural network contains 2 hidden layers; When describing the processing flow of the TPGA attention encoding block, let its input be... First, a variational attention mechanism is introduced: for input features... Perform linear transformations to generate query matrices respectively. Key matrix AND-value matrix ; by calculating the first query vectors With the Key vectors The dot product generates unnormalized base attention weights. The expression is as follows: in, , , Each represents a trainable parameter matrix; , , Both represent bias terms; Indicates the dimension of the key vector; Basic attention weights In the input posterior inference network, the mean of the posterior distribution of the generative attention weights is calculated. with standard deviation As shown below: The posterior inference network comprises three multilayer perceptrons, each using... , , express; Indicates the activation function; Represents the natural exponential function; Represents the natural logarithm function; The first Key vectors The input is fed into the prior inference network to calculate the mean of the prior distribution of the attention weights. Compared with the baseline standard deviation As shown below: The prior inference network comprises three multilayer perceptrons, each using... , , express; Based on risk score Building dynamic confidence gating Adaptive adjustment of the basic prior standard deviation Generate the final time prior standard deviation As shown below: During model training, based on the prior distribution parameters of the attention weights , With attention weight posterior distribution parameters , Calculate the KL divergence loss : Based on the reparameterization technique, dynamic nonnormalized attention weights are generated by sampling from the posterior distribution of attention weights. As shown below: in, This is standard normally distributed noise; Nonnormalized attention weights Normalize the sequence by performing a softmax activation function along the sequence dimension, and then compare it with the sequence's first (i.e., the first) line. Value vectors Perform a weighted summation to obtain the TPGA attention encoding block for the first... Variational attention features at each time step As shown below: Secondly, variational attention features Compared with the original input features Residuals are stacked and then normalized to obtain... As shown below: Then, The input is fed into the first hidden layer of the feedforward neural network, where it undergoes linear upscaling projection to generate features for the intermediate hidden layers. And perform a nonlinear activation operation on it to obtain As shown below: in, , These represent the weight matrix and bias vector of the first hidden layer, respectively. Will In the second hidden layer of the input feedforward neural network, linear dimensionality reduction projection is performed to generate the feedforward output features. : in, , These represent the weight matrix and bias vector of the second hidden layer, respectively. Finally, the feedforward output features Input features of feedforward neural networks Residual stacking and layer normalization are performed to generate the output features of the TPGA attention coding block. As shown below: TPGA encoders include A series of stacked TPGA attention encoding blocks; the input of the first TPGA attention encoding block is... The output is The input to the second TPGA attention encoding block is The output is And so on, the first The input to each TPGA attention encoding block is The output is ,in The final output of the TPGA encoder is: ,Right now ; Will Input the fully connected layer to generate predicted values ​​for quality variables. .

6. The method according to claim 5, characterized in that, In step S3, firstly, the prediction error loss function within the training batch is calculated, as shown below: in, Indicates the training batch size; This represents the true value of the quality variable; Represents the predicted value of a quality variable; The total KL divergence loss function is calculated as follows: in, This indicates the number of TPGA attention coding blocks in the TPGA encoder; Indicates the first KL divergence loss for each TPGA attention encoding block; Then the total loss function for: in, This represents the hyperparameter that controls the strength of KL regularization; Then, the Adam optimizer is used to iteratively update all learnable parameters of the entire model based on the total loss function until the model converges.

7. The method according to claim 6, characterized in that, In step S4, during the model testing phase, the input variable sample sequence is obtained. And keep the random sampling mechanism of the variational posterior distribution in the model after training convergence active; normalize the Input model execution Subindependent Monte Carlo stochastic forward propagation predictions, obtaining Predicted values ​​of each quality variable ,in ; through calculation The mean of the predicted values ​​is used as the final predicted value, by calculating... The uncertainty of the model is assessed by using the standard deviation of each predicted value. calculate The mean of the predicted values ​​of each quality variable is used to eliminate random fluctuations in a single prediction, resulting in the final predicted values ​​of the quality variables, as shown below: based on Predicted values ​​and mean values ​​of each quality variable Calculate the standard deviation of the predicted values. The cognitive standard deviation of the quantification model under non-uniform sampling and complex working conditions is shown below: Based on the mean of the predicted values with standard deviation Construct a 95% dynamic confidence interval for the final prediction result. This provides a basis for reliability assessment for operational decisions and risk warnings in industrial production systems, as shown below: 。 8. An industrial soft measurement system based on a time-prior-guided attention mechanism, characterized in that, include: At least one processor; And at least one memory communicatively connected to the processor, wherein: the memory stores program instructions executable by the processor, and the processor invokes the program instructions to perform the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Personalized short message verification code pushing method based on multi-task learning

    CN121968031A

  • Multimodal scenario risk determination method based on generative ai large language model

    WO2025185005A1