Parallelization-based device state prediction method for cyclic network

By improving the GRU network architecture, combining GELU feedforward network and causal convolution, the problem of the traditional recurrent network being too long to process massive device timing data is solved, parallel training and prediction are realized, and the speed and accuracy of device state prediction are improved.

CN120277847AActive Publication Date: 2025-07-08CHINA YANGTZE POWER
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510328575.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-08
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

Traditional recurrent networks have too long calculation time when processing massive device timing data, which makes it difficult to meet the needs of real-time monitoring and rapid prediction, and the degree of parallelization is low, making it difficult to make full use of hardware resources.

Method used

By improving the GRU network architecture, combining GELU feedforward network, causal convolution and RMSNorm, parallel training and causal convolution are used to capture deep features and long-term information, and the ADAMW optimizer optimizes the model to design a parallelized recurrent network architecture.

Benefits of technology

It significantly improves the training speed and accuracy of device status prediction, can quickly complete model training and parameter optimization, meet the enterprise's rapid deployment and iteration needs, and effectively avoid equipment burst failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277847A_ABST
    Figure CN120277847A_ABST
Patent Text Reader

Abstract

The invention discloses an equipment state prediction method based on a parallelization loop network. The method comprises the following steps: S1, improving a GRU network architecture; s2, extracting an input time sequence deep feature mode by using a GELU feedforward network; s3, adding a noise module to realize consistency of parallel training and sequence prediction; s4, improving the efficiency and the stability of the model by using RMSNorm; s5, capturing long-period causal time sequence information by adopting causal convolution; s6, the modules are deepened, stacked and combined through residual connection; s7, performing optimization by using an ADAMW optimizer, superposition dynamics and prediction loss; s8, inputting historical and future control quantity information, outputting differential historical and future prediction, and calculating mean square error improvement performance; and S9, carrying out equipment state early warning according to a prediction result. Compared with a traditional LSTM model, the method has the advantages that the prediction error is reduced, the single-sequence reasoning time is greatly shortened, and a powerful solution is provided for accurate and efficient prediction of the state of the industrial equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of industrial equipment status prediction, and particularly relates to a method for predicting the status of equipment based on a parallelized recurrent network. Background Art

[0002] When industrial equipment is running, it will continuously generate a large amount of complex time-series data. For example, a wind turbine can generate tens of thousands of time-series records containing parameters such as rotational speed, temperature, and vibration every day. Although traditional recurrent networks (such as LSTM and GRU) have a certain processing ability for time-series data, their architectures have limitations. Their serial computing mode means that each step of calculation depends on the result of the previous step. As the amount of data increases, the calculation time will be significantly extended. When processing a device operation parameter sequence with a length of 10,000, an ordinary GRU model takes several minutes on a conventional configuration CPU. If the sequence length increases to 50,000, the time consumption will soar to more than half an hour, which far cannot meet the requirements of real-time monitoring and rapid prediction.

[0003] At the same time, the parallelization degree of traditional recurrent networks is low. In a multi-core CPU or GPU environment, due to the strong dependence characteristics between their calculation steps, it is difficult to make full use of the hardware parallel computing resources. Taking a certain industrial equipment fault prediction data set as an example, when equipped with an NVIDIA RTX 3090 GPU, the parallel acceleration ratio of the traditional LSTM model is only about 2.5, which is significantly different from the acceleration ratio of more than 10 of the convolutional neural network under the same hardware. This makes it difficult for traditional recurrent networks to meet the industrial actual application in terms of both training and prediction efficiency when facing a large amount of device data, and new technical solutions are urgently needed to improve the processing efficiency of a large amount of time-series data. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method for predicting the status of equipment based on a parallelized recurrent network, which can accurately extract deep features and long-cycle causal information in the time-series data of equipment operation through the synergistic effect of the GELU feed-forward network and causal convolution; effectively avoid production interruptions caused by sudden equipment failures.

[0005] To solve the above technical problems, the technical solution adopted by the present invention is:

[0006] A method for predicting the status of equipment based on a parallelized recurrent network, the steps are as follows:

[0007] S1. Improve the GRU network architecture;

[0008] S2: Use the GELU feed-forward network to extract the deep feature patterns of the input time series;

[0009] S3: Add a noise module to make the parallel training consistent with the sequential prediction;

[0010] S4: Improve the model efficiency and stability with RMSNorm;

[0011] S5: Adopt causal convolution to capture long - period causal temporal information;

[0012] S6: Use residual connections to deepen and stack each module;

[0013] S7: Use the ADAMW optimizer and stack the dynamics and prediction losses for optimization;

[0014] S8: Input historical and future control quantity information, output differential history and future predictions, and calculate the mean square error to improve performance;

[0015] S9: Conduct device status early warning based on the prediction results.

[0016] Preferably, the sub - steps of S1 are:

[0017] S1.1: Define h t The hidden state of the recurrent unit at time step t, and its calculation formula is:

[0018]

[0019] In the formula, z t is the update gate, used to control the combination ratio of the previous - moment hidden state h t-1 and the current candidate hidden state ⊙ is the element - wise multiplication operator, h t-1 is the hidden state of the recurrent unit at time step t - 1, is the candidate hidden state at the current time step t;

[0020] Among them,

[0021] where σ is the logistic sigmoid function, which maps the input value to the range between 0 and 1, is a linear transformation of the input x t with the output dimension of d h , and x t is the input data at time step t;

[0022] S1.2: Use Hesinsen Associative Scan to perform cumulative scanning in the log space to achieve high - speed parallel training and inference.

[0023] Preferably, the sub - steps of S1.2 are:

[0024] S1.2.1: Implement a recursive cumulative operation using the Heinsen Associative Scan Log module to process sequence data with logarithmic operations;

[0025] S1.2.2: Mathematical formula parsing:

[0026] Calculate a * : The sum of the logarithms of all elements of the logarithmic coefficient array log_coeffs, and the calculation formula is:

[0027]

[0028] If log_coeffs is already in logarithmic form, then it is

[0029] where log_coeffs is the logarithmic coefficient array, n is the number of elements in the logarithmic coefficient array log_coeffs, coeffs is the coefficient array, and when log_coeffs is not in logarithmic form, log_coeffs[i] = log(coeffs[i]);

[0030] Calculate the logarithmic weighted sum log(h0 + b * ): The logarithmic weighted sum calculated by combining the initial value h0 and the logarithmic values log_values, and the formula is:

[0031]

[0032] h0 is the initial value, which may be a constant or an initial state, and b * is the exponential translation summation term through log_values[i] - a * , that is, b * = ∑exp(log_values[i] - a * ), log_values is the logarithmic value array, and m is the number of elements in the logarithmic value array log_values;

[0033] S1.2.3: Confirm variable consistency, dimensionality issues, and clarify the initial value h0;

[0034] S1.2.4: Final output formula:

[0035] Formula 1: log(h) = a * + log(h0 + b * );

[0036] log(h): The final logarithmic output, which is the sum of the logarithms a *It consists of two parts: the logarithmically weighted sum log(h0 + b) obtained through exponential translation and summation operations in the second step * ).

[0037] Formula 2: h = exp(log(h));

[0038] h: The original value obtained by taking the exponential of the final result log(h) in the logarithmic domain, keeping the calculation in the numerically stable logarithmic domain and finally outputting the actual numerical value.

[0039] Preferably, the sub-steps of S2 are as follows:

[0040] S2.1: Feedforward network structure: Information transfer and non-linear transformation are achieved through the combination of two linear transformations and activation functions;

[0041] S2.2: Mathematical formula analysis:

[0042] S2.2.1: First step: Perform linear transformation 1:

[0043] h1: The result after the input x passes through linear transformation 1, and the calculation formula is h1 = W1·x + b1, and dim(h1) = dim_inner;

[0044] Where W1 is the weight matrix of linear transformation 1, with dimensions dim_inner×dim, x is the input data of the feedforward network, with dimensions dim, b1 is the bias of linear transformation 1, with dimensions dim_inner, dim is the dimension of the input data x, and dim_inner is the internal dimension, that is, the dimension of the output h1 of linear transformation 1;

[0045] S2.2.2: Second step: Perform non-linear transformation:

[0046] h2: The result after h1 passes through the activation function GELU transformation, and the formula is h2 = GELU(h1); where GELU is the Gaussian Error Linear Unit activation function, and its formula is GELU(x) = x·Φ(x), and Φ(x) is the cumulative distribution function of the standard normal distribution;

[0047] S2.2.3: Third step: Linear transformation 2:

[0048] y: The final output of the feedforward network, with the same dimension as the input x, which is dim, and the calculation formula is y = W2·h2 + b2; here W2 is the weight matrix of linear transformation 2, with dimensions dim×dim_inner, and b2 is the bias of linear transformation 2, with dimensions dim;

[0049] S2.3: It is specified that the dimension of the input x is dim, the dimension of the intermediate state h1 is dim_inner, and the dimension of the output y is restored to dim; GELU provides the non - linear ability to enable the network to learn complex patterns; the dimensions of the weight matrices W1 and W2 are dim_inner×dim and dim×dim_inner respectively.

[0050] Preferably, the sub - steps of S3 are:

[0051] S3.1: The BlockNoise1D module performs the operation of injecting noise into the input data in blocks.

[0052] S3.2: Conduct mathematical formula parsing:

[0053] S3.2.1: The first step: Generate a mask matrix

[0054] mask: The mask matrix, used to determine which blocks need to add noise, and the definition of its element mask[i, j] is

[0055] where i is the row index of the mask matrix and j is the column index of the mask matrix;

[0056] S3.2.2: The second step: Generate the noise for the input data:

[0057] noise i : The noise generated for the input data x i The calculation formula is

[0058] where is a normal distribution with a mean of 0 and a standard deviation of σ, σ is the standard deviation of the normal distribution, controlling the noise intensity, and x i is the i - th element of the input data;

[0059] S3.2.3: The third step: Perform the noise injection operation:

[0060] S3.3: Specify the definition of the noise blocks (such as time steps, feature dimensions, etc.), the noise standard deviation σ (which needs to be adjusted according to the task requirements), and the generation method of the mask matrix (possibly based on random sampling or fixed patterns);

[0061] S3.4: Apply the BlockNoise1D module to data augmentation, regularization, and robust training.

[0062] Preferably, the sub - steps of S4 are:

[0063] S4.1: Use RMSNorm to normalize the input vector and introduce a learnable scaling factor γ to balance the gradient scale and improve the network stability;

[0064] S4.2: Mathematical formula parsing:

[0065] S4.2.1: First step: Normalization:

[0066] The i-th element of the input x after normalization is calculated as where x i is the i-th element of the input vector x, and ‖x‖2 is the L2 norm of the input vector x.

[0067] S4.2.2: Second step: Scaling and adjustment:

[0068] y: The final output of the RMSNorm module, which is calculated as where γ is a learnable scaling factor used to adjust the normalized output, scale is determined by the square root of the input dimension (it may be a constant or a function based on the input dimension, such as ), and is used to further adjust the output scale, and dim is the dimension of the input vector x;

[0069] S4.3: It is specified that the normalization operation is performed on the last dimension and is applicable to multi-dimensional inputs (such as sequence data or feature vectors); the learnable parameter γ allows the model to dynamically adjust the normalized output according to the task requirements; if the formula for scale is not specified, it may be a constant or a function based on the input dimension.

[0070] Preferably, the sub-steps of S5 are:

[0071] S5: Use causal convolution to capture richer long-term causal temporal information:

[0072] S5.1: Use Causal Depthwise Convolution for temporal data, and ensure that the model can only access past information through depthwise separable convolution and causal padding;

[0073] S5.2: Mathematical formula parsing:

[0074] S5.2.1: First step: Causal padding:

[0075] The result of the input x after causal padding is calculated as where x is the input temporal data of the causal convolution module, pad is the padding function, and padding = (k - 1, 0) means padding k - 1 zeros at the start of the sequence, and k is the size of the convolution kernel;

[0076] S5.2.2: Second step: Depthwise separable convolution:

[0077] y: The output result after depthwise separable convolution, with the same dimension as the input x, and the calculation formula is where Convld is the depthwise separable convolution operation;

[0078] S5.3: It is clear that causal padding is used to ensure that the model can only access past information, which is applicable to time series data modeling (such as time series prediction, speech processing, etc.); depthwise separable convolution reduces the number of parameters and computational complexity while maintaining the non-linear ability of the convolution operation; padding k - 1 zeros ensures that the convolution kernel does not access future data points.

[0079] Preferably, the specific content of S8 is: Input all historical information for a period of time: the data size is L1, D;

[0080] Input future control quantity information: the size is L2, D - 1;

[0081] The output is the differential historical prediction of length L1 and the future prediction of length L2. The mean square error loss is calculated for both predicted values to improve performance;

[0082] where: L1, D: The feature dimension of the input historical information data;

[0083] L2: The length dimension of the future control quantity information data;

[0084] D - 1: The feature dimension of the future control quantity information data;

[0085] Preferably, the sub-steps of S9:

[0086] S9.1: Obtain measuring points such as the unit status, active power, outlet current, inlet and outlet temperatures of the main pipe of the air cooler cooling water, and stator winding temperature;

[0087] S9.2: Preprocess the data using a specified method;

[0088] S9.3: Design a new network architecture that can learn the operation rules of the device and adapt to the device dynamics differences brought by different measuring points. After the model training is completed, it can also show good results when applied to the devices of other units;

[0089] S9.4: Test the model performance: Perform short-term and long-term data prediction on the data of the target measuring point and conduct index comparison.

[0090] A device status prediction system based on a parallelized recurrent network adopts the device status prediction method based on a parallelized recurrent network described above.

[0091] The present invention can achieve the following beneficial effects:

[0092] 1. Through the synergistic effect of the GELU feed-forward network and causal convolution, deep features and long-term causal information in the device operation time-series data can be accurately extracted. In the scenario of motor fault prediction, the ability to capture early weak fault features is significantly enhanced, effectively avoiding production interruptions caused by sudden device failures.

[0093] 2. The improved GRU network architecture, combined with the parallel training and inference mechanism of Hesinsen Associative Scan, significantly reduces the training time. When processing large-scale device operation data, the training speed is 3 - 5 times faster than that of traditional recurrent networks, quickly completing model training and parameter optimization to meet the enterprise's needs for rapid model deployment and iteration. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] The present invention will be further described below in conjunction with the drawings and embodiments:

[0095] Figure 1 This is the flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0096] The preferred solution is as Figure 1 shown. A method for predicting the device state based on a parallelized recurrent network, the steps are as follows:

[0097] S1: Improve the GRU network architecture:

[0098] S1.1: Define the hidden state h of the recurrent unit at time step t t , and its calculation formula is

[0099] z t is the update gate, calculated through , where σ is the logistic sigmoid function that maps the input value to between 0 and 1, is a linear transformation of the input xt (the input data at time step t), and the output dimension is d h (set to 50). h t1 is the hidden state of the recurrent unit at time step t - 1, is the candidate hidden state at the current time step, and ⊙ is the element-wise multiplication operator.

[0100] S1.2: Sub-steps:

[0101] S1.2.1: Use the Heinsen Associative Scan Log module to perform a recursive cumulative operation on the sequence data containing logarithmic operations.

[0102] S1.2.2: Calculate a. If the logarithmic coefficient array log_coeffs is not in logarithmic form, its calculation formula is

[0103] If log_coeffs is already in logarithmic form, then

[0104] Here, n is set to 100 (the number of elements in the logarithmic coefficient array), coeffs is the coefficient array, and when log_coeffs is not in logarithmic form, log_coeffs[i] = log(coeffs[i]).

[0105] Next, calculate the logarithmically weighted sum log(h0 + b),

[0106] where h0 is set to 0.5 (the initial value, a constant), b is the exponential translation summation term of log_values[i]a, log_values is the logarithmic value array, and m is set to 200 (the number of elements in the logarithmic value array).

[0107] S1.2.3: Carefully confirm the variable consistency, check the dimension issues (such as the dimension matching between the logarithmic coefficient array and the logarithmic value array, etc.), and clarify the setting of the initial value h0.

[0108] S1.2.4: Obtain the final output formula. Formula 1 is log(h) = a + log(h0 + b), log(h) is the final logarithmic output, which consists of two parts: a^ and log(h0 + b); Formula 2 is h = exp(log(h)), h is the original value obtained by taking the exponent of the final result log(h) in the logarithmic domain, ensuring that the calculation is performed in a numerically stable logarithmic domain, and the final output is the actual value.

[0109] S2: Use the GELU feedforward network to extract the deep feature pattern of the input time series:

[0110] S2.1: Construct the feedforward network structure, and realize information transfer and non-linear transformation through the combination of two linear transformations and activation functions.

[0111] S2.2: The sub-steps are as follows:

[0112] S2.2.1: Perform linear transformation 1. The result h1 after the input x (with dimension dim, set to 20 dimensions) passes through linear transformation 1 is calculated by the formula h1 = W1·x + b1, and dim(h1) = dim_inner (set to 30 dimensions). Here, W1 is the weight matrix of linear transformation 1, with dimension dim_inner × dim (i.e., 30 × 20), and b1 is the bias of linear transformation 1, with dimension dim_inner (30 dimensions).

[0113] S2.2.2: Perform a non-linear transformation. The result h2 after the activation function GELU transforms h1. The formula is h2 = GELU(h1). Here, GELU is the Gaussian Error Linear Unit activation function, and its formula is GELU(x) = x·Φ(x), where Φ(x) is the cumulative distribution function of the standard normal distribution.

[0114] S2.2.3: Perform a linear transformation 2. The final output y of the feed-forward network has the same dimension as the input x, which is dim (20 dimensions). The calculation formula is y = W2·h2 + b2. Here, W2 is the weight matrix of the linear transformation 2, with a dimension of dim×dim_inner (20×30), and b2 is the bias of the linear transformation 2, with a dimension of dim (20 dimensions).

[0115] S2.3: It is clear that the dimension of the input x is dim (20 dimensions), the dimension of the intermediate state h1 is dim_inner (30 dimensions), and the dimension of the output y is restored to dim (20 dimensions). Recognize the non-linear ability provided by GELU, which enables the network to learn complex patterns, and the dimensions of the weight matrices W1 and W2 are 30×20 and 20×30 respectively.

[0116] S3: Noise addition module, achieving consistency between parallel training and sequential prediction

[0117] S3.1: Use the BlockNoise1D module to perform the operation of injecting noise into the input data in blocks. Here, the data is divided into blocks of every 10 time steps according to the time steps.

[0118] S3.2: Sub-steps:

[0119] S3.2.1: Generate a mask matrix mask. The definition of its element mask[i,j] is:

[0120]

[0121] where i is the row index of the mask matrix and j is the column index of the mask matrix.

[0122] S3.2.2: Generate the noise noise for the input data i , and the calculation formula is

[0123] where is a normal distribution with a mean of 0 and a standard deviation of σ (assumed to be 0.1), and x i is the i-th element of the input data.

[0124] S3.2.3: Perform the noise injection operation.

[0125] S3.3: Define the noise block as one block for every 10 time steps. Set the noise standard deviation σ to 0.1 (predetermined according to data characteristics and training effects), and use random sampling for the generation method of the mask matrix.

[0126] S3.4: Apply the BlockNoise1D module to data augmentation, regularization, and robustness training.

[0127] S4: Improve the model efficiency and stability using RMSNorm:

[0128] S4.1: Normalize the input vector (with a dimension of 20) using RMSNorm and introduce a learnable scaling factor γ to balance the gradient scale and improve the network stability.

[0129] S4.2: Sub - steps:

[0130] S4.2.1: Perform normalization. The i - th element of the input x after normalization The calculation formula is where xi is the i - th element of the input vector, is the L2 norm of the input vector.

[0131] S4.2.2: Perform scaling and adjustment. The final output y of the RMSNorm module, the calculation formula is where γ is set to 1.2 (learnable parameter, initial value), and scale is set to (dim is the dimension of the input vector, that is 20, and scale is ), which is used to further adjust the output scale.

[0132] S4.3: Specify that the normalization operation is performed on the last dimension, which is applicable to the current 20 - dimensional input vector. The learnable parameter γ allows the model to dynamically adjust the normalized output according to task requirements. If the formula for scale is not specified, in this embodiment, it is set as a function based on the input dimension

[0133] S5: Use causal convolution to capture long - period causal temporal information:

[0134] S5.1: Use Causal Depthwise Convolution for temporal data. Ensure that the model can only access past information through depth - separable convolution and causal padding.

[0135] S5.2: The sub - steps are:

[0136] S5.2.1: Perform causal padding. The result of the input x after causal padding The calculation formula is Where x is the input time series data of the causal convolution module, pad is the padding function, and k is set to 3 (the size of the convolution kernel).

[0137] S5.2.2: Perform depthwise separable convolution. The output result y after depthwise separable convolution has the same dimension as the input x, and the calculation formula is Where Convld is the depthwise separable convolution operation.

[0138] S5.3: Clearly ensure that the model can only access past information through causal padding, which is applicable to device operation time series data modeling. Depthwise separable convolution reduces the number of parameters and computational complexity while maintaining the non-linear ability of the convolution operation. Pad k1 zeros to ensure that the convolution kernel does not access future data points.

[0139] S6: Deepen and stack each module with residual connections:

[0140] Stack each module processed through the above steps in order, and directly connect the output of the previous layer to the input of the next layer through residual connections to ensure the effective transmission of information in the model and deepen the model structure. For example, perform a residual connection between the output after improving the GRU network and the input of the GELU feedforward network, enabling the model to learn more complex features.

[0141] S7: Use the ADAMW optimizer and stack the dynamics and prediction losses for optimization:

[0142] Use the ADAMW optimizer to train and optimize the model, set the learning rate to 0.001, and the weight decay to 0.01. In terms of the loss function, stack the dynamics loss and the prediction loss. The dynamics loss is calculated based on the physical laws of device operation, and the prediction loss uses the mean squared error loss to calculate the error between the model prediction value and the true value. By continuously adjusting the model parameters, minimize the stacked loss value to improve the model performance.

[0143] S8: Input historical and future control quantity information, output differential history and future prediction, calculate the mean squared error to improve performance. Input the prepared historical information data with dimensions L1(1000), D(20) and future control quantity information data with dimensions L2(10), D1(19) into the model. The model outputs a differential history prediction of length L1(1000) and a future prediction of L2(10). Calculate the mean squared error loss for both of these prediction values, and gradually reduce the mean squared error by continuously adjusting the model parameters to improve the model prediction performance.

[0144] S9: Perform device status warning based on the prediction results:

[0145] S9.1: Obtain the measured data of the unit status, active power, outlet current, inlet and outlet temperatures of the main pipe of the cooling water of the air cooler, stator winding temperature, etc.

[0146] S9.2: Preprocess the data using the normalization method to unify the data to the range of 0-1.

[0147] S9.3: Design a new network architecture including the above-mentioned modules. This network architecture can adapt to the equipment dynamics differences brought by different measurement points by learning the equipment operation rules. After training the current factory equipment, apply the model to the equipment of other units of the same type for testing.

[0148] S9.4: Test the model performance, perform short-term (next 1 hour) and long-term (next 1 day) data predictions on the data of the target measurement point, and conduct index comparisons. Use indicators such as root mean square error (RMSE) and mean absolute error (MAE) to evaluate the prediction accuracy of the model. For example, in the temperature prediction of a certain key measurement point, the RMSE of the short-term prediction is 0.5°C, and the MAE is 0.3°C; the RMSE of the long-term prediction is 1.2°C, and the MAE is 0.8°C. According to the preset threshold (such as the temperature anomaly threshold is the normal temperature ±2°C), when the predicted value exceeds the threshold range, send out an equipment status warning signal in a timely manner.

[0149] Through the above embodiments, the specific implementation process and effect of the equipment status prediction method based on the parallelized recurrent network in the actual equipment status prediction are demonstrated.

[0150] The above embodiments are only the preferred technical solutions of the present invention and should not be regarded as a limitation to the present invention. The protection scope of the present invention should be the technical solutions recorded in the claims, including the equivalent replacement solutions of the technical features in the technical solutions recorded in the claims. That is, the equivalent replacement improvements within this scope are also within the protection scope of the present invention.

Claims

1. A method for predicting device status based on a parallelized recurrent network, characterized in that It includes the following steps: S1. Improve the GRU network architecture; S2: Use the GELU feed-forward network to extract the deep feature patterns of the input time series; S3: Add a noise module to make the parallel training consistent with the sequential prediction; S4: Use RMSNorm to improve the model efficiency and stability; S5: Adopt causal convolution to capture the long-period causal time series information; S6: Use residual connections to deepen and stack and combine each module; S7: Use the ADAMW optimizer and superimpose the dynamics and prediction losses for optimization; S8: Input the historical and future control quantity information, output the differential history and future predictions, and calculate the mean square error to improve the performance; S9: Conduct device status early warning according to the prediction results.

2. The device state prediction method based on a parallelized recurrent network according to claim 1, wherein: The sub-steps of S1 are: S1.

1. Define h t The hidden state of the recurrent unit at time step t, and its calculation formula is as follows: where, z t is the update gate, which is used to control the combination ratio of the previous hidden state h t-1 and the current candidate hidden state , is the element-wise multiplication operator, h t-1 is the hidden state of the recurrent unit at time step t-1, is the candidate hidden state at the current time step t; Among them, where σ is the logical sigmoid function that maps the input value to between 0 and 1, is a linear transformation of the input x t with an output dimension of d h , and x t is the input data at time step t; S1.2: Use Hesinsen Associative Scan to perform cumulative scanning in the log space to achieve high-speed parallel training and inference.

3. The device state prediction method based on a parallelized recurrent network according to claim 2, wherein: The sub-steps of S1.2 are: S1.2.1: Use the Heinsen Associative Scan Log module to implement recursive cumulative operations and process sequence data containing logarithmic operations; S1.2.2: Mathematical formula analysis: Calculate a * : The cumulative sum of the logarithmic values of all elements in the logarithmic coefficient array log_coeffs. The calculation formula is: If log_coeffs is already in logarithmic form, then Where log_coeffs is an array of logarithmic coefficients, n is the number of elements in the array of logarithmic coefficients log_coeffs, coeffs is an array of coefficients, and when log_coeffs is not in logarithmic form, log_coeffs[i] = log(coeffs[i]); Calculate the logarithmically weighted sum log(h0 + b * ): The logarithmically weighted sum calculated by combining the initial value h0 and the logarithmic values log_values, with the formula: h0 is the initial value, which may be a constant or an initial state, b * is the exponential translation summation term obtained by log_values[i] - a * , that is, b * = ∑exp(log_values[i] - a * ), log_values is an array of logarithmic values, and m is the number of elements in the array of logarithmic values log_values; S1.2.3: Confirm the variable consistency, dimension issues, and clarify the initial value h0; S1.2.4: Final output formula: Formula 1: log(h) = a * + log(h0 + b * ) log(h): The final logarithmic output, which is the sum a of the logarithmic coefficients log_coeffs in the first step * and the logarithmically weighted sum log(h0 + b obtained by exponential translation and summation operations in the second step * ; consists of two parts Formula 2: h = exp(log(h)); h: The original value obtained by taking the exponent of the final result log(h) in the logarithmic domain, keeping the calculation in the numerically stable logarithmic domain, and finally outputting the actual numerical value.

4. The device state prediction method based on a parallelized recurrent network according to claim 1, wherein: The sub-steps of S2 are: S2.1: Feed-forward network structure: Realize information transfer and nonlinear transformation through the combination of two linear transformations and activation functions; S2.2: Mathematical formula analysis: S2.2.1: First step: Perform linear transformation 1: h1: The result after the input x passes through linear transformation 1, and the calculation formula is h1 = W1·x + b1, and dim(h1) = dim_inner; Where W1 is the weight matrix of linear transformation 1, with dimension dim_inner × dim, x is the input data of the feed-forward network, with dimension dim, b1 is the bias of linear transformation 1, with dimension dim_inner, dim is the dimension of the input data x, and dim_inner is the internal dimension, that is, the dimension of the output h1 of linear transformation 1; S2.2.2: Second step: Perform nonlinear transformation: h2: The result after h1 passes through the activation function GELU transformation, and the formula is h2 = GELU(h1); where GELU is the Gaussian Error Linear Unit activation function, and its formula is GELU(x) = x·Φ(x), and Φ(x) is the cumulative distribution function of the standard normal distribution; S2.2.3: Third step: Linear transformation 2: y: The final output of the feedforward network, with the same dimension as the input x, which is dim. The calculation formula is y = W2·h2 + b2; here, W2 is the weight matrix of the linear transformation 2, with the dimension of dim×dim_inner, and b2 is the bias of the linear transformation 2, with the dimension of dim. S2.3: It is clear that the dimension of the input x is dim, the dimension of the intermediate state h1 is dim_inner, and the dimension of the output y is restored to dim; GELU provides non - linear capabilities, enabling the network to learn complex patterns; the dimensions of the weight matrices W1 and W2 are dim_inner×dim and dim×dim_inner respectively.

5. A method for predicting device status based on a parallelized recurrent network according to claim 1, characterized in that: The sub - steps of S3 are: S3.1: Use the BlockNoise1D module to perform the operation of injecting noise into the input data in blocks. S3.2: Conduct mathematical formula analysis: S3.2.1: The first step: Generate a mask matrix mask: A mask matrix used to determine which blocks need to be added with noise. The definition of its element mask[i, j] is where i is the row index of the mask matrix and j is the column index of the mask matrix. S3.2.2: The second step: Generate the noise for the input data: noise i : The noise generated for the input data x i is calculated by the formula where is a normal distribution with a mean of 0 and a standard deviation of σ, where σ is the standard deviation of the normal distribution and controls the noise intensity, and x i is the i-th element of the input data; S3.2.3: The third step: Perform the noise injection operation: S3.3: Define the noise block, the noise standard deviation σ, and the generation method of the mask matrix. S3.4: Apply the BlockNoise1D module to data augmentation, regularization, and robust training.

6. The device state prediction method based on a parallelized recurrent network according to claim 1, characterized in that: The sub - steps of S4 are: S4.1: Use RMSNorm to normalize the input vector and introduce a learnable scaling factor γ to balance the gradient scale and improve the network stability. S4.2: Mathematical formula analysis: S4.2.1: The first step: Normalization: The i-th element of the input x after normalization, and the calculation formula is where x i is the i-th element of the input vector x, and ‖x‖2 is the L2 norm of the input vector x S4.2.2: The second step: Scaling and adjustment: y: The final output of the RMSNorm module, calculated as where γ is a learnable scaling factor used to adjust the normalized output, scale is determined by the square root of the input dimension, and dim is the dimension of the input vector x; S4.3: It is clear that the normalization operation is performed on the last dimension and is applicable to multi - dimensional inputs. The learning parameter γ allows the model to dynamically adjust the normalized output according to the task requirements.

7. A method for predicting device status based on a parallelized recurrent network according to claim 1, characterized in that: The sub - steps of S5 are: S5: Use causal convolution to capture richer long - period causal temporal information: S5.1: Use Causal Depthwise Convolution for temporal data. Ensure that the model can only access past information through depth - separable convolution and causal padding. S5.2: Mathematical formula analysis: S5.2.1: The first step: Causal padding: The result after causal padding of the input x, and the calculation formula is where x is the input time-series data of the causal convolution module, pad is the padding function, padding=(k - 1, 0) means padding k - 1 zeros at the start position of the sequence, and k is the size of the convolution kernel; S5.2.2: The second step: Depth - separable convolution: y: The output result after depthwise separable convolution, with the same dimension as the input x, and the calculation formula is where Convld is the depthwise separable convolution operation; S5.3: It is clear that ensuring the model can only access past information through causal padding; maintaining the non - linear capabilities of the convolution operation; Pad k - 1 zeros to ensure that the convolution kernel does not access future data points.

8. A method for predicting device status based on a parallelized recurrent network according to claim 1, characterized in that: The specific content of S8 is: Input all historical information for a period of time: The data size is L1, D; Input future control quantity information: The size is L2, D - 1; The output is the differential historical prediction of length L1 and the future prediction of length L2. Calculate the mean squared error loss for both predicted values to improve performance; where: L1, D: The feature dimension of the input historical information data; L2: The length dimension of the future control quantity information data; D - 1: The feature dimension of the future control quantity information data.

9. A method for predicting device status based on a parallelized recurrent network according to claim 1, characterized in that: The sub - steps of S9: S9.1: Obtain measuring points such as unit status, active power, outlet current, inlet and outlet temperatures of the main pipe of the cooling water of the air cooler, stator winding temperature, etc.; S9.2: Preprocess the data by specified methods; S9.3: Perform short-term and long-term data prediction on the data of the target measuring points and conduct index comparison.

10. A device state prediction system based on a parallelized recurrent network, characterized in that: An equipment status prediction method based on a parallelized recurrent network according to claim 1 is adopted.

Citation Information

Patent Citations

  • MRI segmentation method based on reinforcement learning multi-scale neural network

    CN111784652A

  • Training method of equipment comprehensive efficiency prediction model, storage medium and electronic equipment

    CN115130671A

  • Database index data anomaly prediction method based on gated convolution and graph attention

    CN118585936A

  • Multi-step time sequence prediction method of quantum bidirectional circulation network based on causal convolution

    CN119250115A

  • Short-term traffic flow prediction method based on causal gated-low-pass graph convolutional network

    US20240029556A1