Backpressure curve prediction system and method based on deep learning

Through deep learning technology, combined with Transformer, LSTM and ResNet models, the traditional backpressure curve prediction method has solved the accuracy and real-time shortcomings, and achieved high accuracy prediction of backpressure curves and effective identification of potential problems, ensuring the safe operation of the energy system.

CN120069570APending Publication Date: 2025-05-30INNER MONGOLIA JINGNING THERMAL POWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311611055.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-29
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The traditional backpressure curve prediction method relies on manual experience and simple sensor monitoring, and there are problems of insufficient accuracy and real-time performance, making it difficult to comprehensively and comprehensively evaluate the overall state of the backpressure curve.

Method used

A deep learning-based system is adopted to collect the image and signal data of the system operation, and use Transformer to combine with the LSTM model to predict the backpressure curve time, and predict the backpressure curve under different influencing factors through an improved ResNet network.

Benefits of technology

It realizes highly accurate prediction and judgment of the backpressure curve, improves the ability to identify potential problems and real-time early warning, and ensures the safe operation of the energy system and environmental protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069570A_ABST
    Figure CN120069570A_ABST
Patent Text Reader

Abstract

The invention provides a back pressure curve prediction system and method based on deep learning, and belongs to the technical field of thermal power plant safety monitoring, and the method comprises the steps: enabling a model to better capture abstract features and complex modes through introducing a deeper network structure, improving the sensitivity of subtle changes of a back pressure curve, and achieving the more precise prediction; secondly, attention to different positions is enhanced through a multi-attention mechanism, the model can capture key information more flexibly, and robustness and generalization ability are improved; and finally, components such as a residual block, global average pooling, a convolutional layer and an activation function are integrated, spatial features and a nonlinear relation of an input sequence are learned through organic combination, the adaptability to different back pressure curve modes is improved, and the performance is more excellent. And an operator is guided to carry out real-time adjustment, so that the system is expected to operate in the optimal operation state, and the stability and economy of system operation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of safety monitoring in thermal power plants, and particularly to a backpressure curve prediction system and method based on deep learning. Background Art

[0002] With the rapid development of technology, the technology of the energy system is also constantly advancing. In order to improve the operation efficiency, reduce the operation cost, and ensure safe and stable operation, the backpressure curve prediction technology has gradually become an indispensable part. The backpressure curve is crucial for the operation of the energy system. However, due to the complex processes and equipment inside the system, as well as the extreme environmental conditions such as high temperature and high pressure, the backpressure curve prediction faces certain challenges.

[0003] Traditional prediction methods mainly rely on manual experience and simple sensor monitoring, and these methods have obvious limitations in terms of accuracy and real-time performance. Manual experience requires a large amount of manpower and time, and is easily affected by subjective factors. While traditional sensor monitoring can only focus on certain specific parameters and it is difficult to comprehensively and integrally evaluate the overall state of the backpressure curve.

[0004] In this innovative model, we introduce the backpressure curve prediction technology. By collecting the image and signal data of the system operation and using deep learning algorithms for analysis and processing, we have achieved highly accurate prediction and judgment of the backpressure curve. This technology not only improves the overall accuracy of the backpressure curve prediction, but also strengthens the ability to identify potential problems and real-time early warning. This is of great significance for ensuring the safe operation of the energy system and environmental protection.

[0005] The application of the backpressure curve prediction model will greatly improve the management efficiency of the energy system and enhance the overall safety and reliability. The comprehensiveness, self-adaptability and global nature of the model make it a key tool to ensure the stability of the backpressure curve. This technology not only has the potential to solve the limitations of traditional prediction methods, but also will provide strong support for the intelligent and sustainable development of the energy system, and has broad application prospects and practical value. Summary of the Invention

[0006] In order to accurately predict the backpressure curve of the air-cooled system in a thermal power plant, improve the operation efficiency of the air-cooled system, meet the system operation under the optimal backpressure control condition, thereby reducing the coal consumption of the unit and achieving the requirements of green energy conservation and economy, the present invention provides a backpressure curve prediction system and method based on deep learning.

[0007] In the first aspect, the present invention provides a backpressure curve prediction system based on deep learning to achieve accurate and efficient prediction of the backpressure curve. The device includes the following parts: a data acquisition unit, a data processing unit, a backpressure curve time prediction module, and a backpressure curve influencing factor prediction module.

[0008] Data acquisition unit: The data acquisition unit is equipped with a temperature measurement customized cable, an infrared temperature measurement system, a tube bundle deformation detection system, data acquisition equipment, and a DCS control system. It collects data through the acquisition equipment and extracts and uploads the collected data. The acquisition parameters include temperature, pressure, and deformation data.

[0009] Data processing unit: It cleans and normalizes the collected data to obtain the required data including ambient temperature, unit load, and atmospheric pressure data, and divides the dataset into a test set and a training set.

[0010] Backpressure curve time prediction module: This module is one of the core modules of the present invention. It includes a trained neural network model. In the present invention, a Transformer combined with an LSTM model is adopted.

[0011] The advantages of combining the Transformer and LSTM models in air-cooled backpressure prediction are mainly reflected in two aspects. First, the self-attention mechanism introduced by the Transformer enables the model to better focus on the key parts in the input sequence, enhancing the perception of global information. Second, the parallel computing ability of the Transformer makes it more efficient in processing long sequences. The multi-head self-attention mechanism enables the model to adapt to information at different time scales and capture features in the sequence more comprehensively. This combination can better balance global and local information when predicting the air-cooled backpressure, improving the performance and generalization ability of the model. The LSTM model can effectively process time series data and capture long-term dependencies.

[0012] Backpressure curve influencing factor prediction module: This module is another core module of the present invention. It predicts the backpressure curve under different influencing factors by using an improved ResNet network to find the optimal operation plan.

[0013] By introducing a residual learning structure, the problem of gradient disappearance in deep networks is effectively alleviated through skip connections, enabling the network to easily reach a depth of dozens of layers or more, improving the expressive ability and generalization performance of the model. The shared features and parameter efficiency of this architecture make the network easier to optimize. Through multi-scale feature extraction, with advantages such as high parameter efficiency, obvious hierarchical structure, and strong adaptability, the neural network can capture information at different scales more comprehensively and efficiently, improving the model performance and computational efficiency.

[0014] Secondly, the present invention provides a backpressure curve prediction method combining Transformer+LSTM and ResNet networks to achieve accurate and efficient prediction of the backpressure curve. The method includes:

[0015] Data acquisition steps: The data acquisition unit is equipped with a temperature measurement customized cable, an infrared temperature measurement system, a tube bundle deformation detection system, data acquisition equipment, and a DCS control system. It collects data through the acquisition equipment, extracts and uploads the collected data, and the acquisition parameters include temperature, pressure, and deformation data.

[0016] Processing the collected data to obtain a sample data set: Cleaning and normalizing the collected data to obtain the required data including ambient temperature, unit load, atmospheric pressure data, and dividing the data set into a test set and a training set.

[0017] Backpressure curve time prediction steps: This module is one of the core modules of the present invention. It includes a trained neural network model. In the present invention, the Transformer combined with the LSTM model is adopted.

[0018] The advantages of combining the Transformer and LSTM models in the prediction of the air-cooled backpressure are mainly reflected in two aspects. First, the self-attention mechanism introduced by the Transformer enables the model to better focus on the key parts in the input sequence, enhancing the perception of global information. Second, the parallel computing ability of the Transformer makes it more efficient in processing long sequences. The multi-head self-attention mechanism enables the model to adapt to information at different time scales and capture features in the sequence more comprehensively. This combination can better balance global and local information when predicting the air-cooled backpressure, improving the performance and generalization ability of the model. The LSTM model can effectively process time series data and capture long-term dependencies.

[0019] Backpressure curve influencing factor prediction steps: This module is another core module of the present invention. It predicts the backpressure curve under different influencing factors by using an improved ResNet network to find the optimal operation plan.

[0020] By introducing a residual learning structure, the problem of gradient disappearance in deep networks is effectively alleviated through skip connections, enabling the network to easily reach a depth of dozens of layers or more, improving the expressive ability and generalization performance of the model. The shared features and parameter efficiency of this architecture make the network easier to optimize. Through multi-scale feature extraction, with advantages such as high parameter efficiency, obvious hierarchical structure, and strong adaptability, the neural network can capture information at different scales more comprehensively and efficiently, improving the model performance and computational efficiency.

[0021] The significant advantages of the improved ResNet model proposed by the present invention in predicting backpressure curves are mainly reflected in the following aspects: First, the introduction of a deeper network structure enables the model to better capture abstract features and complex patterns, improves the sensitivity to subtle changes in the backpressure curve, and achieves more accurate predictions; Second, the multi-attention mechanism enhances the attention to different positions, enables the model to more flexibly capture key information, and improves the robustness and generalization ability; Finally, components such as residual blocks, global average pooling, convolutional layers, and activation functions are integrated, and the spatial features and non-linear relationships of the input sequence are learned through an organic combination, improving the adaptability to different backpressure curve patterns and showing more superiority.

[0022] Thirdly, the present application also provides a computing device and a computer-readable storage medium corresponding to a backpressure curve prediction method combining Transformer + LSTM and ResNet networks, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the above-mentioned backpressure curve prediction method. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Appendix Figure 1 is the flowchart of the control algorithm of the present invention;

[0024] Appendix Figure 2 is the data acquisition and processing flow of the present invention;

[0025] Appendix Figure 3 is the structural diagram of the Transformer + LSTM network of the present invention;

[0026] Appendix Figure 4 is the structural diagram of the improved ResNet network of the present invention;

[0027] Appendix Figure 5 is the prediction result diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] The present invention will be further described in detail below with reference to the accompanying drawings:

[0029] The following will be further described in detail with reference to the appendix Figures 1-5 for the present application.

[0030] The present application provides a backpressure curve prediction device and method based on deep learning. Referring to Figure 1 , a backpressure curve prediction method process based on deep learning includes: data acquisition, data preprocessing, providing the prediction result to the staff after being processed by the Transformer + LSTM and improved ResNet networks, and the staff judges whether it is necessary to adjust the unit operation according to the prediction result, and finally completes the backpressure curve prediction and monitoring.

[0031] An air-cooled backpressure curve prediction system based on data deep learning, comprising the following modules:

[0032] Sensor module: Installed at key environmental and equipment locations of the air-cooled system, used to collect various parameter data in real time, such as temperature, pressure, deformation, etc. The sensor module can include multiple sensor units, such as a temperature sensor unit arranged in a customized cable, an external deformation sensor unit, a wind speed sensor unit, and so on. Each unit is responsible for monitoring specific parameters of this part. The measurement results are convenient for transmission, recording, and processing.

[0033] Data acquisition module: Receives the raw data transmitted by modules such as the temperature sensor module and the deformation sensor module, performs analog-to-digital conversion processing on the data, and transmits the processed data to the data processing module.

[0034] Data processing module: Uses data deep learning algorithms to analyze and model the sensor data. Considering the portability of lightweight algorithms, the present invention uses the embedded module Jetson AGX Xavier as a test platform to transplant and test the algorithms of the present invention. The following is the relevant introduction of the test platform:

[0035] This module includes an 8-core NVIDIA Carmel ARMv8.2 64-bit CPU, supporting the parallel computing language CUDA 11. 64 Tensor cores, 16GB 256-bit LPDDR4x, dual deep learning accelerator (DLA) engines, NVIDIA vision accelerator engines, and the Xavier-integrated Volta GPU, which can obtain GPU workstation performance on an embedded module below 30W.

[0036] Figure 2 The specific processes of data acquisition, processing, and dataset division are given.

[0037] Data acquisition: Data acquisition is carried out in a batch import manner to obtain data containing key features and target variables. After data extraction is completed, standardize the data format and store the data in the same CSV format.

[0038] Data processing: Use the interpolation method to process missing values. For outliers in the data, use deep learning algorithms to detect and process the outliers. Perform operations such as data normalization and standardization, and process the stationarity of time series to ensure applicability to the proposed deep learning algorithms.

[0039] Dataset division: When splitting the dataset into the training set and the test set, first randomize the data to prevent the model from overfitting to a certain order. Divide the dataset into the training set, validation set, and test set according to the ratio of 60-20-20, and maintain a similar class distribution to avoid the impact of class imbalance on the model performance. Use 5-fold cross-validation to comprehensively evaluate the model performance.

[0040] Figure 3 The network structure diagram of Transformer+LSTM is given, and the specific prediction steps of the Transformer+LSTM module are as follows:

[0041] (1) Input data and mask padding: Map the input data to the vector space through the embedding layer. When processing variable-length sequences, in order to ensure that the model can correctly process the variable-length input, the padding positions in the input sequence need to be marked and masked during the calculation of attention. Create a matrix with the same shape as the input sequence, where the elements corresponding to the padding positions are set to 0, and the other positions are 1. This matrix is called the mask matrix. When calculating the attention weights, this mask matrix is used to set the weights of the padding positions to negative infinity so as to suppress them in subsequent operations.

[0042] (2) Maximum sequence length and positional encoding: Since Transformer does not handle the order information of the elements in the input sequence, positional encoding needs to be introduced. Here, sine and cosine functions are used to generate positional embeddings. For each position and each dimension of the maximum sequence, by calculating the sine and cosine values, these encoding values are added to the embedding vector. In this way, the model can distinguish the elements at different positions, thus introducing the order information of the sequence.

[0043] (3) Time series dimension and other feature embedding layers: By fusing time series data with other key features in the model input, a more comprehensive model representation is constructed. First, through time series embedding, map the dynamic time characteristics to the low-dimensional vector space to more effectively capture trends and periodicity. At the same time, for other types of features, such as meteorological data or operating status, use the embedding layer to project them into the same low-dimensional vector space to achieve a unified representation of multi-type features. Finally, combine the time series embedding and other feature embeddings to form a comprehensive input representation. This comprehensive representation not only enables the model to understand the input data more comprehensively, but also enables the model to more flexibly handle the complexity and diversity of the task, improving the accuracy of the model in predicting future time series trends.

[0044] (4) Multi-attention mechanism: The multi-attention mechanism improves the model's ability to capture information in the sequence by using multiple parallel attention heads, with each head independently processing the input sequence. In each head, the input sequence is mapped to different subspaces, enabling the model to learn to focus on different aspects of the input sequence. This mechanism allows the model to more comprehensively consider the correlations between different positions when processing sequence data, which helps to handle long-range dependencies and complex sequence structures.

[0045] (5) Residual connection and layer normalization: After the multi-head attention mechanism, residual connection and layer normalization are adopted to construct the output of each sub-layer. The residual connection adds the input to the output, which helps to alleviate the vanishing gradient problem. Layer normalization is used to normalize the output of each sub-layer, making it easier to train.

[0046] (6) Feed-forward neural network: After each attention sub-layer, a fully connected layer is used for non-linear transformation. The parameters of this fully connected layer need to be learned, and its input is the output of the attention sub-layer. Then, the non-linear activation function ReLU is applied to introduce non-linear transformation, and finally the output of the feed-forward neural network is obtained. This output is the model's higher-level abstraction and feature learning of the input sequence.

[0047] (7) Residual connection and layer normalization: Residual connection and layer normalization are introduced after the output of the feed-forward neural network to promote gradient stability and accelerate the training process. The residual connection first adds the output of the feed-forward neural network to the input to form the intermediate output of the residual connection. Then, layer normalization is performed on the intermediate output to make its distribution similar in each feature dimension. Such a design can alleviate the vanishing gradient problem, enabling the gradient to propagate more effectively, thereby improving the training efficiency of the model.

[0048] (8) LSTM layer: The output of the Transformer is connected to the LSTM layer. The LSTM is introduced to utilize its ability to model sequence data, especially long-term dependencies. The LSTM will learn the dynamic patterns and changes in the time series and output the final learning result. The input layer is set with 6 nodes corresponding to six input variables. The LSTM layer is a single layer with 100 nodes added, using the activation function inside the LSTM. The output layer is 1 node for the regression task without using an activation function.

[0049] Set the learning rate to 0.001, use the mean squared error (MSE) as the loss function, the Adam optimizer as the optimizer, and the number of iterations as 100 times.

[0050] The core idea of this module is to learn the global information of the sequence through the Transformer model, then input the learning result into the feature space processed by the LSTM, and finally use the LSTM layer to capture the dynamic features of the time series to achieve the task of predicting the backpressure curve over time. This hybrid model makes full use of the learning ability of the Transformer for global information and the modeling ability of the LSTM for time series, and has a good prediction effect.

[0051] Figure 4 The improved ResNet network structure diagram is given, and the specific prediction steps of the improved ResNet module are as follows:

[0052] (1) Embedding layer: Embed the preprocessed input data through the embedding layer to integrate the time series and other features to form a representation that can be understood by the model. Set the input nodes: 6 (ambient temperature, unit load, atmospheric pressure, fan power consumption, ambient wind speed, operating data). The output node dimension is set to 64, and the activation function ReLU is selected.

[0053] (2) Multi-attention mechanism: Add a multi-attention mechanism to weight different input features through attention weights, enabling the model to more flexibly focus on information of different importance. The number of heads of the multi-head attention is 8, and the number of hidden nodes in the attention head is set to 64.

[0054] (3) Max pooling layer: Use the max pooling layer to extract the key features of the input data, which can help the model better capture the overall trend and reduce the data dimension. The pooling window size is set to 3x3.

[0055] (4) Convolutional layer: Apply a series of convolutional layers and activation functions to learn the spatial features of the input data and introduce non-linear transformations. The number and size of the convolutional kernels: The size of the convolutional kernel is selected as 3x3, and the stride is 1.

[0056] (5) Enhancement module: Add an enhancement module to capture features of different scales through multiple branches to more comprehensively represent the complexity of the input data. Create 3 branches, and the 3 branches use 1x1, 1x3, and 1x5 convolutional kernels respectively to capture features of different scales and concatenate the outputs of each branch in the channel dimension.

[0057] (6) Residual block: The size of the convolutional kernel of the first convolutional layer is 3x3. Normalize the output of the convolutional layer and apply the activation function ReLU. The second convolutional layer also uses a 3x3 convolutional kernel. Normalize the output of the convolutional layer and apply the activation function ReLU. Connect the output of the second convolutional layer with the input residually (i.e., add them). Ensure that the model can more easily learn the residuals during training, which helps the propagation of gradients.

[0058] (7) Average pooling layer: Global average pooling is used to reduce the dimension of the output of the last convolutional layer, converging the spatial dimension into a global feature. In the design of the average pooling layer, a 3x3 pooling window is selected, with a stride of 2 and no padding, to downsample the input feature map. This design can reduce the spatial dimension of the feature map, extract the main features, and reduce the computational burden.

[0059] (8) Fully connected layer: The features are passed to the fully connected layer, and the final feature representation is mapped to the output space through the fully connected layer, that is, the prediction result of the backpressure curve. The number of input nodes is 64, and the number of output nodes is set to 1.

[0060] Through the improved network, the prediction of the backpressure curve under different influencing conditions can be realized, and the prediction of the backpressure curve under the action of composite influencing factors can also be realized. This prediction will provide an effective basis for the decision-making of operators and sufficient guarantee for the efficient operation of the air-cooled system.

[0061] Figure 5 The prediction diagram of the backpressure curve changing with time and the prediction diagram of the backpressure value changing with the fan frequency are given. It can be seen that the proposed model has good prediction effects and can provide guidance for actual production.

[0062] Data acquisition: Through various sensors and data acquisition devices, key parameters such as ambient temperature, unit load, atmospheric pressure, fan power consumption, ambient wind speed, and operation data are obtained in real time. The data acquisition frequency should be high enough to capture the dynamic changes of the system.

[0063] Data preprocessing: The collected raw data is cleaned, and missing values and outliers are processed to ensure the stability and accuracy of model training. Feature engineering is performed to integrate data of different parameters and create new features to better reflect the system state. The data is normalized to ensure that different features are in the same numerical range.

[0064] Deep learning model construction: The Transformer + LSTM and improved ResNet are used to analyze and model the sensor data to predict the backpressure curve, and the deep learning model is trained to accurately predict the condition of the backpressure curve.

[0065] The prediction results are provided to the staff: The prediction results of the backpressure curve processed by the Transformer + LSTM and the improved ResNet, as well as the possible influencing factors, are provided. The results are presented in an easy-to-understand form, such as a curve diagram or a numerical output.

[0066] Staff judgment and adjustment: The staff judge whether the system needs to be adjusted according to the prediction results. The adjustment includes changing the unit operation parameters or other system parameters to optimize the system performance.

[0067] Complete backpressure curve prediction and monitoring: The goal of the entire process is to achieve accurate prediction of the backpressure curve and real-time monitoring of the system operation. By continuously optimizing the model and adjusting the system, ensure that the system can operate efficiently under different working conditions.

[0068] The implementation principle of a method for predicting the backpressure curve based on deep learning in an embodiment of this application is to predict the backpressure curve through deep learning and guide the operators to make real-time adjustments, so that the system operates in the best state, improving the stability and economy of the system operation.

[0069] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0070] In the description of the present invention, unless otherwise stated, the orientation or positional relationship indicated by the terms "upper", "lower", "left", "right", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention.

[0071] Finally, it should be noted that the above technical solution is only one implementation manner of the present invention. For those skilled in the art, based on the disclosed application methods and principles of the present invention, various types of improvements or deformations can be easily made, not limited to the methods described in the above specific implementation manners of the present invention. Therefore, the above-described manner is only preferred and does not have a restrictive meaning.

Claims

1. A backpressure curve prediction system based on deep learning, characterized in that it includes the following parts: a data acquisition unit, a data processing unit, a backpressure curve time prediction module, and a backpressure curve influencing factor prediction module; Data acquisition unit: The data acquisition unit is equipped with a temperature measurement customized cable, an infrared temperature measurement system, a tube bundle deformation detection system, data acquisition equipment, and a DCS control system, collects data through the acquisition equipment, and extracts and uploads the collected data; The acquisition parameters include temperature, pressure, and deformation data; Data processing unit: Clean and normalize the collected data to obtain the required data including ambient temperature, unit load, atmospheric pressure data, and divide the data set into a test set and a training set; Deep learning model construction: Use Transformer+LSTM and improved ResNet to analyze and model the sensor data to predict the backpressure curve, and train the deep learning model to accurately predict the condition of the backpressure curve; Staff judgment and adjustment. The staff judge whether the system needs to be adjusted according to the prediction results. The adjustment includes changing the unit operation parameters or other system parameters to optimize the system performance.

2. The backpressure curve prediction system according to claim 1, characterized in that: The specific prediction steps of the Transformer+LSTM module are as follows: (1) Input data and mask filling: Map the input data to the vector space through the embedding layer. When processing variable-length sequences, in order to ensure that the model can correctly process the variable-length input, the padding positions in the input sequence need to be marked and masked during the calculation of attention; Create a matrix with the same shape as the input sequence, where the elements corresponding to the padding positions are set to 0, and the other positions are 1. This matrix is called the mask matrix; when calculating the attention weights, this mask matrix is used to set the weights of the padding positions to negative infinity so as to suppress them in subsequent operations; (2) Maximum sequence length and position encoding: Since Transformer does not process the order information of the elements in the input sequence, position encoding needs to be introduced; Here, sine and cosine functions are used to generate position embeddings. For each position and each dimension of the maximum sequence, by calculating the sine and cosine values, these encoding values are added to the embedding vector; in this way, the model can distinguish the elements at different positions, thus introducing the order information of the sequence; (3) Time series dimension and other feature embedding layers: By fusing time series data with other key features in the model input to build a more comprehensive model representation; First, through time series embedding, map the dynamic time characteristics to the low-dimensional vector space to more effectively capture trends and periodicity; At the same time, for other types of features, such as meteorological data or operating status, use the embedding layer to project them into the same low-dimensional vector space to achieve unified representation of multiple types of features; Finally, combine the time series embedding and other feature embeddings to form a comprehensive input representation; (4) Multi-attention mechanism: The multi-attention mechanism improves the model's ability to capture information in the sequence by using multiple parallel attention heads, each of which processes the input sequence independently; in each head, the input sequence is mapped to different subspaces, enabling the model to learn to focus on different aspects of the input sequence; This mechanism allows the model to more comprehensively consider the associations between different positions when processing sequence data, which helps to handle long-range dependencies and complex sequence structures; (5) Residual connection and layer normalization: After the multi-head attention mechanism, residual connection and layer normalization are adopted to construct the output of each sub-layer; the residual connection adds the input to the output, which helps to alleviate the vanishing gradient problem; layer normalization is used to normalize the output of each sub-layer, making it easier to train; (6) Feed-forward neural network: After each attention sub-layer, a fully-connected layer is used for non-linear transformation; the parameters of this fully-connected layer need to be learned, and its input is the output of the attention sub-layer; then, the non-linear activation function ReLU is applied to introduce non-linear transformation, and finally the output of the feed-forward neural network is obtained; This output is the model's higher-level abstraction and feature learning of the input sequence; (7) Residual connection and layer normalization: Residual connection and layer normalization are introduced after the output of the feed-forward neural network to promote gradient stability and accelerate the training process; The residual connection first adds the output of the feed-forward neural network to the input to form the intermediate output of the residual connection; then, the intermediate output is layer-normalized to have a similar distribution in each feature dimension; such a design can alleviate the vanishing gradient problem, enabling the gradient to propagate more effectively, thereby improving the training efficiency of the model; (8) LSTM layer: The output of the Transformer is connected to the LSTM layer; the LSTM is introduced to utilize its ability to model sequence data, especially long-term dependencies; The LSTM will learn the dynamic patterns and changes in the time series and output the final learning results; the input layer is set with 6 nodes, corresponding to six input variables; The LSTM layer is a single layer and adds 100 nodes, using the activation function inside the LSTM; the output layer is 1 node for regression tasks and does not use an activation function; The learning rate is set to 0.001, the loss function uses the mean squared error (MSE), the optimizer uses the Adam optimizer, and the number of iterations is 100 times.

3. The backpressure curve prediction system according to claim 1, characterized in that the specific prediction steps of the improved ResNet module are as follows: (1) Embedding layer: The preprocessed input data is subjected to feature embedding through the embedding layer to integrate the time series and other features to form a representation form understandable by the model; (2) Multi-attention mechanism: The multi-attention mechanism is added, and different input features are weighted by attention weights, enabling the model to more flexibly focus on information of different importance; the number of multi-head attention heads is 8, and the number of hidden nodes in the attention head is set to 64; (3) Max pooling layer: The max pooling layer is used to extract the key features of the input data, which can help the model better capture the overall trend and reduce the data dimension; the pooling window size is set to 3x3; (4) Convolutional layer: Apply a series of convolutional layers and activation functions to learn the spatial features of the input data and introduce non-linear transformations; number and size of convolutional kernels: The convolutional kernel size is selected as 3x3, and the stride is 1; (5) Enhancement module: Add an enhancement module to capture features at different scales through multiple branches to more comprehensively represent the complexity of the input data; create 3 branches, and the 3 branches use 1x1, 1x3, and 1x5 convolutional kernels respectively to capture features at different scales and concatenate the outputs of each branch in the channel dimension; (6) Residual block: The convolutional kernel size of the first convolutional layer is 3x3, normalize the output of the convolutional layer, and apply the activation function ReLU; the second convolutional layer also uses a 3x3 convolutional kernel; normalize the output of the convolutional layer and apply the activation function ReLU; connect the output of the second convolutional layer with the input in a residual connection; (7) Average pooling layer: Use global average pooling to reduce the dimension of the output of the last convolutional layer, and converge the spatial dimension into a global feature; in the design of the average pooling layer, select a 3x3 pooling window, the stride is 2, and there is no padding to downsample the input feature map; this design can reduce the spatial dimension of the feature map, extract the main features, and reduce the computational burden; (8) Fully connected layer: Transfer the features to the fully connected layer, and map the final feature representation to the output space through the fully connected layer, that is, the prediction result of the backpressure curve; the number of input nodes is 64, and the number of output nodes is set to 1; through the improved network, the backpressure curve prediction under different influence conditions can be realized, and the backpressure curve prediction under the action of composite influence factors can also be realized.

4. A method for predicting backpressure curve based on deep learning, characterized in that it includes the following steps: Data acquisition: The data acquisition unit is equipped with a temperature measurement customized cable, an infrared temperature measurement system, a tube bundle deformation detection system, data acquisition equipment, and a DCS control system, collect data through the acquisition equipment, and extract and upload the collected data; the acquisition parameters include temperature, pressure, and deformation data; Data processing: Clean and normalize the collected data to obtain the required data including ambient temperature, unit load, atmospheric pressure data, and divide the data set into a test set and a training set; Deep learning model construction: Use Transformer + LSTM and improved ResNet to analyze and model the sensor data to predict the backpressure curve, and train the deep learning model to accurately predict the condition of the backpressure curve; Staff judgment and adjustment, the staff judge whether the system needs to be adjusted according to the prediction result, and the adjustment includes changing the unit operation parameters or other system parameters to optimize the system performance.

5. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the program, the method described in claim 4 can be implemented.

6. A computer-readable storage medium, It is characterized in that the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the method described in claim 4.