Heating pipeline residual life prediction and maintenance method based on deep reinforcement learning
Through a deep reinforcement learning method, combined with multi-source data acquisition and CNN-LSTM network extraction of spatiotemporal features, high-precision life prediction and dynamic maintenance strategy generation of heating pipelines is achieved, solving the static modeling limitations of life prediction methods in the existing technology and the rigid maintenance strategy problems, significantly improving economic and reliability.
Patent Information
- Application Number
- CN202510335791.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The life prediction method of existing heating pipelines has problems such as static modeling limitations, rigid maintenance strategies and insufficient economics. It has failed to effectively consider the multi-factor coupling effect and dynamic operating conditions, and has failed to achieve multi-objective balance.
Using a method based on deep reinforcement learning, multi-source data acquisition and feature fusion, combined with CNN-LSTM network to extract spatiotemporal features, design state space and action space, define reward functions, and perform model training and update to realize dynamic adaptive maintenance strategy generation and multi-objective optimization.
The accuracy of the remaining life prediction of heating pipes has been improved, and the dynamic strategy has reduced the annual maintenance cost by 23%, the system failure rate has been reduced by 18%, and the MTBF has been increased to 2.3 years.
Smart Images

Figure FT_1 
Figure FT_2 
Figure QLYQS_1
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent maintenance of heating pipelines, and particularly relates to a method for predicting the remaining life and maintaining heating pipelines based on deep reinforcement learning, which is particularly applicable to the dynamic life assessment and maintenance decision optimization of aging pipelines in urban central heating pipe networks. Background Art
[0002] As an important part of urban infrastructure, heating pipelines are long-term in an environment of high temperature, high pressure and corrosive media, and are prone to failures such as material deterioration and weld cracking.
[0003] Traditional life prediction methods mainly rely on a single physical model (such as the Larson-Miller creep model) or statistical regression, and have the following defects: First, the limitation of static modeling. For example, Patent CN110309577A describes pipeline degradation using a staged stochastic diffusion process, but does not consider the coupling effect of multiple factors (the synergistic effect of temperature, pressure, and corrosion) and dynamic operating conditions changes; Second, the maintenance strategy is rigid. Existing methods (such as CN110705176A based on artificial neural network classification) mostly use fixed thresholds to trigger maintenance, and cannot dynamically adjust the strategy according to real-time risks, resulting in over-maintenance or under-maintenance; Third, the economy is insufficient. Traditional solutions do not incorporate maintenance costs, system reliability, and life extension benefits into a unified optimization framework, and it is difficult to achieve multi-objective balance.
[0004] In addition, the following key factors are often ignored in the maintenance decisions for heating pipelines in the prior art: The coexistence of multiple failure modes is ignored. Pipelines may simultaneously have corrosion thinning (soft failure) and sudden leakage (hard failure), which need to be differentially addressed; The spatio-temporal heterogeneity is ignored. Different sections of pipelines show non-uniform degradation characteristics due to material and load differences; The feedback closed-loop is missing. The maintenance effect is not fed back to the prediction model in real time, resulting in weak strategy adaptability.
[0005] Therefore, there is an urgent need for an integrated solution that combines multi-source perception, dynamic prediction, and intelligent decision-making to improve the safety and operation and maintenance economy of heating pipe networks. Summary of the Invention
[0006] In view of the above problems, the present invention proposes a method and system for predicting the remaining life and maintaining heating pipelines based on deep reinforcement learning, aiming to achieve: high-precision life prediction under multi-factor coupling; generation of dynamic adaptive maintenance strategies; multi-objective optimization of cost and reliability balance.
[0007] The technical solution of the present invention includes the following core modules: Multi-source data acquisition and feature fusion, including a sensing layer where acoustic emission sensors (detecting crack propagation), thermocouples (monitoring temperature gradients), pressure transmitters (recording pressure fluctuations), and electrochemical probes (measuring corrosion rates) are deployed at key pipeline nodes. The sampling frequency is 1 - 10 Hz, and LoRa wireless transmission is supported.
[0008] And a data preprocessing layer that fills missing data using time series interpolation method and calculates for missing data points using linear interpolation: ; Where and are valid timestamps adjacent to the missing point, and are the corresponding values.
[0009] Furthermore, the dimension difference is eliminated by Z-score standardization: ; Where is the mean value, is the standard deviation.
[0010] Furthermore, a sliding window (window length 60 s, step size 10 s) is used to extract time series statistical features (mean value, variance, kurtosis): ; ; .
[0011] Furthermore, a CNN-LSTM network is fused to extract spatio-temporal features. The CNN layer is designed with 3 convolutional kernels (size 3×1, number of channels 32 / 64 / 128) to capture the spatial patterns of local corrosion waveforms: ; Where is the feature map of the th layer, is the convolutional kernel weight (size 3×1, number of channels 32 / 64 / 128 in sequence), represents the convolutional operation, is the ReLU activation function.
[0012] Furthermore, the LSTM layer is designed as a bidirectional LSTM unit (hidden layer dimension 64) to model the long-term dependencies of temperature - pressure - corrosion rate. The gating mechanism of a single LSTM unit is as follows: ; ; ; ; 。
[0013] The bidirectional LSTM concatenates features through forward and backward propagation, and the final output is: ; where, and are the hidden states of the forward and backward LSTMs respectively, with dimensions of 64 each. After concatenation, the total dimension is 128.
[0014] On this basis, a deep reinforcement learning model is designed, including the following modules: Define the state space, and the state vector contains real-time sensing data (temperature T, pressure P, corrosion rate V); The predicted remaining useful life value (output by the LSTM sub-network); Historical maintenance records (the types, times, and costs of the last 3 maintenances).
[0015] Furthermore, define the action space: The action a ∈ {Preventive Maintenance (PM), Overhaul (OH), Minimal Repair (MR)}, and the execution cost ranges of each action are [2000, 5000], [10000, 20000], and [500, 1500] yuan respectively; PM can delay the degradation rate, OH resets the pipeline state to "almost new", and MR only repairs the current defect.
[0016] Furthermore, design the reward function: Reward Comprehensively consider economy, reliability, and life benefits: ; where, is the expected loss of no failure (positively correlated with the pipeline failure probability), is the increment of the remaining useful life after maintenance (calculated by the predicted difference of the LSTM), is the reliability improvement coefficient, quantified according to the change in the system MTBF (Mean Time Between Failures) after maintenance.
[0017] Furthermore, for model training and updating, it includes: In the offline stage, use the historical dataset (including 100,000 state-action-reward samples) to pre-train the DQN network model, with an initial learning rate of 0.001 and an experience replay buffer capacity of 10,000; In the online stage, the Q network is updated every 24 hours, and a double Q network architecture is adopted to reduce the overestimation bias.
[0018] Furthermore, the last module, multi-objective maintenance decision optimization, is introduced, including: First, policy generation is carried out. According to the current state select the action a with the maximum Q value; Secondly, conflict resolution is carried out. If multiple actions meet the reliability threshold (system MTBF ≥ 1 year), Monte Carlo tree search (MCTS) is used to simulate the state transition in the next 5 steps, and the path with the lowest long-term cost is selected; Finally, digital twin integration is carried out. A three-dimensional model of the pipeline is built on the Unity3D platform to map the stress distribution and maintenance suggestions in real time, supporting the interactive decision-making of operation and maintenance personnel.
[0019] Verified by the actual pipe network, the present invention has the following advantages compared with the traditional method: Improved in prediction accuracy. The multi-factor fusion reduces the remaining life prediction error from ±15% of the traditional method to ±7%; Optimized in economy: The dynamic policy reduces the average annual maintenance cost by 23% and avoids over-maintenance; Enhanced in reliability entropy: The system failure rate drops by 18%, and the MTBF is increased to 2.3 years. Description of the Drawings
[0020] Figure 1 is the flow chart of the method of the present invention.
[0021] Figure 2 is the architecture diagram of the deep reinforcement learning model. Detailed Embodiments
[0022] Step 1: Multi-source data collection and preprocessing.
[0023] Through sensor deployment, install at key nodes (elbows, welds, valves) of the heating pipeline: Acoustic emission sensors (AE sensors, accuracy ±0.1dB) to monitor crack propagation; Armored thermocouples (accuracy ±0.5°C) to record the temperature field distribution; Piezo-resistive pressure transmitters (accuracy ±0.2% FS) to collect pressure fluctuations; Three-electrode electrochemical probes (accuracy ±0.01mm / year) to measure the corrosion rate; The sensors upload data to the edge computing node in real time through the LoRa wireless module (transmission distance 3km).
[0024] Based on this, data cleaning and feature engineering are carried out. For missing value handling: Cubic spline interpolation is used to fill in the missing data. For example, when a certain section of temperature data is missing: ; Among them, It is determined by fitting through the adjacent 3 valid points.
[0025] Based on this, standardization processing is carried out, and Z-score standardization is performed on parameters such as pressure and temperature: .
[0026] Based on this, feature extraction is carried out: The mean, variance, and kurtosis are calculated using a sliding window (window length 60s, step size 10s); The CNN-LSTM network is used to extract spatio-temporal features.
[0027] Step 2: Construction of the deep reinforcement learning model.
[0028] Calculate the state space. The state vector is: ; Among them, is the current temperature (in degrees Celsius), is the current pressure (in MPa), is the corrosion rate ( ), is the predicted remaining life value (in years), is the last 3 maintenance records (type, time, cost).
[0029] Calculate the action space. The action set , and the parameters of each action: Action type PM, execution cost 2000 - 5000 yuan, life gain 0.5 - 1.2 years, reliability improvement 1.2 times; Action type OH, execution cost 10000 - 20000 yuan, life gain 3.0 - 5.0 years, reliability improvement 2.0 times; Action type MR, execution cost 500 - 1500 yuan, life gain 0.1 - 0.3 years, reliability improvement 1.1 times.
[0030] Calculate the reward function: ; Among them, is 10000 yuan, is 10 years, is 1 year.
[0031] Based on this, model training is carried out, including: Offline training, pre-trained using historical data of the past 10 years (including 100,000 sets of samples), learning rate 0.001, and the capacity of the experience replay buffer is 10,000; Online update, fine-tuned with new data every 24 hours, adopting a double Q-network architecture.
[0032] Step 3: Maintain decision generation and optimization.
[0033] Generate a policy, based on the current state , select the action with the largest Q value: .
[0034] Based on this, perform multi-objective optimization. If there are multiple actions that meet the reliability threshold (MTBF ≥ 1 year), use Monte Carlo tree search (MCTS) to simulate the decision-making path for the next 5 steps.
[0035] Based on this, perform digital twin verification: Build a 3D model of the pipeline on the Unity3D platform for real-time mapping: Stress distribution nephogram (based on ANSYS finite element analysis); Thermogram of maintenance suggestions (red areas need to be processed first); Life prediction timeline (showing the degradation trend in the next 5 years).
[0036] Verification of the embodiment: Select 10 typical pipelines in the heating pipeline network of a certain area for testing: The average prediction error of the traditional method is 12.3%, and the average annual maintenance cost is ¥186,000; The average prediction error of the method of the present invention is 6.8%, and the average annual maintenance cost is ¥143,000; The economic improvement is to save 23.1% of the cost, and the system failure rate is reduced by 18%.
[0037] It should be noted that: The specific implementation parameters can be adjusted according to the actual pipeline material (such as Q235B steel), design pressure (1.6 MPa), etc.; Reward function weight coefficient Needs to be optimized through orthogonal experiments; A safety threshold needs to be set during online learning to prevent the model from outputting dangerous decisions.
[0038] Finally, it should be noted that: The above embodiments are only used to illustrate the present invention, rather than limiting the protection scope of the present invention; Those skilled in the art can adjust the parameters and algorithms according to actual needs, which should be regarded as equivalent implementations of the present invention.
Claims
1. A method for predicting and maintaining the remaining life of heating pipes based on deep reinforcement learning, characterized in that: The following steps are involved: Collect multi-source sensor data of heating pipes, including temperature, pressure, corrosion rate, vibration signals and historical maintenance records, to build a multi-dimensional time series data set; Preprocessing the multidimensional time series data, including missing value filling, normalization and feature extraction, to generate a state feature vector; A deep reinforcement learning model is constructed, and the state space, action space and reward function are designed based on the deep Q network (DQN) framework, where: the state space is composed of the real-time operating parameters of the pipeline, the remaining life prediction value and the historical maintenance records; the action space includes three maintenance operations: preventive maintenance, overhaul and minimum repair; the reward function is dynamically calculated by combining maintenance cost, system reliability and remaining life extension benefits; Optimize model parameters by combining offline training with online updates to generate the optimal maintenance strategy; Maintenance decisions are made based on the current state feature vector output, and the model is continuously iterated and optimized based on execution feedback.
2. The method according to claim 1, characterized in that The feature extraction in step (2) includes: The sliding window method is used to extract the mean, variance and trend characteristics of time series data; The local spatial features are extracted by convolutional neural network (CNN), combined with long short-term memory network (LSTM) to capture long-term dependencies.
3. The method according to claim 1, characterized in that The reward function of the deep reinforcement learning model is designed as: ; in, is the expected loss if no failure occurs, To maintain operating costs, is the remaining life extension, is the system reliability improvement factor, is the weight parameter.
4. The method according to claim 1, characterized in that The offline training described in step (4) uses historical data sets to initialize the model, and the online update dynamically adjusts the Q value network parameters based on real-time collected data and maintenance feedback.
5. The method according to claim 1, characterized in that When the maintenance decision is generated, the solution that meets the reliability threshold and has the lowest cost is selected first. If there is a conflict, multi-objective optimization is performed through Monte Carlo Tree Search (MCTS).
6. A heating pipe remaining life prediction and maintenance system based on deep reinforcement learning, characterized in that: include: The first module includes a data acquisition module, which deploys multi-source sensors to collect pipeline temperature, pressure, corrosion rate and vibration signals in real time; The second module, including the preprocessing module, cleans, normalizes and fuses the original data to generate the state feature vector; The third module includes a model training module, which dynamically updates the Q-value network parameters by building and optimizing a deep reinforcement learning model; The fourth module, including the decision optimization module, generates maintenance strategies based on the model output and makes multi-objective decisions in combination with cost and reliability constraints; The fifth module includes an execution feedback module, which records the effects of maintenance operations and feeds them back to the model to achieve closed-loop optimization.
7. The system according to claim 6, characterized in that The data acquisition module includes an acoustic emission sensor, a thermocouple and a pressure transmitter, and supports wireless transmission and edge computing.
8. The system according to claim 6, characterized in that The decision optimization module is embedded in the digital twin platform to display the pipeline status and maintenance plan through three-dimensional visualization.
Citation Information
Patent Citations
Submarine pipeline remaining life prediction method based on IM and LMLE-BU algorithms
CN110309577A
Gas pipeline residual life prediction method and device
CN110705176A
Method for predicting service life of heat distribution pipeline
CN119004118A
Transformer life prediction method and system based on machine learning and Internet of Things
CN119441993A
Cited By
Machine vision and big data fused linear slide rail defect identification and prediction system
CN120411794A
A linear slide defect recognition and prediction system based on the fusion of machine vision and big data
CN120411794B