Buried pipeline corrosion rate risk assessment method and system based on streaming data
By integrating streaming data and multi-source sensor data, a corrosion acceleration model was constructed, which solved the problems of accuracy and timeliness in predicting corrosion of buried pipelines, and enabled dynamic monitoring of corrosion spread and refined management of high-risk areas.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGDONG SPECIAL EQUIP TESTING INST FOSHAN TESTING INST
- Filing Date
- 2025-08-28
- Publication Date
- 2026-06-09
AI Technical Summary
Existing technologies struggle to effectively utilize streaming data to efficiently integrate multi-source heterogeneous data and construct accurate corrosion diffusion models, resulting in insufficient accuracy and timeliness in predicting corrosion of buried pipelines and an inability to adapt to dynamic changes in complex environments.
By preprocessing multi-source sensor data, aligning and fusing heterogeneous data, using a streaming data processing framework for batch analysis, constructing a corrosion acceleration model, deeply mining key time nodes, and adjusting dynamic model parameters, real-time corrosion risk assessment and early warning can be achieved.
It improves the accuracy and timeliness of corrosion prediction for buried pipelines, enables refined monitoring of high-risk areas, and supports timely maintenance decisions.
Smart Images

Figure CN121117468B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of pipeline corrosion monitoring technology, and in particular to a method and system for risk assessment of corrosion rates of buried pipelines based on flow cytometry data. Background Technology
[0002] Buried pipelines, as critical infrastructure for energy and resource transportation, are essential for economic development and environmental protection due to their safety and long-term stability. However, pipelines operate in complex soil environments, making them susceptible to corrosion, which can lead to decreased structural integrity and even leaks, causing severe economic losses and environmental pollution. Current methods for assessing corrosion risks in buried pipelines largely rely on periodic inspections and single-parameter analysis, making them ill-suited to the dynamic changes in complex environments. These methods often suffer from insufficient data fusion and poor real-time performance when processing multi-source heterogeneous data, limiting the accuracy of corrosion predictions. Furthermore, existing technologies do not adequately address the dynamic evolution of corrosion diffusion processes and lack the ability to continuously track the spread of corrosion from localized areas to the overall network.
[0003] In the field of pipeline corrosion monitoring, the core challenge lies in how to effectively utilize streaming data to achieve dynamic monitoring and accurate prediction of pipeline corrosion diffusion. The multidimensional data streams of soil environmental parameters, electrochemical properties, and pipeline material conditions are characterized by high frequency, heterogeneity, and dynamic changes. Existing technologies struggle to efficiently integrate this data to construct accurate corrosion diffusion models. Insufficient real-time data processing leads to delays in predicting corrosion diffusion paths and speeds, thus affecting the timeliness of maintenance decisions. A deeper challenge is that the complex nonlinear process of corrosion diffusion requires mathematical models to characterize its evolution from point-like to surface-like patterns. However, current models often lack sufficient accuracy and adaptability in capturing this dynamic process, resulting in deviations between predicted results and actual corrosion behavior.
[0004] Therefore, how to efficiently integrate multi-source environmental and material information based on streaming data, construct a mathematical model that can accurately describe the dynamic process of corrosion diffusion, and achieve real-time monitoring and prediction has become a key issue in the risk assessment of corrosion rates of buried pipelines. Summary of the Invention
[0005] This invention provides a method and system for risk assessment of corrosion rates of buried pipelines based on streaming data, in order to improve the accuracy and timeliness of corrosion prediction for buried pipelines.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for risk assessment of corrosion rate of buried pipelines based on streaming data, comprising the following steps:
[0007] Multi-source sensor data of the soil environment around the buried pipeline were acquired and preprocessed to obtain an environmental feature dataset;
[0008] Obtain pipeline material status data, and perform heterogeneous data alignment and fusion operations based on the environmental feature dataset and pipeline material status data to obtain a comprehensive dataset;
[0009] A streaming data processing framework is used to parse the high-frequency data stream of the comprehensive dataset in batches to obtain the feature change sequence;
[0010] By constructing a corrosion acceleration model based on the feature change sequence, the key time points that may trigger corrosion acceleration can be identified.
[0011] In-depth analysis of soil environmental parameters and pipeline status data at key time points was conducted to determine the evolutionary pattern of corrosion diffusion.
[0012] Based on the evolution of corrosion diffusion, the parameter configuration of the dynamic model is adjusted to obtain a corrosion prediction model;
[0013] Flow cytometry data is input into the corrosion prediction model to determine the corrosion risk level of the pipeline.
[0014] If the corrosion risk level exceeds the preset threshold, an automatic early warning mechanism will be triggered to determine the pipeline segment information to be processed first.
[0015] Based on the pipeline segment information that is prioritized, the configuration of key areas for streaming data acquisition is updated and continuously fed into the corrosion prediction model for monitoring.
[0016] Furthermore, multi-source sensor data is acquired from the soil environment surrounding the buried pipeline and preprocessed to obtain an environmental feature dataset, including the following steps:
[0017] Acquire multi-source sensor data on the soil environment around buried pipelines, including soil moisture, pH, and electrochemical properties; perform noise reduction and standardization on the multi-source sensor data to obtain an environmental feature dataset.
[0018] Furthermore, pipeline material condition data is acquired, and heterogeneous data alignment and fusion operations are performed based on the environmental feature dataset and pipeline material condition data to obtain a comprehensive dataset, including the following steps:
[0019] The process involves acquiring environmental feature datasets and pipeline material status data, parsing data streams from different sources using a preset timestamp format to obtain standardized time series. If timestamp discrepancies exist in the standardized time series, linear interpolation is used to adjust the timestamps, resulting in an aligned time series. Based on the aligned time series, a timestamp-based hash mapping method is used to associate and match the environmental feature data and pipeline material status data, yielding a set of matched data points. Multidimensional features are extracted from the matched data point set, and principal component analysis is used to reduce the dimensionality of these features, resulting in a dimensionality-reduced feature set. Based on the dimensionality-reduced feature set, K-means clustering is used to group the data points, determining the feature clustering results. The multidimensional feature data is then fused using the feature clustering results to generate a fused dataset. If missing values exist in the fused dataset, imputation using the mean value is used to complete the data, resulting in a comprehensive dataset. If no missing values exist in the fused dataset, then the fused dataset is the comprehensive dataset.
[0020] Furthermore, a streaming data processing framework is used to parse the high-frequency data stream of the comprehensive dataset in batches to obtain the feature change sequence, including the following steps:
[0021] A streaming data processing framework is used to parse the high-frequency data stream of the comprehensive dataset in batches to obtain preliminary sequences of environmental features and pipeline status change trends. If the feature dimension of the preliminary sequence exceeds a preset threshold, principal component analysis is used to reduce the dimensionality and obtain a simplified change sequence. Based on the simplified change sequence, dynamic features of the soil environment are extracted to obtain feature vectors. Through correlation analysis between feature vectors and pipeline status, K-means clustering is used to cluster the status trends to obtain classification sequences of pipeline corrosion and pressure. If the real-time update frequency of the classification sequence is lower than the data stream input frequency, the batch size of the batch parsing is adjusted to obtain synchronously updated status trends. Based on the synchronously updated status trends, a dynamic correlation sequence between the soil environment and pipeline status is generated to obtain a real-time changing feature sequence. A sliding window mechanism is used to process the real-time changing feature sequence to obtain a smooth feature change sequence.
[0022] Furthermore, by analyzing the feature change sequence, a corrosion acceleration model is constructed to identify key time points that may trigger accelerated corrosion, including the following steps:
[0023] By collecting characteristic change sequence data, continuous corrosion-related characteristic change data are obtained from monitoring equipment to obtain the original time series. The original time series is smoothed using the moving average method to obtain the smoothed time series. Through difference operations, characteristic change trends are extracted from the smoothed time series to obtain the change trend sequence. If a point in the change trend sequence exceeds twice the standard deviation compared to a preset threshold, the point is marked as an abnormal fluctuation point, and an abnormal fluctuation point set is obtained. Through cluster analysis, time cluster nodes are identified from the abnormal fluctuation point set to determine the abnormal time node set. Regression analysis is used to model the correlation between the abnormal time node set and the corrosion acceleration factor to obtain the corrosion acceleration model. Through the corrosion acceleration model, possible abnormal fluctuation points in future time series are predicted to obtain the predicted fluctuation point set, and the key time nodes that may trigger corrosion acceleration are identified.
[0024] Furthermore, in-depth analysis of soil environmental parameters and pipeline status data corresponding to key time points is conducted to determine the evolutionary pattern of corrosion diffusion, including the following steps:
[0025] The process involves acquiring soil environmental parameters and pipeline status data corresponding to key time nodes, and using time series analysis to determine the time series characteristics of the node data. Based on these time series characteristics, principal component analysis (PCA) is used to extract environmental impact factors from the soil environmental parameters, resulting in a set of key environmental factors. If the factor values in the key environmental factor set exceed a preset threshold, local corrosion features are extracted from the pipeline status data, and support vector machine (SVM) is used to classify the severity of corrosion, resulting in a local corrosion feature distribution. Based on this local corrosion feature distribution and pipeline material characteristics, a random forest algorithm is used to predict corrosion diffusion paths, resulting in potential corrosion diffusion paths. For these potential corrosion diffusion paths, the corrosion rate distribution is analyzed, and a weighted average method is used to calculate the mean corrosion rate along the potential corrosion diffusion paths, determining the preliminary patterns of corrosion diffusion. If the preliminary patterns of corrosion diffusion indicate potential risk areas, cluster analysis is used to divide these risk areas, resulting in a high-risk area distribution. Based on the high-risk area distribution, a dynamic evolution model of corrosion diffusion is generated to determine the evolutionary patterns of corrosion diffusion.
[0026] Furthermore, based on the evolution of corrosion diffusion, the parameters of the dynamic model are adjusted to obtain a corrosion prediction model, including the following steps:
[0027] Historical corrosion data records are acquired, and preliminary analysis is conducted on corrosion rate and evolution patterns to obtain initial corrosion trend characteristics. Based on these initial characteristics, a pre-established dynamic model is used, combined with soil environment and environmental condition data, to calculate and adjust the parameters of the dynamic model, determining the adjusted parameter set. If the adjusted parameter set does not match a preset threshold range, the dynamic model is iteratively updated using an adaptive calibration method to obtain calibrated model parameters. Based on the calibrated model parameters, the dynamic model is run for environmental condition data under different soil environments to determine the distribution of predicted corrosion rates. By analyzing the distribution of predicted corrosion rates and combining the rate analysis results, a random forest algorithm is used to optimize the dynamic model, obtaining optimized prediction results. If the optimized prediction results deviate from the evolution patterns of historical data beyond a preset range, the final corrosion prediction model is output after optimizing the dynamic model.
[0028] Furthermore, the flow cytometry data is input into the corrosion prediction model to determine the corrosion risk level of the pipeline. This includes the following steps: real-time soil environmental data and pipeline material status data during pipeline operation are acquired through the flow cytometry data acquisition system, processed using the corrosion prediction model, and the spread rate and range of pipeline corrosion are calculated to determine the corrosion risk level of the pipeline.
[0029] Furthermore, based on the prioritized pipeline segment information, the configuration of key areas for streaming data acquisition is updated and continuously input into the corrosion prediction model for monitoring. This includes the following steps: obtaining high-risk area identifiers from the pipeline segment information; using a preset risk assessment algorithm to determine the priority of high-risk areas and obtain a list of high-risk areas; adjusting the configuration parameters of the data acquisition equipment based on the list of high-risk areas, increasing the acquisition frequency of high-risk areas, and acquiring environmental change data; extracting dynamic features from the environmental change data, continuously inputting them into the corrosion prediction model, and continuously updating the monitoring results.
[0030] A risk assessment system for the corrosion rate of buried pipelines based on streaming data includes a preprocessing module, an alignment and fusion module, an analysis module, a key time node judgment module, an evolution law determination module, a corrosion prediction model construction module, a corrosion risk level judgment module, an early warning module, and a cyclic monitoring module. The preprocessing module acquires multi-source sensor data from the soil environment surrounding the buried pipeline and preprocesses it to obtain an environmental feature dataset. The alignment and fusion module acquires pipeline material state data and performs heterogeneous data alignment and fusion operations based on the environmental feature dataset and pipeline material state data to obtain a comprehensive dataset. The analysis module uses a streaming data processing framework to analyze the high-frequency data stream of the comprehensive dataset in batches to obtain a feature change sequence. The key time node judgment module determines the feature change sequence... The system comprises the following modules: a corrosion acceleration model and a corrosion prediction model; an evolution law determination module, which performs in-depth analysis of soil environmental parameters and pipeline status data corresponding to key time points to determine the evolution law of corrosion diffusion; a corrosion prediction model construction module, which adjusts the parameter configuration of the dynamic model according to the evolution law of corrosion diffusion to obtain the corrosion prediction model; a corrosion risk level judgment module, which inputs the flow cytometry data into the corrosion prediction model to determine the corrosion risk level of the pipeline; an early warning module, which triggers an automatic early warning mechanism if the corrosion risk level exceeds a preset threshold and determines the pipeline segment information to be prioritized; and a cyclic monitoring module, which updates the key area configuration of the flow cytometry data acquisition based on the pipeline segment information to be prioritized and cyclically inputs it into the corrosion prediction model for continuous monitoring.
[0031] The technical effects and advantages provided by this invention in the above technical solution are as follows: This invention collects multidimensional data on the soil environment surrounding the pipeline, performs preprocessing and heterogeneous data alignment to construct a comprehensive dataset. It utilizes a streaming data processing framework to extract the dynamic changing trends of the environment and pipeline status, establishing a corrosion acceleration model. In-depth analysis is conducted at key time points to determine the evolution law of corrosion diffusion. Model parameters are optimized by combining historical data to achieve adaptive prediction of corrosion rates under different environments. This invention continuously inputs streaming data, calculates corrosion risk levels in real time, and triggers an early warning mechanism when thresholds are exceeded, generating targeted maintenance recommendations. By dynamically adjusting the data collection strategy, it achieves refined monitoring of high-risk areas, thereby improving the accuracy and timeliness of buried pipeline corrosion prediction and providing strong support for pipeline maintenance decisions. Attached Figure Description
[0032] Figure 1 This is a flowchart illustrating the risk assessment method for the corrosion rate of buried pipelines based on streaming data provided in an embodiment of the present invention.
[0033] Figure 2 This is a schematic diagram of the construction process of the comprehensive dataset provided in the embodiments of the present invention;
[0034] Figure 3 This is a schematic diagram of the refined monitoring process for high-risk areas provided in an embodiment of the present invention. Detailed Implementation
[0035] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0036] Example 1:
[0037] Reference Figure 1 This invention provides a method for risk assessment of corrosion rates in buried pipelines based on streaming data, comprising the following steps:
[0038] S101: Obtain multi-source sensor data from the soil environment around the buried pipeline and preprocess it to obtain an environmental feature dataset;
[0039] S102, acquire pipeline material status data, and perform heterogeneous data alignment and fusion operations based on the environmental feature dataset and pipeline material status data to obtain a comprehensive dataset;
[0040] S103 uses a streaming data processing framework to parse the high-frequency data stream of the comprehensive dataset in batches to obtain the feature change sequence;
[0041] S104. By constructing a corrosion acceleration model through the feature change sequence, the key time nodes that may trigger corrosion acceleration are identified.
[0042] S105, conduct in-depth analysis of soil environmental parameters and pipeline status data corresponding to key time nodes to determine the evolution law of corrosion diffusion;
[0043] S106. Based on the evolution law of corrosion diffusion, the parameter configuration of the dynamic model is adjusted to obtain the corrosion prediction model.
[0044] S107. Input the flow cytometry data into the corrosion prediction model to determine the corrosion risk level of the pipeline.
[0045] S108 If the corrosion risk level exceeds the preset threshold, an automatic early warning mechanism will be triggered to determine the pipeline section information to be processed first.
[0046] S109, based on the priority pipeline segment information, updates the key area configuration of the streaming data acquisition and cyclically inputs it into the corrosion prediction model for continuous monitoring.
[0047] In a preferred embodiment, step S101 involves acquiring multi-source sensor data from the soil environment surrounding the buried pipeline and preprocessing it to obtain an environmental feature dataset, including the following steps:
[0048] Multi-source sensor data on the soil environment surrounding buried pipelines, including soil moisture, pH, and electrochemical properties, are acquired. Multi-source sensor data fusion technology is used to obtain the original environmental dataset. Median filtering is applied to denoise the original environmental dataset, resulting in a denoised environmental dataset. Z-score normalization is then applied to normalize the denoised environmental dataset, resulting in a standardized environmental dataset. If missing values exist in the standardized environmental dataset, k-nearest neighbor interpolation is used to impute the missing data, resulting in a complete environmental dataset. If no missing values exist in the standardized environmental dataset, then the standardized environmental dataset is the complete environmental dataset. Based on the complete environmental dataset, statistical features, including mean, variance, and skewness, are extracted to obtain an environmental feature set. Principal component analysis is used to reduce the dimensionality of the environmental feature set, resulting in a low-dimensional environmental feature set. Based on the low-dimensional environmental feature set, the correlation coefficients between soil moisture, pH, and electrochemical properties are calculated to determine the changing trends of the soil environment surrounding the pipeline.
[0049] Specifically, when acquiring multi-source sensor data on the soil environment around buried pipelines, it is assumed that soil temperature sensors, soil moisture sensors, pH sensors, and electrochemical characteristic sensors were deployed to collect data. The soil temperature data ranged from -5℃ to 40℃, the soil moisture data ranged from 20% to 80%, the pH data ranged from pH 4.5 to 8.5, and the electrochemical conductivity data ranged from 0.1 to 2.0 mS / cm. The sampling frequency was once every half hour for 24 hours, resulting in 48 sets of data. The data then entered the data preprocessing stage. First, the raw data underwent noise reduction using a wavelet transform algorithm. The wavelet basis was set to Daubechies4, and the decomposition level was 3. A soft thresholding method was used to filter out high-frequency noise. For example, an outlier of 75% (significantly deviating from the mean of 45%) detected in the soil moisture data was smoothed to 46% to ensure data continuity. Next, standardization was performed using the Z-score standardization method, converting the data of each dimension into a distribution with a mean of 0 and a standard deviation of 1. If missing values exist, k-nearest neighbor interpolation is used to impute the missing data, resulting in a complete environmental dataset. If no missing values exist, the complete environmental dataset is obtained directly. Based on the complete environmental dataset, statistical features (mean, variance, skewness) are extracted to obtain an environmental feature set. Principal component analysis (PCA) is used to reduce the dimensionality of the environmental feature set (statistical features), resulting in a low-dimensional environmental feature set (such as PC1 and PC2, reflecting only the core information of environmental features). Based on the low-dimensional environmental feature set, the correlation coefficients between various soil parameters are calculated to determine the trend of soil environmental changes. The above process is implemented through automated scripts. Data processing and analysis are both written in Python, using the SciPy library for wavelet denoising and the NumPy library for standardization. Logically, a closed loop is formed from data acquisition to feature extraction, ensuring data quality and analytical reliability.
[0050] In a preferred embodiment, refer to Figure 2 Step S102 involves acquiring pipeline material condition data, performing heterogeneous data alignment and fusion operations based on the environmental feature dataset and the pipeline material condition data to obtain a comprehensive dataset, including the following steps:
[0051] The process involves acquiring environmental feature datasets and pipeline material status data, parsing data streams from different sources using a preset timestamp format to obtain standardized time series. If timestamp discrepancies exist in the standardized time series, linear interpolation is used to adjust the timestamps, resulting in an aligned time series. Based on the aligned time series, a timestamp-based hash mapping method is used to associate and match the environmental feature data and pipeline material status data, yielding a set of matched data points. Multidimensional features are extracted from the matched data point set, and principal component analysis is used to reduce the dimensionality of these features, resulting in a dimensionality-reduced feature set. Based on the dimensionality-reduced feature set, K-means clustering is used to group the data points, determining the feature clustering results. The multidimensional feature data is then fused using the feature clustering results to generate a fused dataset. If missing values exist in the fused dataset, imputation using the mean value is used to complete the data, resulting in a comprehensive dataset. If no missing values exist in the fused dataset, then the fused dataset is the comprehensive dataset.
[0052] Suppose we have two sets of data: an environmental characteristics dataset and a pipe material state dataset. The goal is to align and fuse these data to extract useful information. Environmental characteristics dataset (PC1_env (temperature-humidity principal component), PC2_env (pH-electrochemical principal component)):
[0053]
[0054] Pipeline material condition dataset (corrosion level and pressure):
[0055]
[0056] First, timestamps need to be standardized to the same format, and any timestamp discrepancies should be checked. It was noted that the timestamps for environmental characteristic data and pipeline material status data were not entirely consistent; in particular, environmental characteristic data was recorded at 10:30:00, while pipeline material status data was recorded at 11:00:00.
[0057] Linear interpolation can be used to extrapolate the 10:30:00 data for pipeline material condition. For example, assuming the corrosion rate is 15% at 10:00:00 and 20% at 11:00:00, the corrosion rate at 10:30:00 can be calculated using linear interpolation. If the pipeline data lacks a 10:30 record, interpolation can be performed based on a linear trend from 10:00 to 11:00.
[0058] Corrosion level: (15 + 20) ÷ 2 = 17.5%
[0059] Pressure: (800 + 750) ÷ 2 = 775 Pa
[0060] Environmental data is missing records for 11:00; interpolation is performed based on the linear trend from 10:00 to 10:30.
[0061] PC1_env interpolation calculation: 10:00 is 2.1, 10:30 is 2.5, increasing by 0.4 every 30 minutes, therefore the value at 11:00 = 2.5 + 0.4 = 2.9;
[0062] PC2_env interpolation calculation: 10:00 is 0.8 to 10:30 is 1.1, increasing by 0.3 every 30 minutes, therefore the value at 11:00 = 1.1 + 0.3 = 1.4.
[0063] Using the timestamp hash mapping method, each timestamp can be mapped to a unique hash value, and matching can be performed based on this. For example, 2025-03-01 10:00:00 is mapped to hash value H(10:00)=12345, 2025-03-01 10:30:00 is mapped to H(10:30)=12346, and 2025-03-01 11:00:00 is mapped to H(11:00)=12347.
[0064] By comparing hash values, the corresponding environmental feature data can be matched with the pipeline material status data. For 10:30:00, since the pipeline status data has already been obtained through linear interpolation, it can match the 10:30:00 timestamp in the environmental feature data. Aligned data point set:
[0065]
[0066] Fusion features (low-dimensional environmental features + pipeline state features) are extracted from the matched data point set. PCA is then used to reduce the dimensionality of these fusion features, resulting in a fusion-reduced feature set (e.g., PC1_fusion, PC2_fusion, reflecting the joint information of the environment and pipeline). Fusion-reduced feature set:
[0067]
[0068] Next, the K-means clustering algorithm is used to cluster the dimensionality-reduced features. Assuming K=2, the algorithm will divide the data into two groups. For example, 2025-03-01 10:00:00 and 2025-03-01 10:30:00 might be clustered in one group, while 2025-03-01 11:00:00 might be clustered in the other group.
[0069] Clustering results showed that the first two data points belonged to one group, and the last data point belonged to another group. Data fusion was then performed based on these clustering results. During this process, if any missing data appeared (e.g., missing pressure data), imputation using the mean could be used. For example, if pressure data for a certain time point was missing, the average pressure data for other time points could be calculated and used to impute the missing data.
[0070] Following the steps above, the final comprehensive dataset will look like this:
[0071]
[0072] In this way, the entire process from data alignment and dimensionality reduction to clustering and missing value completion is completed, and the final comprehensive dataset can be used for subsequent analysis.
[0073] In a preferred embodiment, step S103 involves using a streaming data processing framework to parse the high-frequency data stream of the comprehensive dataset in batches to obtain a feature change sequence, including the following steps:
[0074] A streaming data processing framework is used to parse the high-frequency data stream of the comprehensive dataset in batches to obtain preliminary sequences of environmental features and pipeline status change trends. If the feature dimension of the preliminary sequence exceeds a preset threshold, principal component analysis is used to reduce the dimensionality and obtain a simplified change sequence. Based on the simplified change sequence, dynamic features of the soil environment are extracted to obtain feature vectors. Through correlation analysis between feature vectors and pipeline status, K-means clustering is used to cluster the status trends to obtain classification sequences of pipeline corrosion and pressure. If the real-time update frequency of the classification sequence is lower than the data stream input frequency, the batch size of the batch parsing is adjusted to obtain synchronously updated status trends. Based on the synchronously updated status trends, a dynamic correlation sequence between the soil environment and pipeline status is generated to obtain a real-time changing feature sequence. A sliding window mechanism is used to process the real-time changing feature sequence to obtain a smooth feature change sequence.
[0075] In the processing of high-frequency data streams, using a streaming data processing framework for batch parsing can effectively analyze and extract useful feature change sequences. The following is a detailed explanation of this process, along with explanations of specific steps and algorithms.
[0076] First, a streaming data processing framework is needed to parse the high-frequency data stream of the comprehensive dataset in batches. High-frequency data streams mean that data is generated continuously at a very high frequency, thus requiring timely and batch processing. Streaming frameworks (such as Apache Kafka and Apache Flink) can divide the data stream into multiple batches for processing. Each batch contains data within a certain time period, effectively avoiding the waste of computational resources caused by data flooding. During parsing, data from each batch is extracted and timestamped to ensure all data is on the same timeline. Joint features (such as PC1_fusion (reflecting the correlation between environment and corrosion) and PC2_fusion (reflecting the correlation between environment and pressure)) and core pipeline status data are extracted from the comprehensive dataset to form a preliminary trend sequence. By processing and comparing this data in real time, the preliminary correlation and trends between environmental changes and pipeline status can be captured.
[0077] If the feature dimension of the initial sequence exceeds a preset threshold, dimensionality reduction is required. For example, to retain six features simultaneously: PC1_env, PC2_env, PC1_fusion, PC2_fusion, corrosion level, and pressure, with a threshold of 4. Principal Component Analysis (PCA) is a commonly used dimensionality reduction algorithm. It identifies the main directions of change in the data, retains the most important features, and reduces redundant information. Specifically, PCA calculates the covariance matrix of the data and extracts principal components through eigenvalue decomposition. By retaining the first few principal components, the dimensionality of the data can be significantly reduced, thus simplifying subsequent processing. For example, suppose the original dataset contains multiple features, but due to the high correlation between some features, the feature dimension may be too high. After dimensionality reduction using PCA, only the first two principal components may be retained, thus reducing the data dimension from multiple features to two principal components while still retaining most of the useful information.
[0078] After dimensionality reduction, a simplified sequence of changes was obtained, containing key information about environmental and pipeline condition changes. Next, we will focus on the dynamic characteristics of the soil environment. The soil environment has multifaceted effects on pipelines, particularly regarding the corrosive effects of factors such as moisture and temperature. By analyzing the changing trends of the soil environment (e.g., fluctuations in humidity and temperature), we can extract key dynamic features, such as the rate of change in soil moisture and temperature fluctuations. These dynamic features will serve as feature vectors to describe changes in the soil environment. By calculating the rate of change for each feature, a feature vector can be generated for each time point. For example, if soil temperature changes significantly while humidity changes relatively little, the feature vector might represent a combination of high temperature fluctuations and low humidity fluctuations.
[0079] Based on the dynamic feature vectors of the soil environment, correlation analysis of pipeline conditions is performed. Specifically, the K-means clustering algorithm is used to cluster pipeline conditions (such as corrosion level and pressure). K-means clustering is an unsupervised learning method used to divide data points into K clusters, minimizing the distance between data points within a cluster and maximizing the distance between data points in different clusters. By performing K-means clustering on pipeline condition data, different condition trends can be identified. For example, K-means clustering can divide pipeline condition data into several categories, such as slight corrosion and severe corrosion, with each category corresponding to a pipeline condition trend. After clustering, the resulting classification sequence can intuitively reflect the changing trends of pipeline corrosion level and pressure status over different time periods.
[0080] In real-time processing of high-frequency data streams, situations may arise where the update frequency of the classification sequence is lower than the input frequency of the data stream. In such cases, to ensure timely updates to the pipeline state trend, the batch size for parsing needs to be adjusted. If each parsing batch is too small, it may lead to a lag in pipeline state classification updates; if the batch is too large, it may cause processing delays. Therefore, adjusting the batch size is to optimize the balance between data processing speed and real-time updates. For example, suppose the data stream generates one data point per second, and the classification sequence updates once per minute. If the current batch size is once every ten seconds, it may not be able to capture the changing trend of each data point in a timely manner. Therefore, the batch size needs to be adjusted according to the actual situation so that each update can be as close as possible to the input frequency of the data stream.
[0081] After adjusting the batch size, the next step is to generate a dynamic correlation sequence between the soil environment and pipeline condition. This process tracks changes in both the soil and pipeline in real time by combining the dynamic characteristics of the soil environment with pipeline condition trends. The correlation sequence is continuously updated over time, providing a dynamic relationship between the soil and pipeline, which helps in further analyzing the long-term impact of environmental changes on pipeline condition. For example, over a certain period, as soil temperature and humidity increase, an increase in pipeline corrosion may be observed. Through correlation analysis, the interaction between the soil environment and pipeline condition can be clearly revealed.
[0082] To smooth the feature variation sequence and reduce the impact of noise, a sliding window mechanism is used for smoothing. The basic idea of the sliding window mechanism is to slide a fixed-size window across the time series, calculating the average or weighted average of the data within the window to obtain a smoothed feature sequence. This helps remove short-term fluctuations and noise, improving the stability of the analysis. For example, assuming a sliding window size of 30 minutes, the feature variation data within each 30-minute period will be smoothed to obtain a more stable trend. This smoothing process effectively removes the influence of sudden events or abnormal fluctuations on the feature sequence, making the final feature variation sequence more stable and reliable.
[0083] Through the above steps, batch parsing and processing of high-frequency data streams were achieved. Combining algorithms such as PCA dimensionality reduction, K-means clustering, and feature vector extraction, a dynamic correlation sequence between soil environment and pipeline status was obtained. By adjusting the batch size and applying a sliding window mechanism, the real-time performance and stability of data processing were further ensured. This process can provide strong data support for pipeline health monitoring, corrosion prediction, and maintenance decisions.
[0084] In a preferred embodiment, step S104, constructing a corrosion acceleration model through the feature change sequence and identifying key time points that may trigger accelerated corrosion, includes the following steps:
[0085] By collecting characteristic change sequence data, continuous corrosion-related characteristic change data are obtained to obtain the original time series. The original time series is smoothed using the moving average method to obtain a smoothed time series. Through difference operations, characteristic change trends are extracted from the smoothed time series to obtain a change trend sequence. If a point in the change trend sequence exceeds twice the standard deviation compared to a preset threshold, the point is marked as an abnormal fluctuation point, and an abnormal fluctuation point set is obtained. Through cluster analysis, time cluster nodes are identified from the abnormal fluctuation point set to determine the abnormal time node set. Regression analysis is used to model the correlation between the abnormal time node set and the corrosion acceleration factor to obtain a corrosion acceleration model. Through the corrosion acceleration model, possible abnormal fluctuation points in future time series are predicted to obtain a predicted fluctuation point set, and key time nodes that may trigger corrosion acceleration are identified.
[0086] Specifically, a corrosion acceleration model based on time series analysis is constructed using feature change sequences. The steps are as follows:
[0087] First, continuous corrosion-related characteristic change data are acquired through feature change sequence data collection. These characteristics need to be transformed into core features of the original corrosion-related time series, including corrosion thickness loss, soil pH difference (ΔpH), and chloride ion concentration difference (ΔCl). -The conversion process relies entirely on the basic data from S101-S102: the corrosion thickness loss (unit: mm / day) needs to be calculated by combining the "corrosion degree %" of S102 with the inherent parameters of the pipeline, and the formula is as follows.
[0088]
[0089] In the formula For the first Corrosion thickness loss over the day For S102, the first The degree of corrosion of the sky (substitute into decimal form, such as 15% is 0.15). This refers to the pipe wall thickness (e.g., 8mm is commonly used for DN600 gas pipelines). This is the time coefficient (k=1 here because "daily loss" is being calculated). ΔpH is the... The difference between the soil pH and the reference pH (usually neutral 6.5) is calculated using the following formula:
[0090]
[0091] in pH data from S101 pretreatment ;
[0092] ΔCl - Let be the difference between the chloride ion concentration on day t and the reference concentration (e.g., 200 mg / L calibrated on-site), expressed by the formula:
[0093]
[0094] Electrochemical characteristic sensor data from S101, Simultaneously, PC1_fusion (environment-corrosion associated principal component) in S103 is retained as an auxiliary feature, ultimately forming a feature including "timestamp-corrosion thickness loss-ΔpH-ΔCl". - The original time series of "-PC1_fusion" ensures that all data can be traced back to previous steps, with no independent new data sources.
[0095] Next, the original time series is smoothed using a moving average method to obtain a smoothed time series. This step aims to eliminate high-frequency noise in the original series (such as abrupt changes in corrosion thickness loss caused by occasional sensor fluctuations). A moving average algorithm with a window size of 3 (the optimal window verified by field data volatility) is used, and the formula is as follows:
[0096]
[0097] In the formula Let be the smoothed value for day t. , , These represent the corrosion thickness loss on days t-2, t-1, and t in the original time series. For example, if the original corrosion thickness loss sequence is [1.0, 1.2, 2.3, 3.8, 3.5, 3.7, 2.5, 4.2, 3.3, 3.0] (10 days of data).
[0098] The smoothed value on the 3rd day Smoothing value on day 4 By analogy, a complete smooth time series is obtained, which retains the overall trend of corrosion thickness loss while filtering out local noise interference.
[0099] Subsequently, "through differencing, the characteristic change trend is extracted from the smoothed time series to obtain the change trend sequence." The core of differencing is to quantify the "acceleration or deceleration trend" of the smoothed features, avoiding misjudgment of the trend caused by directly using the smoothed value. The calculation formula is as follows:
[0100]
[0101] In the formula The characteristic trend value on day t. Let be the smoothed value for day t. This is the smoothed value for day t-1.
[0102] Taking the smoothed corrosion thickness loss sequence [1.5, 2.4, 3.2, 3.7, 3.2, 3.5, 3.3, 3.5] (corresponding to days 3 to 10) as an example, the trend value on day 4 is... Day 5 Based on this, a trend sequence containing "timestamp-trend value" is calculated. This sequence directly reflects the daily fluctuation direction and magnitude of corrosion thickness loss, which is the core basis for subsequent anomaly judgment.
[0103] Abnormal fluctuation points are detected using the Z-score method. The formula for calculating the Z-score is as follows: ,in, This is the trend value on day t. It is the mean of the sequence. It is the standard deviation;
[0104] For example, the trend sequence is: [2.3,3.2,3.3,3.5,4.0,2.7,3.1,3.5,3.2];
[0105] mean ( ;
[0106] Standard deviation for ;
[0107] Anomaly detection:
[0108] Day 4 : Value > 2 (marked as abnormal);
[0109] Day 8 : Value > 2 (marked as abnormal);
[0110] Outliers: Day 4 and Day 8 are outliers.
[0111] Then, through cluster analysis, nodes with time clusters are identified from the set of abnormal fluctuation points, and the set of abnormal time nodes is determined. This step aims to screen outliers with "long duration and strong correlation" (excluding single-point accidental anomalies). The K-means clustering algorithm is used, and the clustering dimension is "time difference + trend value difference".
[0112] First, we set the number of clusters K=2 (based on the two common patterns in field data: "short-term fluctuations" and "continuous anomalies"), and define the distance metric formula as follows: ,
[0113] In the formula Let be the distance between the i-th and j-th outliers. The timestamps are two points (in days). The trend value of the two points; if Clusters with a time interval ≤ 2 days and a trend value difference ≤ 0.5 (based on a field-verified clustering threshold) are considered to belong to the same category. For example, the set of abnormal fluctuation points is...
[0114] The clustering result of the values {(Day 3, 1.8, 3.02), (Day 4, 1.3, 2.04), (Day 8, 1.1, 1.84), (Day 9, 0.9, 1.65)} yields two classes: the first class contains outliers from days 3 to 4 (time interval 1 day, trend difference 0.5). The second category includes outliers on days 8-9 (time interval 1 day, trend value difference 0.2). Then, categories with "2 or more data points" are selected (excluding single-point anomalies) to form an abnormal time node set, such as {[3,4], [8,9]}. Each set represents a continuous period of corrosion anomaly, providing a batch of effective abnormal samples for subsequent modeling.
[0115] Next, regression analysis is used to model the correlation between the abnormal time point set and the corrosion acceleration factor, resulting in a corrosion acceleration model. Here, the core variable needs to be clearly defined: the dependent variable is the corrosion acceleration factor (unit: mm / a), which is the difference between the corrosion rate during the abnormal period and the normal corrosion rate. The formula is as follows:
[0116]
[0117] in The average corrosion rate within the set of anomalous time points. The normal corrosion rate calibrated on-site (e.g., 0.1 mm / a); the independent variable is the average environmental characteristics within the set of abnormal time points, including... (Mean value of ΔpH within the set) (within the set) , (Mean of PC1_fusion within the set).
[0118] Construct a corrosion acceleration model, the formula is as follows:
[0119]
[0120] In the formula The predicted value of the corrosion acceleration factor at time t. For model parameters, The residual term (with a mean of 0 and a variance of ) (Normal distribution). Parameter calibration was performed using a Bayesian optimization algorithm, with historical data from S101-S103 over the past 90 days (including 120 sets of field monitoring data, corrosion rate 0.05~0.35 mm / a, pH value 4.5~7.8, Cl⁻ concentration 50~800 mg / L) as the training set. The optimization objective was to minimize the mean square error (MSE), and the MSE formula is: (m is the number of training samples, For predicted values, (These are measured values).
[0121] The parameter prior distribution is set as follows , After 200 iterations, the parameter values converged (e.g.) , , The prediction error of the model validation set (20% historical data) is ≤ ±5% and the coefficient of determination R² = 0.91, which meets the requirements for real-time early warning accuracy. Moreover, all modeling data comes from previous steps and there is no independently collected data.
[0122] Finally, using a corrosion acceleration model, we predict potential anomalous fluctuation points in future time series, obtaining a set of predicted fluctuation points to identify key time nodes that may trigger accelerated corrosion. This step requires obtaining future environmental characteristics as input, specifically future ΔpH and ΔCl. - ,PC1_fusion.
[0123] Future ΔpH and ΔCl - It can be inferred from PC2_env (acidity-electrochemical principal component) of S103.
[0124] The future PC2_env value is predicted using the ARIMA time series model, as shown in the formula.
[0125]
[0126] in To predict the step size, , These are the model coefficients. To express the future The predicted value of the "PC2_env" feature at time step 1. , They represent Time (the step before the predicted target) and The actual observed value of "PC2_env" at time (two steps before the predicted target). Then, the future ΔpH and ΔCl are inferred by using the calibration relationship. - The formula is
[0127]
[0128]
[0129] Where k1, b1, k2, and b2 are on-site calibration coefficients, such as k1=0.5 and b1=-0.3.
[0130] The future PC1_fusion prediction will use an LSTM neural network, trained on historical PC1_fusion sequences from S103, to output predicted values for the next τ days. .
[0131] The predicted ΔpH and ΔCl - , Substituting into the corrosion acceleration model, if ( If the standard deviation of the normal corrosion rate is 0.02 mm / a (i.e., the abnormal threshold is set to 0.14 mm / a), then t+τ is marked as the "predicted abnormal fluctuation point" and included in the predicted fluctuation point set.
[0132] To avoid misjudgment based on "single prediction anomalies," it is necessary to combine the CUSUM change point detection algorithm to verify trend persistence. The CUSUM cumulative bias formula is:
[0133]
[0134] In the formula The cumulative deviation on day t is... (Initial conditions) (Reference value, balancing sensitivity and false positive rate). The mean of the historical trend sequence; set the cumulative deviation threshold. (Optimal threshold verified on-site) If the cumulative deviation corresponding to a certain predicted abnormal fluctuation point If so, then this time point is confirmed as the "critical time node for accelerated corrosion".
[0135] In a preferred embodiment, step S105 involves in-depth analysis of soil environmental parameters and pipeline status data corresponding to key time points to determine the evolution pattern of corrosion diffusion, including the following steps:
[0136] The process involves acquiring soil environmental parameters and pipeline status data corresponding to key time nodes, and using time series analysis to determine the time series characteristics of the node data. Based on these time series characteristics, principal component analysis (PCA) is used to extract environmental impact factors from the soil environmental parameters, resulting in a set of key environmental factors. If the factor values in the key environmental factor set exceed a preset threshold, local corrosion features are extracted from the pipeline status data, and support vector machine (SVM) is used to classify the severity of corrosion, resulting in a local corrosion feature distribution. Based on this local corrosion feature distribution and pipeline material characteristics, a random forest algorithm is used to predict corrosion diffusion paths, resulting in potential corrosion diffusion paths. For these potential paths, the corrosion rate distribution is analyzed, and a weighted average method is used to calculate the mean corrosion rate along the potential path, determining the preliminary pattern of corrosion diffusion. If the preliminary pattern of corrosion diffusion indicates potential risk areas, cluster analysis is used to divide these risk areas, resulting in a high-risk area distribution. Based on the high-risk area distribution, a dynamic evolution model of corrosion diffusion is generated to determine the evolution pattern of corrosion diffusion. Visualization techniques are used to present this evolution pattern and determine the final corrosion evolution trend.
[0137] Specifically, when detecting key time points, time series analysis algorithms, such as sliding window detection, can be used. With a window size of 7 days, the rate of change of soil environmental parameters (such as pH, humidity, and chloride ion concentration) and pipeline status data (such as wall thickness and potential) can be calculated. Thresholds (such as pH change rate > 0.5 / day or wall thickness reduction rate > 0.1 mm / month) can be set to identify anomalous nodes. For example, if a node is detected where the pH drops sharply from 6.5 to 5.8 and the humidity rises from 40% to 60%, it is marked as a key node. Deep data mining is then performed on the node data, using principal component analysis (PCA) to reduce the 10-dimensional feature vectors of soil parameters and pipeline status to 3 dimensions, retaining 80% of the variance, and extracting the main influencing factors (such as pH and chloride ion concentration).
[0138] If the values of factors in the set of key environmental factors exceed preset thresholds, localized corrosion characteristics are extracted from the pipeline condition data. These thresholds are calibrated based on industry standards and historical field data. For example, the pH threshold is set to 5.5 (below this value indicates strong acidity), the chloride ion concentration threshold is set to 500 mg / L (above this value indicates high corrosion risk), and the soil moisture threshold is set to 55% (above this value accelerates electrochemical corrosion).
[0139] Localized corrosion feature extraction requires the pipeline surface inspection data obtained by S102 as a basis, including the three-dimensional coordinates of localized corrosion points, the depth and area of corrosion pits, and the wall thickness loss at the corresponding locations. For example, if 100 localized corrosion points are detected at locations such as (10cm, 20cm) and (15cm, 25cm) on the pipeline surface, these unlabeled corrosion points need to be pre-grouped in an unsupervised manner. Typically, the K-means clustering algorithm (K=3, corresponding to the three expected corrosion categories of "mild, moderate, and severe") is used. The clustering dimensions are "wall thickness loss, corrosion pit depth, and corrosion area". By calculating the Euclidean distance, the corrosion points are divided into 3 clusters. For example, cluster 1 has a wall thickness loss ≤0.1mm and a pit depth ≤0.05mm (preliminarily judged as potential mild corrosion), cluster 2 has a wall thickness loss of 0.1~0.3mm and a pit depth of 0.05~0.1mm (potential moderate corrosion), and cluster 3 has a wall thickness loss >0.3mm and a pit depth >0.1mm (potential severe corrosion). Then, the training dataset for SVM is constructed based on these cluster features.
[0140] The central features of each cluster are combined with key environmental factor values (such as pH=5.8, chloride ion=550mg / L) to form input features, and then combined with a small number of manually labeled corrosion severity labels (such as cluster 1 corresponding to "mild", cluster 2 corresponding to "moderate", and cluster 3 corresponding to "severe", the label data comes from historical pipeline dismantling and inspection records).
[0141] The SVM model is trained using the Radial Basis Function (RBF) kernel, a commonly used kernel function in Support Vector Machines (SVMs) and other machine learning algorithms. It is a non-linear kernel function that calculates similarity based on the distance between data points. The mathematical expression for the RBF kernel is:
[0142]
[0143] There are two data points. and The kernel function values between It is an adjustable parameter that controls the width of the RBF core.
[0144] The penalty parameter C and kernel parameter γ are optimized using 5-fold cross-validation (typically C=10 and γ=0.1 after optimization). After training, the features of all local corrosion points are input into the SVM, which outputs the corrosion severity classification result for each corrosion point. Finally, a heatmap is drawn according to "pipe surface coordinates + corrosion severity" to form the distribution of local corrosion features. For example, the first section of the pipe (0~50cm) is mainly lightly corroded, the middle section (50~150cm) is concentrated with moderate corrosion, and the last section (150~200cm) has two severely corroded points. This distribution result intuitively reflects the spatial differences of local corrosion in the pipe.
[0145] The "corrosion diffusion initiation point" is determined based on the distribution of local corrosion characteristics. Typically, the center coordinates of a severely corroded point or a moderately corroded concentrated area are selected (e.g., a severely corroded point at (100cm, 80cm) in the middle section of the pipe). Then, a meshed model of the pipe surface is constructed, discretizing the outer surface of the pipe into several mesh nodes. Each node needs to be associated with key environmental factor values and pipe material characteristic parameters.
[0146] Then, a graph theory model was used to transform the meshed nodes into vertices of a graph. The connectivity between nodes was defined according to the principle of physical adjacency (each node is connected to the four nodes above, below, left, and right). The weight of an edge was set to the absolute value of the difference in corrosion rate between two nodes; the smaller the weight, the less resistance there is to corrosion diffusion between the two nodes. Based on this graph theory model, Dijkstra's algorithm was used to calculate the "shortest weighted path" from the diffusion origin to all other nodes, resulting in 5-8 potential diffusion path candidates at the physical level. However, these candidate paths only consider physical resistance and do not take into account the dynamic effects of the environment and materials. Therefore, a random forest algorithm was introduced for probabilistic screening.
[0147] The input features of the random forest are the path-level features of each candidate path, including the average critical environmental factor value of all nodes on the path, the average corrosion resistance coefficient of the pipeline material, the corrosion severity at the starting point of the path, and the path length (e.g., 0.05m corresponding to 5 grid nodes). The training data uses pipeline corrosion diffusion cases from the past 3 years (paths that actually diffused are labeled as "positive samples" and those that did not diffuse are labeled as "negative samples"). By constructing 100 decision trees and voting to output the "corrosion diffusion probability" of each candidate path, the paths with a diffusion probability ≥70% are finally selected as potential corrosion diffusion paths.
[0148] Subsequently, the corrosion rate of each node along the potential diffusion path was obtained. For example, the corrosion rates of the five nodes on path 2 were 0.2 mm / a, 0.25 mm / a, 0.3 mm / a, 0.28 mm / a, and 0.22 mm / a, respectively. Since the weight of corrosion impact varies among different nodes along the path (e.g., high chloride ion nodes contribute more to the overall diffusion), a simple arithmetic average cannot be used; a weighted average method needs to be designed. The weighting coefficient is determined by the combination of "node key environmental factor value + wall thickness loss," as shown in the formula.
[0149]
[0150] in The chemical corrosivity principal component value of the i-th node (the higher the value, the greater the weight). This represents the relative proportion of wall thickness loss at that node (the maximum wall thickness loss is the maximum value among all nodes on the path). For example, if the weights of the nodes on path 2 are 0.8, 1.0, 1.2, 1.1, and 0.9, the weighted average corrosion rate is... .
[0151] By combining time series prediction models (such as ARIMA, parameters p=2, d=1, q=1) to analyze the time trend of corrosion rate, and inputting wall thickness loss data for the past 30 days on the potential path (e.g., from 8.0 mm to 7.85 mm), the corrosion rate change for the next 15 days is predicted. If the prediction results show that the rate increases from 0.26 mm / a to 0.31 mm / a (monthly increase of 0.05 mm), then the preliminary law of corrosion diffusion can be determined as "accelerated diffusion"—that is, the corrosion rate on the potential path gradually increases over time, and the diffusion range expands with the increase of rate. This preliminary law provides the core basis for subsequent risk area delineation.
[0152] If the initial pattern of corrosion diffusion indicates potential risk areas, cluster analysis is used to delineate risk areas by combining soil environmental parameters and pipeline condition data, resulting in a distribution of high-risk areas. When the initial pattern indicates that the corrosion rate on the potential path is accelerating, risk area delineation needs to be initiated. The input features for cluster analysis should cover "multi-dimensional risk indicators" to avoid the one-sidedness of a single dimension, typically including six dimensions: key environmental factor values of each node on the path, local corrosion severity (SVM classification results, mild=1, moderate=2, severe=3), diffusion probability of the potential diffusion path, weighted average corrosion rate, and wall thickness loss rate. The K-means clustering algorithm (K=3, corresponding to "high, medium, and low" risk categories) is used to divide all nodes on the pipeline surface into three clusters by calculating Euclidean distance: high-risk cluster, medium-risk cluster, and low-risk cluster. The risk area distribution map of the pipeline surface can be drawn based on the clustering results.
[0153] A dynamic evolution model for corrosion diffusion is constructed by integrating time-series prediction (the rate of corrosion increasing over time) and spatial expansion simulation (the spatial evolution of high-risk areas). This requires integrating temporal, spatial, and environmental data, with the core logic of spatial expansion driven by rate increases. Initial corrosion rate, monthly increase (e.g., 0.05 mm / a), time step (10 days), pipeline grid data, and soil chloride ion concentration distribution are input to establish a basic correlation: "every 0.05 mm / a increase in rate corresponds to an expansion of 2 cm / month," and a directional rule: "expansion accelerates along the direction of increasing chloride ion concentration and decelerates along the direction of decreasing chloride ion concentration." Through time iteration, the current rate is calculated according to the step size, and the basic expansion distance (e.g., 0.67 cm / 10 days) is calculated based on the rate difference. Then, the actual expansion distance (e.g., 0.8 cm along the direction of high concentration) is corrected using a directional coefficient. Based on this, the boundary grid of the high-risk area is updated, incorporating adjacent grids that meet the criteria. A random forest is introduced, and correction coefficients are output based on rate, environmental parameters, and historical errors to optimize the expansion results. After 3-6 months of iteration, a time-space evolution matrix and dynamic heat map are generated, showing the expansion process of high-risk areas and changes in key parameters.
[0154] In a preferred embodiment, step S106, adjusting the parameter configuration of the dynamic model according to the evolution law of corrosion diffusion to obtain a corrosion prediction model, includes the following steps:
[0155] Historical corrosion data records are acquired, and preliminary analysis of corrosion rate and evolution patterns is performed to obtain initial corrosion trend characteristics. Based on these initial corrosion trend characteristics, a pre-established dynamic model is used, combined with soil environment and environmental condition data, to calculate and adjust the parameters of the dynamic model, determining the adjusted parameter set. If the adjusted parameter set does not match the preset threshold range, the dynamic model is iteratively updated using an adaptive calibration method to obtain calibrated model parameters. Based on the calibrated model parameters, the dynamic model is run for environmental condition data under different soil environments to determine the distribution of predicted corrosion rates. By analyzing the distribution of predicted corrosion rates and combining the rate analysis results, a random forest algorithm is used to optimize the dynamic model, obtaining optimized prediction results. If the optimized prediction results deviate from the evolution patterns of historical data beyond a preset range, the final corrosion prediction model is output after optimizing the dynamic model. Using the final corrosion prediction model, the model is run for data input under different environmental conditions to obtain long-term corrosion rate trend predictions.
[0156] Specifically, when constructing a corrosion prediction model, the first step is to collect historical corrosion data records. This data typically includes corrosion thickness loss or corrosion rate at multiple time points, covering pipeline corrosion under different environmental conditions. For each time point, environmental characteristics (such as temperature, humidity, pH value, salinity, etc.) and the state of the pipeline material are recorded. The initial analysis aims to identify the evolution pattern of corrosion. Through time-series analysis of corrosion rate data, some key trend characteristics can be extracted, such as the acceleration or deceleration of the corrosion rate over time. This step typically involves calculating the rate of change of the corrosion rate, identifying the time points of corrosion aggravation, and visualizing the dynamics of corrosion development using graphical methods (such as time series graphs, trend lines, etc.).
[0157] Based on the corrosion trend characteristics derived from the preliminary analysis, the next step is to adjust the parameters according to the pre-established dynamic corrosion model. Dynamic corrosion models typically involve the influence of environmental conditions such as temperature, humidity, chemical composition, and soil type, which have a significant impact on the corrosion rate. The parameter adjustment process for the dynamic model includes the following steps:
[0158] Historical corrosion data and soil environmental characteristics (such as temperature, humidity, pH, etc.) are input into the dynamic model. Based on the model's operation, initial corrosion rate predictions are obtained. The deviation between the predicted and actual corrosion rates is compared to calculate the necessity of parameter adjustments and make corresponding adjustments. For example, corrosion constants or environmental influence factors in the model can be adjusted until the corrosion rate output by the model closely matches the actual observed data. If the initially adjusted parameters match the actual corrosion data well, subsequent steps can proceed; otherwise, the adaptive calibration phase begins.
[0159] When the adjusted model parameters do not meet the preset threshold range, an adaptive calibration method is used to iteratively update the model. The core idea of adaptive calibration is to iteratively optimize the model parameters through continuous feedback between the model output and actual corrosion data until the set performance standards are met. This calibration method typically uses optimization algorithms (such as gradient descent, genetic algorithms, etc.) to dynamically update the model parameters based on the prediction error. Each iteration adjusts the parameter set to make the model output as close as possible to the actual corrosion data. For example, if the model's predicted corrosion value is greater than the actual value, it may be necessary to adjust the weights of environmental factors in the model; if the predicted value is too small, it may be necessary to increase the influence values of certain environmental factors. Through multiple iterations, an optimal parameter set is finally found.
[0160] To prevent the model from iterating indefinitely, the convergence criteria for dynamic model parameter calibration are defined in this embodiment as follows:
[0161] Error convergence criterion: After three consecutive iterations, convergence is considered achieved if the mean squared error (MSE) of the validation set decreases by less than 1%. Maximum iteration limit: If the error criterion is not met after 50 iterations, the process is forcibly stopped, and the current optimal parameters are output. Parameter stability determination: When any parameter update amount... Furthermore, the overall error no longer decreased, and it was marked as stable. Real-time constraint: the time for a single parameter update is less than or equal to 15 seconds (edge computing node, Intel i7-1185G7 + 16GB RAM), ensuring the feasibility of field deployment. The above conditions were implemented using Python 3.9 + Scikit-learn 0.24, and met the industrial real-time requirements in actual operation.
[0162] The calibrated dynamic model parameters are used to further run the model and determine the predicted corrosion rates. In this stage, input data from different soil environmental conditions (such as humidity, temperature, and salinity) are used to predict corrosion rates under different conditions. The model run typically yields a distribution of predicted corrosion rates. This distribution helps to understand the trends in corrosion rates under different environments. For example, certain soil environments may lead to faster corrosion rates, while other environmental conditions may lead to slower corrosion rates. By analyzing these distributions, it is possible to identify which environmental factors have a significant impact on the corrosion rate over a given period.
[0163] After obtaining the predicted distribution of corrosion rates, the residuals between these predicted values and historical data for the corresponding time period are calculated (Residual = Historical Actual Value - Dynamic Model Predicted Value). Combined with the analysis of corrosion rate trends, the random forest algorithm is used to optimize the dynamic model. Random forest is a supervised learning method based on ensemble learning. It optimizes prediction results by constructing multiple decision trees and voting on their results. Specifically, random forest optimizes the dynamic model through the following steps:
[0164] Basic environmental features, dynamic model-related features (such as predicted values, residuals, and deviation rates between predicted values and historical means of the dynamic model), and interaction features (such as chloride ion concentration × humidity, and prediction residuals × environmental features) are used as inputs to the random forest.
[0165] Using the absolute residuals of the dynamic model as the target variable (i.e., the magnitude of the bias to be optimized), a random forest model (containing 100 decision trees, with features selected using the Gini coefficient) is trained. The model outputs two key results:
[0166] 1. Bias Prediction Value: For different combinations of environmental characteristics, output the amount of bias that the dynamic model may produce (e.g., "when chloride ion concentration > 500 mg / L and humidity > 55%, the bias is -0.035 mm / a", indicating that the dynamic model's prediction value is too low in this scenario).
[0167] 2. Feature importance ranking: Identify the core features that cause bias (e.g., the "chloride ion concentration × humidity" interaction item accounts for 32% of the importance and is the main source of bias).
[0168] Based on the bias patterns in the random forest output, the dynamic model is specifically adjusted by increasing the weight of high-importance features identified by the random forest in the dynamic model to enhance its sensitivity to high-risk scenarios. The bias value of the random forest predictions is then incorporated as a correction term into the prediction formula of the dynamic model.
[0169] Optimized dynamic model prediction = Original dynamic model prediction + Random forest biased prediction;
[0170] The optimized predicted values from the random forest output are used as pseudo-labels to supplement the training set of the dynamic model (especially for high-risk scenarios where historical data is scarce), and 30 new synthetic samples are added to improve the model's generalization ability.
[0171] After optimization using the random forest, the optimized corrosion rate prediction results are obtained. At this point, it's necessary to compare the deviation between the optimized results and the historical data. If the deviation exceeds a preset tolerance range, further model updates are required. Specifically, the historical corrosion data's evolution patterns are used as a benchmark to check the reasonableness of the model's predictions. For example, if the predicted corrosion rate fluctuates too much or too little, it may indicate poor model fit under certain environmental conditions, requiring further adjustment of model parameters. After multiple optimizations and updates, a corrosion prediction model that meets expectations is finally obtained. This model can predict corrosion rate trends over a certain timeframe based on data input under different environmental conditions. Using this model, engineers can develop reasonable anti-corrosion measures for different environmental conditions to extend pipeline service life.
[0172] It should be noted that the nonlinear regression model in S104 focuses on the local quantitative description of key nodes in corrosion acceleration; the model in S105 extends to the global time series, capturing the dynamic trend of the rate over time (such as seasonal fluctuations and long-term cumulative effects), compensating for the neglect of time-series correlation by nonlinear regression; the optimized model in S106 further integrates multi-dimensional data such as environmental characteristics and spatial distribution, and corrects biases through random forests to achieve a three-dimensional prediction of "time + space + environment", solving the limitations of the single dimension in the first two steps. The local prediction of S104 supports risk early warning of key nodes, the time-series trend of S105 is used for medium- and long-term maintenance planning, and the optimization results of S106 serve real-time monitoring and dynamic decision-making. The three are respectively adapted to different engineering scenarios of "emergency early warning - medium-term planning - real-time control".
[0173] In a preferred embodiment, step S107, inputting flow cytometry data into the corrosion prediction model to determine the corrosion risk level of the pipeline, includes the following steps:
[0174] A streaming data acquisition system is used to acquire real-time soil environmental data and pipeline material condition data during pipeline operation, determining a preliminary corrosion-related feature set. Based on this preliminary feature set, a corrosion prediction model is used to calculate the spread rate and range of pipeline corrosion, obtaining the dynamic trend of corrosion changes. For this dynamic trend, a random forest algorithm is used to further analyze the spread rate and range, determining the severity level of corrosion. If the severity level exceeds a preset threshold, a risk assessment module is triggered, combining the spread rate and range data to determine the potential risk level of the pipeline. Based on the potential risk level output by the risk assessment module, historical data and environmental variables related to the pipeline condition are obtained to determine if there are external factors accelerating corrosion. Based on the results of the external factor assessment, a preset rule base is used to comprehensively evaluate the pipeline condition, obtaining the final corrosion risk level and corresponding response strategy recommendations.
[0175] Specifically, the optimized corrosion prediction model processes the streaming data. First, real-time data streams are acquired from pipeline sensors, including parameters such as temperature (30°C), pressure (2.5 MPa), fluid pH (6.8), and flow velocity (1.2 m / s). This data is transmitted to edge computing nodes via the MQTT protocol at a frequency of 10 times per second. At the edge nodes, a sliding window algorithm (60-second window size) is used to preprocess the data, calculating the mean and standard deviation of parameters within the window. For example, the temperature mean is 30.2°C with a standard deviation of 0.3°C, and outliers (exceeding 3 times the standard deviation) are removed. Then, the preprocessed feature set is input into the optimized corrosion prediction model (a dynamic evolution model optimized with S106 random forest). The model predicts the current corrosion rate based on real-time environmental characteristics and historical corrosion trends. The model incorporates a "rate-time-space correlation module," which combines pipeline material parameters, the spatial gradient of the real-time rate, and the spatial decay coefficient fitted from historical diffusion data to calculate the corrosion depth, forming a dynamic corrosion trend of "stable increase in diffusion rate and expansion along the high chloride ion region."
[0176] To address this dynamic trend, a random forest algorithm was employed for further analysis: using diffusion speed, rate increase, diffusion range, and cumulative depth as input features, the output is a classification of the severity of corrosion. The classification rules were derived from historical fault data.
[0177] If the severity of corrosion exceeds a preset threshold, the risk assessment module is triggered. A fuzzy logic-based risk assessment algorithm is used, taking corrosion rate, corrosion depth, and pipeline operating years (5 years) as inputs. The fuzzy rules are defined as follows: high risk if corrosion rate > 0.2 mm / year or corrosion depth > 5% of wall thickness; medium risk if corrosion rate 0.1–0.2 mm / year and corrosion depth 2%–5% of wall thickness; and low risk for all others. Calculations show the current corrosion rate is 0.15 mm / year and corrosion depth is 4.5%, classifying it as medium risk. Results are pushed to the monitoring system in real-time via a Kafka streaming platform, and the data for the past 24 hours is stored in a Redis cache for trend analysis, ensuring real-time performance and traceability. All calculations are performed at edge nodes with a latency of less than 50 ms, meeting industrial real-time requirements.
[0178] In a preferred embodiment, step S108 involves acquiring corrosion risk data of the pipeline from the monitoring system and, in conjunction with a risk assessment method, quantifying the corrosion risk of each pipeline segment to obtain a preliminary risk level result. If the preliminary risk level result exceeds a preset threshold, an automatic early warning mechanism is triggered, transmitting relevant data to the analysis module to identify high-risk pipeline segments. Based on the high-risk pipeline segment information and soil environmental characteristic data, key environmental impact factors are extracted through a feature analysis module to obtain targeted environmental impact assessment data. If the environmental impact assessment data indicates that a specific pipeline segment has a high corrosion acceleration factor, corresponding maintenance recommendation data is generated to determine a priority list of pipeline segments. By comparing the priority list of pipeline segments with historical maintenance records and using a logical matching method, information on unprocessed or insufficiently processed pipeline segments is obtained, resulting in a final maintenance priority ranking. Based on the final maintenance priority ranking and the targeted maintenance recommendation data, a specific maintenance task allocation scheme is generated through a decision support system to determine the task execution order and resource allocation information. If there is a conflict between the generated task execution order and resource allocation information, the resource allocation is optimized and adjusted using a support vector machine algorithm to obtain an adjusted task execution scheme.
[0179] Specifically, corrosion risk data for each pipeline segment is collected through a pipeline monitoring system (such as online corrosion sensors and environmental monitoring equipment). Corrosion risk is defined as the corrosion rate. If the corrosion risk exceeds a preset rate threshold, the system determines that the pipeline segment is a high-risk segment and automatically issues an early warning. The warning includes the pipeline segment number, risk level value, and current corrosion-related monitoring data. The data is then pushed to the data analysis module for the next step of feature analysis.
[0180] The analysis module integrates soil environmental characteristic data (such as pH, moisture content, temperature, chloride ion concentration, and organic matter content) and performs joint analysis with data from high-risk pipe sections. Feature analysis functions (such as information gain ranking and random forest feature importance ranking) are used to extract environmental variables that significantly affect corrosion. Processing logic:
[0181] Input: Corrosion rate of high-risk pipe sections and its environmental variables; calculate the Pearson correlation coefficient of each environmental variable with the corrosion rate or the importance score obtained by training the model; output: a list of variables with significance greater than a set threshold as key influencing factors.
[0182] Example results: Chloride ion concentration score 0.82, pH value score 0.73, moisture content score 0.4; Extraction results: Chloride ion concentration and pH value are the key influencing factors.
[0183] If the indicators of key environmental factors are in the high-risk range for accelerated corrosion (e.g., chloride ion concentration > 300 mg / kg), the system generates maintenance recommendations for that section of pipeline. Recommended measures include: coating replacement, enhanced cathodic protection, and remediation of the surrounding soil. Time priority: emergency / urgent handling / routine review.
[0184] The priority list is generated by sorting according to risk level and the intensity of accelerator factors, for example:
[0185]
[0186] After the priority list is generated, the system connects to the historical maintenance record database and uses a logical matching algorithm to determine which pipe sections are: never maintained (no record), last maintained more than the maintenance cycle, or maintained but the measures were not aimed at the current risk factor (e.g., welded but no cathodic protection was performed). The matching logic is as follows:
[0187] If the current pipe segment number has a historical record and the maintenance time is greater than the current date – recommended cycle, and the historical maintenance measures do not include the current recommended measures, then it is marked as “insufficient treatment”. Example: A102 maintenance record only repaired welding, no soil remediation was performed ⇒ marked as “insufficient treatment”. Finally, the maintenance priority list and recommended measures are input into the decision support system (DSS), which generates task execution order and resource configuration information according to the following logic: schedule maintenance tasks according to priority, allocate maintenance resources (personnel, equipment, materials) based on current availability and the resources required by the task, and if the same resource is repeatedly allocated in multiple high-priority tasks, scheduling optimization is required.
[0188] In a preferred embodiment, refer to Figure 3 Step S109 involves updating the key area configuration for flow cytometry data acquisition based on the prioritized pipeline segment information, and cyclically inputting this data into the corrosion prediction model for continuous monitoring. This includes the following steps:
[0189] High-risk area identifiers are obtained from pipeline segment information. A pre-defined risk assessment algorithm is used to prioritize high-risk areas, resulting in a high-risk area list. Based on this list, the configuration parameters of the data acquisition equipment are adjusted to increase the acquisition frequency of high-risk areas, acquiring high-frequency streaming data. Environmental change characteristics are extracted from the high-frequency streaming data, and a time series analysis algorithm is used to determine environmental change trends. Based on these trends, the input parameters of the corrosion prediction model are updated, generating monitoring results. Abnormal data is extracted from the monitoring results; if abnormal data exceeds a pre-defined threshold, an alarm signal is generated. Based on the alarm signal, the acquisition frequency of high-risk areas is adjusted to acquire environmental change data. Dynamic features are extracted from the acquired environmental change data and continuously input into the corrosion prediction model to update the monitoring results.
[0190] Specifically, firstly, high-risk area identifiers are extracted from the pipeline segment information for priority processing. A pre-defined risk assessment algorithm (consistent with the S107 fuzzy logic rule) is used to determine the priority. Then, combining the corrosion rate, cumulative corrosion depth, and years of operation of the pipeline segment, a priority score is calculated. The scoring rules are as follows:
[0191] The input parameters (corrosion rate, corrosion depth, and years of operation) are mapped to a fuzzy score of 0-100, defined by a membership function:
[0192] Corrosion rate: ≤0.1mm / a corresponds to 0-30 points (low risk), 0.1~0.2mm / a corresponds to 30-70 points (medium risk), >0.2mm / a corresponds to 70-100 points (high risk);
[0193] Corrosion depth: ≤2% wall thickness corresponds to 0-30 points, 2%~5% corresponds to 30-70 points, >5% corresponds to 70-100 points;
[0194] Operating years: ≤5 years corresponds to 0-30 points, 5~10 years corresponds to 30-70 points, >10 years corresponds to 70-100 points.
[0195] Overall score =
[0196] Integrate with other high-risk pipeline sections to generate a list of high-risk areas.
[0197] The data acquisition configuration was adjusted based on the list of high-risk areas: For high-risk areas, the sensor acquisition parameters were updated from once per hour to once every 15 minutes. In addition to pipe surface temperature and pressure, the acquired data was supplemented with soil environmental parameters (chloride ion concentration, pH value, and humidity). The data was transmitted to the edge node (not in the cloud, maintaining the same edge computing architecture as S107) via the MQTT protocol at a frequency of 10 times per second. Each data transmission contained 20-dimensional features (environment + status) and was approximately 512KB in size.
[0198] Environmental change characteristics (such as the average chloride ion concentration of 530 mg / L, average pH value of 6.1, and average humidity of 60% in segment A102 within 1 hour) were extracted from high-frequency flow cytometry data. The trend was determined by time series analysis algorithm: by calculating the characteristic change rate through a sliding window (window size of 30 minutes), it was found that the chloride ion concentration increased by 5 mg / L per hour and the humidity increased by 2% per hour, which was determined to be a trend of "continuously increasing environmental corrosivity".
[0199] Based on this environmental change trend, the input parameters of the corrosion prediction model were updated: real-time features such as "chloride ion concentration 530 mg / L (original input 520 mg / L) and humidity 60% (original input 58%)" were substituted into the model. The model updated the corrosion rate prediction based on ARIMA time series logic (corrected from 0.18 mm / a to 0.19 mm / a), and at the same time updated the diffusion range through the built-in "rate-spatial correlation module" (expanded from the original 100-150m segment to the 95-155m segment), generating monitoring results that include "current rate, diffusion range, and environmental trend".
[0200] Extract abnormal data from the monitoring results: If the corrosion rate output by the model (0.19 mm / a) has not exceeded the preset threshold (0.2 mm / a), maintain the sampling frequency of every 15 minutes; after continuous monitoring for 2 hours, if the chloride ion concentration rises to 550 mg / L and the model predicted rate increases to 0.21 mm / a (exceeding the threshold of 0.2 mm / a), then generate an alarm signal (content: corrosion rate of pipe section A102 exceeds the threshold, current diffusion range 90-160 m, triggering factor "sudden increase in chloride ion concentration").
[0201] Based on the alarm signals, the data acquisition configuration was further adjusted: the acquisition frequency of section A102 was increased from once every 15 minutes to once every 5 minutes, focusing on acquiring dynamic changes in chloride ion concentration and pH value in this area; dynamic features (such as chloride ion concentration fluctuation of 8 mg / L within 5 minutes, pH value dropping to 6.0) were extracted from the newly acquired environmental change data, and these features were cyclically input into the corrosion prediction model. The model updated the monitoring results in real time (such as predicting that the rate will reach 0.22 mm / a after 1 hour and the diffusion range will expand to 85-165 m), and the updated results were pushed to the decision support system of S108 to provide a basis for whether to start maintenance tasks in advance, forming a continuous monitoring closed loop of "acquisition-analysis-prediction-adjustment".
[0202] All steps are automated using scripts in Python. Pandas is used to process the data, TensorFlow is used to build the model, and the scheduling system uses Airflow to check the task execution status daily to ensure the continuity of data flow and model prediction.
[0203] Example 2: This example provides a buried pipeline corrosion rate risk assessment system based on streaming data, including a preprocessing module, an alignment and fusion module, an analysis module, a key time node judgment module, an evolution law determination module, a corrosion prediction model construction module, a corrosion risk level judgment module, an early warning module, and a cyclic monitoring module;
[0204] Preprocessing module: Acquires multi-source sensor data from the soil environment surrounding the buried pipeline and performs preprocessing to obtain an environmental feature dataset;
[0205] Alignment and Fusion Module: Acquires pipeline material status data, performs heterogeneous data alignment and fusion operations based on environmental feature dataset and pipeline material status data, and obtains a comprehensive dataset;
[0206] Parsing module: Employs a streaming data processing framework to parse the high-frequency data stream of the comprehensive dataset in batches, obtaining the feature change sequence;
[0207] Key time node identification module: By constructing a corrosion acceleration model through feature change sequences, it identifies key time nodes that may trigger accelerated corrosion.
[0208] Evolution Pattern Determination Module: Deeply analyzes soil environmental parameters and pipeline status data corresponding to key time nodes to determine the evolution pattern of corrosion diffusion;
[0209] Corrosion prediction model construction module: Based on the evolution law of corrosion diffusion, the parameter configuration of the dynamic model is adjusted to obtain the corrosion prediction model;
[0210] Corrosion Risk Level Assessment Module: Input flow cytometry data into the corrosion prediction model to determine the corrosion risk level of the pipeline;
[0211] Early warning module: If the corrosion risk level exceeds the preset threshold, an automatic early warning mechanism is triggered to determine the pipeline segment information to be processed first.
[0212] Circulating monitoring module: Based on the pipeline segment information that is prioritized, it updates the configuration of key areas for flow data acquisition and cyclically inputs it into the corrosion prediction model for continuous monitoring.
[0213] Sensor Failure Tolerance Mechanism: To prevent single-point sensor failures from causing monitoring blind spots, the system implements automatic fault tolerance according to the following process:
[0214] Drift / Fault Detection: Perform isolated forest anomaly detection (n_estimators=100, contamination=0.05) on the real-time data of each sensor channel. If a channel is marked as anomaly for 3 consecutive sampling periods, it is determined to be a fault.
[0215] Backup channel switching: The faulty channel is automatically taken offline, and redundant sensors on the same node are activated (each node is equipped with 2 sets of humidity, 1 set of pH, and 1 set of conductivity backup). The switching delay is ≤5 s, and data gaps are compensated by linear interpolation.
[0216] Model degradation operation: When redundant sensors are insufficient, the system uses the data from the remaining channels to fill in missing values through random forest (based on the similarity of neighboring nodes) to maintain prediction accuracy, with an error increment of <3%.
[0217] Fault reporting: Fault information is pushed to the operation and maintenance platform via the MQTT protocol, triggering the automatic generation of a work order (including node number, fault type, and recommended replacement time).
[0218] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A method for risk assessment of corrosion rate of buried pipelines based on streaming data, characterized in that, Includes the following steps: Multi-source sensor data of the soil environment around the buried pipeline were acquired and preprocessed to obtain an environmental feature dataset; Obtain pipeline material status data, and perform heterogeneous data alignment and fusion operations based on the environmental feature dataset and pipeline material status data to obtain a comprehensive dataset; A streaming data processing framework is used to parse the high-frequency data stream of the comprehensive dataset in batches to obtain the feature change sequence; By constructing a corrosion acceleration model based on the feature change sequence, the key time points that may trigger corrosion acceleration can be identified. In-depth analysis of soil environmental parameters and pipeline status data at key time points was conducted to determine the evolutionary pattern of corrosion diffusion. Based on the evolution of corrosion diffusion, the parameter configuration of the dynamic model is adjusted to obtain a corrosion prediction model; Flow cytometry data is input into the corrosion prediction model to determine the corrosion risk level of the pipeline. If the corrosion risk level exceeds the preset threshold, an automatic early warning mechanism will be triggered to determine the pipeline segment information to be processed first. Based on the pipeline segment information that is prioritized, the configuration of key areas for flow cytometry data acquisition is updated and continuously input into the corrosion prediction model for monitoring. The process involves using a streaming data processing framework to parse the high-frequency data stream of the comprehensive dataset in batches to obtain the feature change sequence, including the following steps: A streaming data processing framework is used to parse the high-frequency data stream of the comprehensive dataset in batches to obtain a preliminary sequence of environmental characteristics and pipeline status change trends. If the feature dimension of the initial sequence exceeds a preset threshold, then principal component analysis algorithm is used to reduce the dimension and obtain a simplified variation sequence. Based on the simplified change sequence, the dynamic characteristics of the soil environment are extracted to obtain the feature vector; By analyzing the correlation between feature vectors and pipeline conditions, and using the K-means clustering algorithm to cluster the condition trends, a classification sequence of pipeline corrosion and pressure is obtained. If the real-time update frequency of the classification sequence is lower than the data stream input frequency, the batch size of the batch parsing is adjusted to obtain the synchronous update status trend. Based on the synchronously updated state trends, a dynamic correlation sequence between the soil environment and the pipeline state is generated, resulting in a real-time changing feature sequence. A sliding window mechanism is used to process the real-time changing feature sequence to obtain a smooth feature change sequence.
2. The method for risk assessment of corrosion rate of buried pipelines based on streaming data according to claim 1, characterized in that, Multi-source sensor data on the soil environment surrounding buried pipelines are acquired and preprocessed to obtain an environmental feature dataset, including the following steps: Acquire multi-source sensor data on the soil environment around buried pipelines, including soil moisture, pH, and electrochemical properties; The environmental feature dataset is obtained by denoising and standardizing the multi-source sensor data.
3. The method for risk assessment of corrosion rate of buried pipelines based on streaming data according to claim 1, characterized in that, To obtain pipeline material condition data, heterogeneous data alignment and fusion operations are performed on the environmental feature dataset and pipeline material condition data to obtain a comprehensive dataset, including the following steps: Acquire environmental feature datasets and pipeline material status data, and parse data streams from different sources using a preset timestamp format to obtain standardized time series; If there is a timestamp bias in the standardized time series, the timestamps are adjusted by linear interpolation to determine the aligned time series. Based on the aligned time series, a timestamp-based hash mapping method is used to associate and match environmental feature data with pipeline material status data to obtain a set of matching data points; Multidimensional features are extracted from the set of matched data points, and the features are reduced in dimensionality using principal component analysis to obtain a dimensionality-reduced feature set. Based on the dimensionality reduction feature set, the K-means clustering algorithm is used to group the data points and determine the feature clustering results; By fusing multidimensional feature data based on feature clustering results, a fused dataset is generated. If there are missing values in the fused dataset, the mean imputation method is used to complete the data to obtain the comprehensive dataset; if there are no missing values in the fused dataset, the fused dataset is the comprehensive dataset.
4. The method for risk assessment of corrosion rate of buried pipelines based on streaming data according to claim 1, characterized in that, By constructing a corrosion acceleration model based on the feature change sequence, the key time points that may trigger accelerated corrosion are identified, including the following steps: Continuous corrosion-related characteristic change data are obtained from monitoring equipment to obtain the original time series; The original time series was smoothed using the moving average method to obtain a smoothed time series; By using difference operations, feature change trends are extracted from smoothed time series to obtain change trend sequences; If a point in the trend sequence exceeds twice the standard deviation compared to a preset threshold, then that point is marked as an abnormal fluctuation point, and a set of abnormal fluctuation points is obtained. Cluster analysis was used to identify time-clustered nodes from the set of abnormal fluctuation points, thus determining the set of abnormal time nodes. Regression analysis was used to model the correlation between the set of abnormal time points and corrosion acceleration factors, resulting in a corrosion acceleration model. By using a corrosion acceleration model, we can predict the abnormal fluctuation points that may occur in future time series, obtain a set of predicted fluctuation points, and identify the key time nodes that may trigger corrosion acceleration.
5. The method for risk assessment of corrosion rate of buried pipelines based on streaming data according to claim 1, characterized in that, In-depth analysis of soil environmental parameters and pipeline status data at key time points was conducted to determine the evolutionary pattern of corrosion diffusion, including the following steps: Obtain soil environmental parameters and pipeline status data corresponding to key time nodes, and use time series analysis methods to determine the time series characteristics of the node data; Based on the time series characteristics, principal component analysis algorithm was used to extract environmental impact factors from soil environmental parameters, and a set of key environmental factors was obtained. If the factor values in the set of key environmental factors exceed the preset threshold, local corrosion features are extracted from the pipeline status data, and the severity of corrosion is classified using the support vector machine algorithm to obtain the distribution of local corrosion features. Based on the distribution of local corrosion characteristics and the properties of the pipeline material, the random forest algorithm is used to predict the corrosion diffusion path and obtain the potential corrosion diffusion path. For potential corrosion diffusion paths, the corrosion rate distribution is analyzed, and the mean corrosion rate on the potential corrosion diffusion path is calculated using a weighted average method to determine the preliminary law of corrosion diffusion. If the preliminary pattern of corrosion spread indicates potential risk areas, cluster analysis can be used to divide the risk areas and obtain the distribution of high-risk areas by combining soil environmental parameters and pipeline condition data. Based on the distribution of high-risk areas, a dynamic evolution model of corrosion diffusion is generated to determine the evolution law of corrosion diffusion.
6. The method for risk assessment of corrosion rate of buried pipelines based on streaming data according to claim 1, characterized in that, Based on the evolution of corrosion diffusion, the parameters of the dynamic model are adjusted to obtain a corrosion prediction model, including the following steps: Historical corrosion data records are obtained, and preliminary analysis is conducted on corrosion rate and evolution patterns to obtain initial corrosion trend characteristics. Based on the initial corrosion trend characteristics, a pre-established dynamic model was used, combined with data on soil environment and environmental conditions, to calculate the parameter adjustment of the dynamic model and determine the adjusted parameter set. If the adjusted parameter set does not match the preset threshold range, the dynamic model is iteratively updated using an adaptive calibration method to obtain the calibrated model parameters. Based on the calibrated model parameters, a dynamic model is run for environmental conditions under different soil environments to determine the distribution of predicted corrosion rates. By analyzing the distribution of predicted corrosion rates and combining the rate analysis results, the dynamic model is optimized using the random forest algorithm to obtain optimized prediction results. If the deviation between the optimized prediction results and the evolution patterns of historical data exceeds the preset range, the final corrosion prediction model will be output after optimizing the dynamic model.
7. The method for risk assessment of corrosion rate of buried pipelines based on streaming data according to claim 1, characterized in that, Inputting flow cytometry data into a corrosion prediction model to determine the corrosion risk level of the pipeline includes the following steps: By using a streaming data acquisition system, real-time soil environmental data and pipeline material condition data are obtained during pipeline operation. Corrosion prediction models are then used to process the data, calculate the spread rate and range of pipeline corrosion, and determine the corrosion risk level of the pipeline.
8. The method for risk assessment of corrosion rate of buried pipelines based on streaming data according to claim 1, characterized in that, Based on the prioritized pipeline segment information, the key area configuration for flow cytometry data acquisition is updated and continuously fed into the corrosion prediction model for monitoring, including the following steps: High-risk area identifiers are obtained from pipeline segment information. A preset risk assessment algorithm is used to determine the priority of high-risk areas and obtain a list of high-risk areas. Based on the list of high-risk areas, adjust the configuration parameters of the data acquisition equipment, increase the collection frequency in high-risk areas, and obtain environmental change data. Dynamic features are extracted from environmental change data and cyclically input into the corrosion prediction model to continuously update monitoring results.
9. A risk assessment system for the corrosion rate of buried pipelines based on streaming data, used to implement the assessment method described in any one of claims 1-8, characterized in that, It includes a preprocessing module, an alignment and fusion module, a parsing module, a key time node judgment module, an evolution law determination module, a corrosion prediction model construction module, a corrosion risk level judgment module, an early warning module, and a cyclic monitoring module; Preprocessing module: Acquires multi-source sensor data from the soil environment surrounding the buried pipeline and performs preprocessing to obtain an environmental feature dataset; Alignment and Fusion Module: Acquires pipeline material status data, performs heterogeneous data alignment and fusion operations based on environmental feature dataset and pipeline material status data, and obtains a comprehensive dataset; Parsing module: Employs a streaming data processing framework to parse the high-frequency data stream of the comprehensive dataset in batches, obtaining the feature change sequence; Key time node identification module: By constructing a corrosion acceleration model through feature change sequences, it identifies key time nodes that may trigger accelerated corrosion. Evolution Pattern Determination Module: Deeply analyzes soil environmental parameters and pipeline status data corresponding to key time nodes to determine the evolution pattern of corrosion diffusion; Corrosion prediction model construction module: Based on the evolution law of corrosion diffusion, the parameter configuration of the dynamic model is adjusted to obtain the corrosion prediction model; Corrosion Risk Level Assessment Module: Input flow cytometry data into the corrosion prediction model to determine the corrosion risk level of the pipeline; Early warning module: If the corrosion risk level exceeds the preset threshold, an automatic early warning mechanism is triggered to determine the pipeline segment information to be processed first. Circulating monitoring module: Based on the pipeline segment information that is prioritized, it updates the configuration of key areas for flow data acquisition and cyclically inputs it into the corrosion prediction model for continuous monitoring.