Pipeline steel corrosion mode identification method based on random forest algorithm

Through the pipeline steel corrosion pattern recognition method based on random forest algorithm, the problem of relying on manual reliance on electrochemical corrosion data analysis in the prior art is solved, and accurate automatic judgment of pipeline corrosion patterns and real-time online monitoring are realized.

CN120102667APending Publication Date: 2025-06-06SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510201672.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the prior art, the electrochemical impedance data collected by electrochemical corrosion sensors require manual analysis, which results in the accuracy of the pipeline corrosion state and corrosion pattern relying on professionals, with large errors and low efficiency, and real-time online monitoring cannot be achieved.

Method used

The pipeline steel corrosion pattern recognition method based on the random forest algorithm is adopted. By collecting the electrochemical impedance information of pipeline corrosion electrochemical sensors, a corrosion pattern database is established, and a classification model is trained using the random forest algorithm to realize automatic analysis of electrochemical impedance data and accurate judgment of corrosion patterns.

Benefits of technology

It improves the accuracy of pipeline corrosion mode judgment, reduces the error of manual judgment, realizes real-time online monitoring of pipeline corrosion status, and improves the efficiency of corrosion mode judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120102667A_ABST
    Figure CN120102667A_ABST
Patent Text Reader

Abstract

The invention relates to a pipeline steel corrosion mode identification method based on a random forest algorithm, and belongs to the field of intelligent detection. The method comprises the following steps: firstly, collecting various electrochemical impedance signals of the pipeline corrosion electrochemical sensor; then processing the collected electrochemical impedance information, and establishing a pipeline corrosion mode database according to an electrochemical impedance spectroscopy equivalent circuit theory; training by using the pipeline corrosion mode database and adopting a random forest algorithm to obtain a pipeline corrosion mode classification model; and finally, substituting pipeline operation monitoring data into the trained pipeline corrosion mode classification model to obtain a current pipeline corrosion mode and a corrosion state. The problem that the corrosion state of the long-distance buried oil and gas conveying pipeline cannot be monitored online in real time can be solved, and the method has important significance in improving the corrosion monitoring technology level of the long-distance buried oil and gas conveying pipeline, guaranteeing safe operation of the pipeline and saving the maintenance cost of the pipeline.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of pipeline corrosion detection, and in particular relates to a pipeline steel corrosion pattern recognition method based on a random forest algorithm. Background Art

[0002] Pipeline transportation is an emerging and economical mode of transportation. Compared with traditional modes of transportation such as roads, railways, and aviation, pipeline transportation of oil has the advantages of large transportation capacity, low cost, less terrain restrictions, stability and safety, and has become the main mode of oil and gas transportation in my country. However, since the medium in the pipeline is mostly a mixture of oil and gas containing acids, alkalis, salts and other corrosive substances, and the outside of the pipeline is corrosive media such as the atmosphere and soil, the oil and gas pipeline is prone to electrochemical corrosion reactions, and severe corrosion can easily cause safety hazards and lead to accidents.

[0003] At present, the traditional pipeline corrosion status detection methods mainly include corrosion hanging method, thickness measurement method, electrochemical method, etc. Among them, the corrosion hanging method is to place a piece of metal sheet with the same material as the pipeline into the system for monitoring, and observe the pipeline corrosion status by regularly measuring the degree of change of the metal sheet. However, this method requires regular excavation and detection of metal sheets for long-distance buried pipelines, and the detection cost is high; the thickness measurement method uses ultrasonic, magnetic induction and other methods to measure the thickness of the pipeline wall on the ground to monitor the pipeline corrosion. This method cannot be used to excavate underground pipelines, but this method cannot monitor the pipeline corrosion status in real time online. The electrochemical method uses an electrochemical corrosion sensor to detect the electrochemical impedance information of the pipeline in real time, and obtains the pipeline corrosion status by analyzing the electrochemical impedance data of the pipeline. This method has low cost, high detection accuracy and can be monitored online in real time, so it is widely used in pipeline corrosion status detection.

[0004] However, the main problem with electrochemistry at present is that the electrochemical impedance data collected by the electrochemical corrosion sensor needs to be manually analyzed. The accuracy of the pipeline corrosion status and corrosion mode obtained by electrochemical impedance analysis is heavily dependent on professionals. Not only is there a large error in the manual judgment of the pipeline corrosion status and the accuracy is low, but the judgment time of the pipeline corrosion mode is also long, resulting in low efficiency in pipeline corrosion judgment and inability to achieve real-time online monitoring of the pipeline corrosion status. Summary of the invention

[0005] In view of the above-mentioned deficiencies in the prior art, the purpose of the present invention is to provide a pipeline steel corrosion pattern recognition method based on a random forest algorithm, which is mainly used to accurately and automatically identify the pipeline corrosion pattern according to the pipeline electrochemical impedance information collected by the electrochemical corrosion sensor. The method first collects various electrochemical impedance signals of the pipeline corrosion electrochemical sensor; then processes the collected electrochemical impedance information, and establishes a pipeline corrosion pattern database based on the electrochemical impedance spectrum equivalent circuit theory; uses the pipeline corrosion pattern database and uses the random forest algorithm to train a pipeline corrosion pattern classification model; finally, uses the trained pipeline corrosion pattern classification model to analyze and process the electrochemical impedance data, thereby obtaining the pipeline corrosion state and corrosion pattern.

[0006] The present invention provides a pipeline steel corrosion pattern recognition method based on a random forest algorithm, which comprises the following steps:

[0007] Step 1, collecting electrochemical impedance information of pipeline corrosion electrochemical sensor;

[0008] Step 2: Process the collected electrochemical impedance information, classify and calibrate the electrochemical impedance data of each pipeline according to the electrochemical impedance spectroscopy equivalent circuit theory, and establish a pipeline corrosion mode database;

[0009] Step 3: Using the pipeline corrosion pattern database, a pipeline corrosion pattern classification model is obtained by training using a random forest algorithm;

[0010] Step 4: Substitute the pipeline operation monitoring data into the trained pipeline corrosion pattern classification model to obtain the current pipeline corrosion pattern and corrosion status.

[0011] In step 1, the electrochemical impedance information includes frequency and corresponding features, wherein the features include the real part of electrochemical impedance, the imaginary part of electrochemical impedance, the modulus of electrochemical impedance, the phase angle of electrochemical impedance, the collected current, and the collected voltage.

[0012] In step 1, in the electrochemical impedance information of each pipeline corrosion electrochemical sensor, the frequency points are F points selected in a uniform interval from 0.01 Hz to 1 MHz, and the characteristic data in the electrochemical impedance information is collected once at each frequency point. The electrochemical impedance information of a single pipeline corrosion electrochemical sensor contains F*7 dimensional features.

[0013] In step 2, each piece of pipeline electrochemical impedance information is classified and calibrated. According to the electrochemical impedance spectroscopy equivalent circuit theory, combined with the characteristics of the Bode diagram and Nyquist diagram of the pipeline electrochemical impedance information and the physical meaning of the corresponding pipeline corrosion process, it is divided into three modes: no corrosion, corrosion and loose rust layer, and corrosion and tight rust layer, and characterized by classification labels.

[0014] In step 3, the pipeline corrosion pattern database is used to train the pipeline corrosion pattern classification model using the random forest algorithm, including the following steps:

[0015] Step 3-1: Normalize and preprocess the data in the pipeline corrosion pattern database:

[0016]

[0017] Among them, D ij It represents the jth characteristic value in the collected pipeline electrochemical impedance data at the ith frequency point, Min j Represents the minimum value of feature j in the data set, Max j represents the maximum value of feature j in the data set, represents the pipeline electrochemical impedance data after normalization preprocessing;

[0018] Step 3-2, dividing the pipeline corrosion pattern database according to a set ratio to obtain a model training sample set and a test sample set;

[0019] Step 3-3, using a recursive feature elimination method to process the normalized pipeline electrochemical impedance data, removing irrelevant feature items in the data to reduce the data dimension;

[0020] Step 3-4: Use the random forest classification algorithm to train the training sample set, then use the test sample set to verify, and optimize the random forest algorithm parameters to finally obtain the optimized pipeline corrosion pattern classification model.

[0021] In step 3-3, the recursive feature elimination method is used to process the normalized pipeline electrochemical impedance data to remove irrelevant feature items in the data, specifically: by calculating the Pearson correlation coefficient between each feature in the data and the label value To filter:

[0022]

[0023] in, Represents the jth characteristic value of the normalized pipeline electrochemical impedance data The covariance between the output label value Y, σ Y Respectively and the standard deviation of Y;

[0024] Then the feature whose absolute value of the Pearson correlation coefficient is less than the threshold is calculated. Remove and realize feature screening of pipeline electrochemical impedance data set.

[0025] In step 3-4, the random forest classification algorithm is used to train the training sample set, and the following method is used to select features for decision tree training:

[0026] 1) First, sort the features in the training sample set in descending order according to the Pearson correlation coefficient of each feature;

[0027] 2) Then, starting from the feature with the largest Pearson correlation coefficient, M feature quantities are selected to construct a data subset and input into the random forest algorithm to train the first decision tree;

[0028] 3) Before training the next decision tree, add N new features in the order of the Pearson correlation coefficient, remove the N features with the largest Pearson correlation coefficient in the original data subset, and then train a new decision tree. Repeat step 3) until all the features in the training sample set have been traversed and eliminated.

[0029] In step 3-4, when the trained decision tree is constructed into a pipeline corrosion pattern classification model using the random forest algorithm, the weighted voting mechanism is used instead of the average voting mechanism, that is, the decision tree trained with the data subset constructed with the feature quantity whose Pearson correlation coefficient is greater than the threshold has a greater weight in the process of constructing the pipeline corrosion pattern classification model, as follows:

[0030]

[0031] Among them, T i (x) represents the classification result of the i-th decision tree, ω i represents the weight of the i-th decision tree, Y out Represents the final output classification result, where the order of decision trees is sorted from large to small according to the Pearson correlation coefficient of the feature quantity used to train the decision tree, and the weight of each decision tree ω i As follows:

[0032]

[0033] Among them, K is the number of decision trees.

[0034] A pipeline steel corrosion pattern recognition system based on random forest algorithm, comprising:

[0035] Electrochemical impedance information acquisition module, used to collect electrochemical impedance information of pipeline corrosion electrochemical sensors;

[0036] The database construction module is used to process the collected electrochemical impedance information, classify and calibrate the electrochemical impedance data of each pipeline based on the electrochemical impedance spectroscopy equivalent circuit theory, and establish a pipeline corrosion mode database;

[0037] A classification model building module is used to obtain a pipeline corrosion pattern classification model by using a pipeline corrosion pattern database and training with a random forest algorithm;

[0038] The corrosion pattern recognition module is used to substitute the pipeline operation monitoring data into the trained pipeline corrosion pattern classification model to obtain the current pipeline corrosion pattern and corrosion status.

[0039] The present invention has the following beneficial effects and advantages:

[0040] 1. High accuracy. In the method of the present invention, the pipeline corrosion pattern classification model obtained by training with the random forest algorithm is used to distinguish the pipeline corrosion pattern, which can effectively avoid the errors introduced by manual discrimination, and the pipeline corrosion pattern classification model obtained by training with the random forest algorithm can further improve the accuracy of recognition and judgment by adjusting the details of the classification model. In addition, when using the random forest algorithm for model training, the Pearson correlation coefficient of each feature quantity in the training sample set is sorted in order from large to small, and the phenomenon that some features are not extracted due to the random replacement extraction method is eliminated by traversal, which can improve the accuracy of the final training model. In the process of constructing the classification model, a weighted voting mechanism is adopted, so that the decision tree obtained by training the data subset constructed with the feature quantity with a larger Pearson correlation coefficient has a greater weight in the process of constructing the random forest model, which can truly reflect the influence of different feature quantities in the final decision, thereby improving the accuracy of label prediction.

[0041] 2. Real-time performance. Compared with the traditional electrochemical method, which manually analyzes and judges each frame of electrochemical impedance data, the method of the present invention can simultaneously analyze and process the electrochemical impedance data collected by multiple electrochemical corrosion sensors, greatly reducing the time for distinguishing pipeline corrosion patterns and providing a basis for real-time online monitoring of pipeline corrosion status.

[0042] 3. Low cost and easy to implement. In the method of the present invention, the pipeline corrosion pattern classification model obtained by training with the random forest algorithm is small in size and simple to deploy. It can be installed and deployed on a traditional industrial monitoring host without adding additional hardware. In addition, the electrochemical impedance spectroscopy measurement technology and random forest algorithm used in the present invention are mature application technologies, and the difficulty of implementation is relatively low. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 This is a schematic diagram of the process of a pipeline steel corrosion pattern recognition method based on a random forest algorithm according to the present invention;

[0044] Figure 2 This is a schematic diagram of a specific process of a pipeline steel corrosion pattern recognition method based on a random forest algorithm according to the present invention;

[0045] Figure 3 The electrochemical impedance data of the electrochemical sensor for corrosion of a single pipeline in the implementation example of the method of the present invention;

[0046] Figure 4 is a schematic diagram of the data structure of an example pipeline corrosion data set implemented in the method of the present invention;

[0047] Figure 5a The Bode diagram and Nyquist diagram of the electrochemical impedance data of the pipeline in the non-corrosion mode in the implementation example of the method of the present invention;

[0048] Figure 5b The Bode diagram and Nyquist diagram of the electrochemical impedance data of the pipeline in the embodiment of the method of the present invention in the mode where corrosion has occurred and the rust layer is loose;

[0049] Figure 5c The Bode diagram and Nyquist diagram of the electrochemical impedance data of the pipeline in the implementation example of the method of the present invention in the mode where corrosion has occurred and the rust layer is compact;

[0050] Figure 6 It is a schematic diagram of the structure of a random forest algorithm model implemented in the method of the present invention;

[0051] Figure 7 This is a comparison chart of the test results of the models obtained by training with common machine learning algorithms implemented in the method of the present invention;

[0052] Figure 8 Schematic diagram of the confusion matrix in the method of the present invention. DETAILED DESCRIPTION

[0053] The present invention is further described in detail below in conjunction with the accompanying drawings.

[0054] like Figure 1 As shown, a pipeline steel corrosion pattern recognition method based on a random forest algorithm is used to accurately and automatically identify the pipeline corrosion pattern according to the pipeline electrochemical impedance information collected by the electrochemical corrosion sensor, including the following steps:

[0055] First, various electrochemical impedance information of pipeline corrosion electrochemical sensors is collected; then, the collected electrochemical impedance information is processed, and the electrochemical impedance data of each pipeline is classified and calibrated according to the equivalent circuit theory of electrochemical impedance spectroscopy, and a pipeline corrosion pattern database is established; in the next step, the random forest algorithm is used to train the pipeline corrosion pattern database to obtain a pipeline corrosion pattern classification model; finally, the trained pipeline corrosion pattern classification model is used to process the pipeline operation monitoring data to identify and obtain the current pipeline corrosion pattern and corrosion status.

[0056] The collected electrochemical impedance information of the pipeline corrosion electrochemical sensor includes frequency, real part of electrochemical impedance, imaginary part of electrochemical impedance, electrochemical impedance modulus, electrochemical impedance phase angle, collected current, and collected voltage. The frequency points are 60 points selected from 0.01Hz to 1MHz in a proportional interval. At each frequency point, another six electrochemical impedance information are collected once. The electrochemical impedance data of a single pipeline corrosion electrochemical sensor contains 60*7=420 dimensional features. In order to build a pipeline corrosion pattern database, at least 3,000 electrochemical impedance data of pipeline corrosion electrochemical sensors need to be collected.

[0057] The classification and calibration of the collected pipeline electrochemical impedance data is based on the equivalent circuit theory of electrochemical impedance spectroscopy. It combines the characteristics of the Bode diagram and Nyquist diagram of the pipeline electrochemical impedance data and the physical meaning of the corresponding pipeline corrosion process to divide it into three modes: no corrosion, corrosion with loose rust layer, and corrosion with tight rust layer.

[0058] The pipeline corrosion pattern database is trained using the random forest algorithm to obtain a pipeline corrosion pattern classification model. The specific process is divided into the following four steps:

[0059] Step 1: The original data D of pipeline electrochemical impedance in the pipeline corrosion model database ij , the normalized pipeline electrochemical impedance data is obtained by processing according to the following formula:

[0060]

[0061] Step 2: Divide the pipeline corrosion pattern database according to the ratio of 8:2 to obtain the model training sample set and test sample set.

[0062] Step 3: Use recursive feature elimination method to calculate the Pearson correlation coefficient between each feature and the classification calibration value in the normalized pipeline electrochemical impedance data Then the features with an absolute value of Pearson correlation coefficient less than 0.7 will be calculated. Remove and realize feature screening of pipeline electrochemical impedance data set.

[0063] Step 4: Use the random forest classification algorithm to train the training sample set, then verify it with the test sample set, and optimize the random forest algorithm parameters to finally obtain the optimized pipeline corrosion pattern classification model.

[0064] The random forest classification algorithm trains the training sample set. When selecting features for decision tree training, the features in the training sample set are first sorted in descending order according to the Pearson correlation coefficient of each feature. Then, starting from the feature with the largest Pearson correlation coefficient, M feature quantities are selected to construct a data subset and input into the random forest algorithm for decision tree training. Before training the next decision tree, N new feature quantities are added in the order of the Pearson correlation coefficient, and the N feature quantities with larger Pearson correlation coefficients in the original data subset are removed. The above process is repeated until all feature quantities in the training sample set after feature elimination are traversed.

[0065] The random forest classification algorithm trains the training sample set, and uses the training to obtain each decision tree. When constructing the random forest model, a weighted voting mechanism is adopted. That is, the decision tree trained with the data subset constructed with a larger feature value of the Pearson correlation coefficient has a greater weight in the process of constructing the random forest model.

[0066] In order to further introduce the method of the present invention, the implementation process of each step of the method is described in detail below in conjunction with specific implementation examples. Figure 2 This is a schematic diagram of the specific process of the method of the present invention. This implementation example selects the pipeline electrochemical impedance data actually collected by the electrochemical corrosion sensor of a long-distance buried transmission pipeline, a total of 3052 data, and the electrochemical impedance data of a single pipeline corrosion electrochemical sensor is as follows Figure 3 As shown. The electrochemical impedance data of a single pipeline corrosion electrochemical sensor contains impedance information of 60 frequency points Freq. The impedance information of each frequency point includes the real part of electrochemical impedance Zreal, the imaginary part of electrochemical impedance Zimag, the electrochemical impedance modulus Zmod, the electrochemical impedance phase angle Zphz, the collected current Idc, and the collected voltage Vdc. The frequency points are 60 points selected from 0.01Hz to 1MHz in a geometric interval. Another six electrochemical impedance information are collected once at each frequency point. The electrochemical impedance data of a single pipeline corrosion electrochemical sensor contains 60*7=420 dimensional features. Then, in order to facilitate the subsequent classification and calibration of pipeline corrosion modes, the electrochemical impedance data of a single pipeline corrosion electrochemical sensor is expanded from the original 60*7 to 1*420 data. Then, the pipeline electrochemical impedance data actually collected by 3052 electrochemical corrosion sensors constitute the pipeline corrosion data set. The data set size is 3052*420. Figure 4 This is a schematic diagram of the data structure of the pipeline corrosion dataset.

[0067] After obtaining the pipeline corrosion data set, according to the electrochemical theory of metal corrosion, when the metal is in different corrosion modes, the equivalent circuit corresponding to the corrosion process is different, and the Bode plot and Nyquist plot of the corresponding pipeline electrochemical impedance data also have different characteristics. Figure 5a , Figure 5b , Figure 5c The Bode plots and Nyquist plots of typical pipeline electrochemical impedance data under three corrosion modes: no corrosion, corrosion and loose rust layer, and corrosion and tight rust layer. Then, the Bode plots and Nyquist plots of the pipeline electrochemical impedance data of each data in the pipeline corrosion data set are manually analyzed to determine which of the three corrosion modes it is closer to and mark the corresponding corrosion mode type on the data. Finally, the 3052 data in the pipeline corrosion data set are manually classified and calibrated to obtain the pipeline corrosion mode database.

[0068] Next, the pipeline corrosion pattern classification model is trained. First, the data in the pipeline corrosion pattern database is processed. Since the data values ​​of each dimension in the pipeline electrochemical impedance data are very different in dimension, in order to prevent the problem of gradient explosion and non-convergence during the training process, the data in the pipeline corrosion pattern database is normalized and preprocessed. The specific processing process is as follows:

[0069]

[0070] Where D ij Indicates the collected raw data of pipeline electrochemical impedance, Min j Represents the minimum value of feature j in the data set, Max j represents the maximum value of feature j in the data set, It represents the pipeline electrochemical impedance data after normalization preprocessing. Then the pipeline corrosion pattern database is divided according to the ratio of 8:2 to obtain the model training sample set and test sample set.

[0071] The dimension of each data in the training sample set and the test sample set is 420 dimensions. If the random Morin algorithm is used directly for training, not only will the model training speed be slow and the training time be long, but it will also easily cause the model to not converge, and ultimately an accurate training model cannot be obtained. Therefore, it is necessary to process the data in the training sample set and the test sample set to remove irrelevant feature items in the data and reduce the data dimension. This method chooses to use the recursive feature elimination method to screen the data. First, the Pearson correlation coefficient between each feature and the label value in the data is calculated according to the following formula:

[0072]

[0073] in Represents the characteristics of pipeline electrochemical impedance data The covariance between the output label value Y, σ Y Respectively The Pearson correlation coefficient can be used to characterize the correlation between the feature quantity and the output label value. The absolute value of the Pearson correlation coefficient is close to 1, indicating that the higher the correlation between the feature quantity and the output label value, the closer the absolute value of the Pearson correlation coefficient is to 0, the lower the correlation between the feature quantity and the output label value. Therefore, the feature quantity with an absolute value of the Pearson correlation coefficient less than 0.7 is calculated. After feature screening and dimensionality reduction, the data dimension in the data set is reduced from 420 dimensions to 176, which can effectively reduce the subsequent model training time and improve the model training accuracy.

[0074] The next step is to use the training sample set to train the model according to the random forest algorithm. The random forest algorithm model structure is as follows: Figure 6 As shown, the traditional random forest algorithm training process is as follows:

[0075] The first step is to select Q samples and M features for training a decision tree;

[0076] The second step is to repeat K times, selecting different samples and features each time, and training K different decision trees;

[0077] In the third step, for a new input sample, K decision trees are used for classification, and finally a voting mechanism is used to determine its category.

[0078] In the traditional random forest algorithm training process, the feature selection used to train the decision tree adopts the random replacement extraction method, which may cause some features not to be extracted, thereby resulting in a decrease in the discrimination accuracy of the model obtained by the final training. In addition, when using the decision tree to construct a random forest training model, a voting mechanism is generally used to determine the category, but this mechanism cannot truly reflect the influence of feature quantities of different correlation degrees in the final decision. Therefore, in view of the problems existing in the traditional random deep forest model training method, this method first sorts the feature quantities in the training sample set according to the Pearson correlation coefficient of each feature quantity in the training sample set from large to small, and then starts from the feature quantity with the largest Pearson correlation coefficient, selects M feature quantities to construct a data subset and inputs it into the random forest algorithm for decision tree training; before training the next decision tree, N new feature quantities are added in the order of the size of the Pearson correlation coefficient, and the N feature quantities with larger Pearson correlation coefficients in the original data subset are removed, and the above process is repeated until all feature quantities in the training sample set after feature elimination are traversed. In this embodiment, the remaining 176-dimensional features after the 420-dimensional feature reduction are traversed. The selection of M and N is based on comprehensive considerations such as the feature dimension of the data, the time required for training the model, and the accuracy of the training model. In this implementation example, M is selected as 40 and N is 4. After training with the random forest algorithm, a total of 34 decision trees are obtained.

[0079] After training multiple decision trees, a voting mechanism is needed to combine the classification results of multiple decision trees to obtain the final output result. The traditional random algorithm uses an average voting mechanism, that is, the classification results of each decision tree have the same impact on the final output result. This method uses a weighted voting mechanism. The specific formula is as follows:

[0080]

[0081] Where T i (x) represents the classification result of the i-th decision tree, ω i represents the weight of the i-th decision tree, Y out Represents the final output classification results, including no corrosion, corrosion and loose rust layer, corrosion and tight rust layer. The order of decision trees is sorted from large to small according to the Pearson correlation coefficient of the feature quantity used in training the decision tree. The weight of each decision tree is ω i As shown in the following formula, in this embodiment, K=34.

[0082]

[0083] After the above steps, the corrosion mode classification model trained by the training sample set is obtained, and then the test sample set is used to evaluate and optimize the model. Figure 7This is a comparison chart of the test results of the model trained by this method and the models trained by several other common machine learning algorithms using the test sample set. Figure 7 The RF on the far left of the horizontal axis represents the model trained by this method. The four curves represent the accuracy, recall, F1 score and precision effects respectively. Figure 7 It can be found that the accuracy of the model trained by this method not only meets the application requirements, but also has better performance indicators than the models trained by other machine learning algorithms. Finally, the newly acquired 159 pipeline operation monitoring data are input into the pipeline corrosion pattern classification model, and then the pipeline corrosion pattern classification model identification results and manual calibration results are used to draw Figure 8 The confusion matrix diagram shown is Figure 8 It can be found that the three types of corrosion patterns identified by the pipeline corrosion pattern classification model are completely consistent with the results of manual calibration.

Claims

1. A pipeline steel corrosion pattern recognition method based on random forest algorithm, characterized in that: The following steps are involved: Step 1, collecting electrochemical impedance information of pipeline corrosion electrochemical sensor; Step 2: Process the collected electrochemical impedance information, classify and calibrate the electrochemical impedance data of each pipeline according to the electrochemical impedance spectroscopy equivalent circuit theory, and establish a pipeline corrosion mode database; Step 3: Using the pipeline corrosion pattern database, a pipeline corrosion pattern classification model is obtained by training using a random forest algorithm; Step 4: Substitute the pipeline operation monitoring data into the trained pipeline corrosion pattern classification model to obtain the current pipeline corrosion pattern and corrosion status.

2. The pipeline steel corrosion pattern recognition method based on random forest algorithm according to claim 1 is characterized in that: In step 1, the electrochemical impedance information includes frequency and corresponding features, wherein the features include the real part of electrochemical impedance, the imaginary part of electrochemical impedance, the modulus of electrochemical impedance, the phase angle of electrochemical impedance, the collected current, and the collected voltage.

3. The pipeline steel corrosion pattern recognition method based on random forest algorithm according to claim 1 is characterized in that: In step 1, in the electrochemical impedance information of each pipeline corrosion electrochemical sensor, the frequency points are F points selected in a uniform interval from 0.01 Hz to 1 MHz, and the characteristic data in the electrochemical impedance information is collected once at each frequency point. The electrochemical impedance information of a single pipeline corrosion electrochemical sensor contains F*7 dimensional features.

4. The pipeline steel corrosion pattern recognition method based on random forest algorithm according to claim 1 is characterized in that: In step 2, each piece of pipeline electrochemical impedance information is classified and calibrated. According to the electrochemical impedance spectroscopy equivalent circuit theory, combined with the characteristics of the Bode diagram and Nyquist diagram of the pipeline electrochemical impedance information and the physical meaning of the corresponding pipeline corrosion process, it is divided into three modes: no corrosion, corrosion and loose rust layer, and corrosion and tight rust layer, and characterized by classification labels.

5. The pipeline steel corrosion pattern recognition method based on random forest algorithm according to claim 1 is characterized in that: In step 3, the pipeline corrosion pattern database is used to train the pipeline corrosion pattern classification model using the random forest algorithm, including the following steps: Step 3-1: Normalize and preprocess the data in the pipeline corrosion pattern database: Among them, D ij It represents the jth characteristic value in the collected pipeline electrochemical impedance data at the ith frequency point, Min j Represents the minimum value of feature j in the data set, Max j represents the maximum value of feature j in the data set, represents the pipeline electrochemical impedance data after normalization preprocessing; Step 3-2, dividing the pipeline corrosion pattern database according to a set ratio to obtain a model training sample set and a test sample set; Step 3-3, using a recursive feature elimination method to process the normalized pipeline electrochemical impedance data, removing irrelevant feature items in the data to reduce the data dimension; Step 3-4: Use the random forest classification algorithm to train the training sample set, then use the test sample set to verify, and optimize the random forest algorithm parameters to finally obtain the optimized pipeline corrosion pattern classification model.

6. The pipeline steel corrosion pattern recognition method based on random forest algorithm according to claim 5 is characterized in that: In step 3-3, the recursive feature elimination method is used to process the normalized pipeline electrochemical impedance data to remove irrelevant feature items in the data, specifically: by calculating the Pearson correlation coefficient between each feature in the data and the label value To filter: in, Represents the jth characteristic value of the normalized pipeline electrochemical impedance data The covariance between the output label value Y, σ Y Respectively and the standard deviation of Y; Then the feature whose absolute value of the Pearson correlation coefficient is less than the threshold is calculated. Remove and realize feature screening of pipeline electrochemical impedance data set.

7. The pipeline steel corrosion pattern recognition method based on random forest algorithm according to claim 5 is characterized in that: In step 3-4, the random forest classification algorithm is used to train the training sample set, and the following method is used to select features for decision tree training: 1) First, sort the features in the training sample set in descending order according to the Pearson correlation coefficient of each feature; 2) Then, starting from the feature with the largest Pearson correlation coefficient, M feature quantities are selected to construct a data subset and input into the random forest algorithm to train the first decision tree; 3) Before training the next decision tree, add N new features in the order of the Pearson correlation coefficient, remove the N features with the largest Pearson correlation coefficient in the original data subset, and then train a new decision tree. Repeat step 3) until all the features in the training sample set have been traversed and eliminated.

8. The pipeline steel corrosion pattern recognition method based on random forest algorithm according to claim 5 is characterized in that: In step 3-4, when the trained decision tree is constructed into a pipeline corrosion pattern classification model using the random forest algorithm, the weighted voting mechanism is used instead of the average voting mechanism, that is, the decision tree trained with the data subset constructed with the feature quantity whose Pearson correlation coefficient is greater than the threshold has a greater weight in the process of constructing the pipeline corrosion pattern classification model, as follows: Among them, T i (x) represents the classification result of the i-th decision tree, ω i represents the weight of the i-th decision tree, Y out Represents the final output classification result, where the order of decision trees is sorted from large to small according to the Pearson correlation coefficient of the feature quantity used to train the decision tree, and the weight of each decision tree ω i As follows: Among them, K is the number of decision trees.

9. A pipeline steel corrosion pattern recognition system based on random forest algorithm, characterized in that: include: Electrochemical impedance information acquisition module, used to collect electrochemical impedance information of pipeline corrosion electrochemical sensors; The database construction module is used to process the collected electrochemical impedance information, classify and calibrate the electrochemical impedance data of each pipeline based on the electrochemical impedance spectroscopy equivalent circuit theory, and establish a pipeline corrosion mode database; A classification model building module is used to obtain a pipeline corrosion pattern classification model by using a pipeline corrosion pattern database and training with a random forest algorithm; The corrosion pattern recognition module is used to substitute the pipeline operation monitoring data into the trained pipeline corrosion pattern classification model to obtain the current pipeline corrosion pattern and corrosion status.

Citation Information

Cited By

  • Municipal pipeline defect classification method and device based on intelligent algorithm and electronic equipment

    CN120726406A