Data-driven intelligent identification method for corrosion failure causes in steel oil and gas pipelines
By using a data-driven approach, internal detection data is collected and processed to construct a multi-source dataset. Mechanism and machine learning models are then used to identify the causes of corrosion in steel oil and gas pipelines. This solves the problem of bias in identification results under the coupling of multiple factors, enabling rapid and accurate recommendations for protective measures and improving pipeline maintenance efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING YUNZHIYONGLI TECHNOLOGY CO LTD
- Filing Date
- 2025-09-12
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies, especially in scenarios involving multiple coupled factors, tend to produce biased results when identifying the causes of corrosion in steel oil and gas pipelines, and lack the ability to quickly analyze and automatically recommend protective measures, resulting in low pipeline maintenance efficiency.
Using a data-driven approach, a multi-source dataset is constructed by collecting, processing, and aligning internal inspection data and working condition media data. Typical internal corrosion defects and working condition characteristics are extracted, and the causes of corrosion are identified using mechanism and machine learning models. Protective measures are then automatically recommended.
It enables rapid and accurate identification of corrosion causes in steel oil and gas pipelines and automatic recommendation of protective measures, thereby improving the level of intelligence in pipeline integrity management.
Smart Images

Figure CN121211933B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of identification technology for the causes of corrosion failure in steel pipelines in ground engineering, and specifically to a data-driven intelligent identification method for the causes of corrosion failure in steel oil and gas pipelines. Background Technology
[0002] Steel oil and gas pipelines used in long-distance transportation of oil and natural gas inevitably experience corrosion, leaks, and even breakage. Timely detection of corrosion and understanding of its causes are crucial for the safe operation and routine maintenance of pipelines. Influenced by factors such as temperature, pressure, corrosive gases, media, and flow rate, the causes of internal pipeline corrosion are diverse and often involve multiple factors coupled together. During routine pipeline inspection and maintenance, pipeline personnel conduct internal inspections to obtain information on the morphology and spatial distribution of defects within the pipeline. The information contained in internal inspections is an external manifestation of the effects of various factors within the pipeline. Therefore, pipeline experts use internal inspection information and service environment information to analyze and identify the causes of internal corrosion and formulate corrosion protection strategies based on these causes. This process is time-consuming and the identification results can be biased, especially in scenarios involving multiple coupled factors. Driven by the evolving needs of pipeline integrity management, there is an urgent need for a systematic approach to rapidly analyze, identify, and formulate countermeasures for internal corrosion. This approach should automatically recommend internal corrosion control strategies based on the different types and severity of internal corrosion failures, enabling timely prevention and control of internal corrosion occurrence and development, and further supporting the intelligent enhancement of pipeline integrity management. Summary of the Invention
[0003] This invention aims to identify the causes of internal corrosion failure in steel oil and gas pipelines and recommend internal corrosion protection strategies, addressing the limitations and biases inherent in human intervention when facing complex situations involving multiple coupled factors. The main contents (or steps) of this invention include: acquisition of internal inspection data and operating medium data; processing of internal inspection data, alignment of operating medium data, and construction of multi-source datasets; analysis of internal corrosion causes and classification of cause types; extraction of typical internal corrosion defect features and service environment features; construction of a pipeline internal corrosion failure cause identification model; construction of pipeline internal corrosion control strategies; and software tool development and application. This method uses pipeline internal inspection data and pipeline service data as input, utilizes a feature extraction module to extract important features of internal corrosion defects and the service environment, and uses an internal corrosion failure cause module to identify the causes of internal corrosion. Based on the identification results, it recommends internal corrosion protection strategies, achieving rapid identification of corrosion causes and accurate recommendation of protection strategies.
[0004] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A data-driven intelligent identification method for the causes of corrosion failure in steel oil and gas pipelines, comprising the following steps: S1: Acquisition of internal detection data and operating condition media data; S2: Internal detection data processing, working condition media data alignment, and multi-source dataset construction: The internal detection data processing involves unifying and standardizing the collected internal detection data, including the following: 1) Filter out data with internal defects based on defect type; 2) Calculate the distance (m) to the nearest circumferential weld based on the distance (m) from the upstream circumferential weld, the distance (m) from the downstream circumferential weld, and the length (m) of the pipe section; 3) For internal inspection data that lacks defect depth information, the defect depth information is roughly estimated using the wall thickness reduction (%) and pipe wall thickness. Defect depth = pipe wall thickness × wall thickness reduction. S3: Internal Corrosion Cause Analysis and Cause Type Classification: The internal corrosion cause analysis and cause type classification refers to obtaining all failure cause types of internal corrosion by analyzing and classifying the internal corrosion causes of oil and gas pipelines. S4: Typical Internal Corrosion Defect Features, Operating Condition Medium Feature Extraction, and Feature Extraction Module Construction: The typical internal corrosion defect features and operating condition medium feature extraction include the number of defects per kilometer, defect spatial distribution features, corrosion severity features, defect morphology features, scour corrosion features, waterline corrosion features, weld corrosion features, top corrosion features, and C. O2 Corrosion characteristics, CO2 / H2S corrosion characteristics, bacterial corrosion characteristics, and characteristics of the working medium; S5: Construction of Pipeline Corrosion Failure Cause Identification Model: The construction of the pipeline corrosion failure cause identification model includes the construction of a corrosion failure cause identification model for purification pipelines and a corrosion failure cause identification model for gathering and transportation pipelines; among them, the corrosion failure cause identification model for gathering and transportation pipelines includes a mechanism-based failure cause identification model and a machine learning-based failure cause identification model, for a total of three internal corrosion failure cause identification models are constructed, all of which are constructed using Python; S6: Construction of Corrosion Control Strategies in Pipelines: The construction of corrosion control strategies in pipelines involves building corresponding control strategies based on the failure cause type and its severity; when the corresponding indicator is triggered, control strategies are automatically recommended.
[0005] Preferably, the internal detection data and operating condition medium data acquisition in S1 includes the following data: The internal inspection data collection includes: mileage location (km), elevation location, clock position, distance from upstream circumferential weld (m), distance from downstream circumferential weld (m), pipe section length (m), defect length (mm), width (mm), depth (mm), and wall thickness reduction (%). The data collected for the operating medium includes: pipeline name, pipeline type, transported medium, pipeline length (km), pipeline wall thickness (mm), starting temperature (°C), ending temperature (°C), starting pressure (MPa), ending pressure (MPa), and daily liquid production (m³). 3 / day), daily gas production (m³) 3 / day), water content (%), gas-liquid ratio (%), flow rate (m / s), CO2 content (mol%), H2S content (ppm), O2 content (mol%), mineralization (mg / L), K + Content (mg / L), Na + Content (mg / L), Ba 2+ Content (mg / L), Sr 2+ Content (mg / L), Ca 2+ Content (mg / L), Mg 2+ Content (mg / L), Fe 2+ Content (mg / L), Cl - Content (mg / L), CO3 2- Content (mg / L), HCO3 3- Content (mg / L), SO4 2- Content (mg / L), pH value, bacterial count (cFU / mL), presence or absence of scale.
[0006] Preferably, the four possible scenarios in step 2) of S2 and their corresponding calculation methods are as follows: ① Only information on the distance (m) from the upstream circumferential weld and the distance (m) from the downstream circumferential weld is available; Calculation method: Take the minimum value of the two; ② Only the distance (m) from the upstream circumferential weld and the length (m) of the pipe section are available; Calculation method: Subtract the distance from the upstream circumferential weld from the pipe section length to obtain the distance from the downstream circumferential weld, and then take the minimum value between the distances from the upstream and downstream circumferential welds; ③ Only information on the distance (m) from the downstream circumferential weld and the length (m) of the pipe section is available; Calculation method: Subtract the distance to the downstream circumferential weld from the pipe section length to obtain the distance to the upstream circumferential weld, and then take the minimum value between the distances to the upstream and downstream circumferential welds; ④ Includes information on the distance (m) from the upstream circumferential weld, the distance (m) from the downstream circumferential weld, and the length (m) of the pipe section; Calculation method: The calculation method of case ① shall be adopted.
[0007] Preferably, the extraction method for each step in step S4 is as follows: 1) Defect Quantity Characteristics per Kilometer (N_L): The defect quantity characteristic per kilometer is equal to the ratio of the total number of corrosion defects (N) in the pipeline to the total pipeline length (L), that is: 2) Spatial distribution characteristics of defects: Among all defects, calculate the probability of a defect being located at the 1 to 12 o'clock positions sequentially, denoted as P1 to P12. For example, the 1 o'clock position is 0:30 to 1:30. The calculation formula is as follows: 3) Corrosion severity characteristics: Among all defects, the proportions of defects with wall thickness reduction of less than 10%, 20%, 30%, 40%, and 50% are calculated sequentially and denoted as P_10, P_20, P_30, P_40, and P_50, respectively. The calculation formulas are as follows: 4) Defect morphology characteristics: Three morphologies are used to describe the morphology of the defects: equiaxed (0.5 <= length / width <= 2), elongated (length / width >= 10), and vertical (length / width <= 0.1). The proportion of each of the three morphologies is calculated using the following formula: in, , , These represent the number of equiaxed defects, elongated defects, and vertical defects, respectively. 5) Erosion corrosion characteristics: Based on the three-dimensional morphological characteristics of erosion corrosion and the clock position, calculate the proportion of defects with a length / width > 10 near the 4-8 o'clock position. , Defined as a characteristic of erosion corrosion, the calculation formula is as follows: in, The number of defects with a length / width greater than 10 around points 4-8. This represents the total number of defects between 4 and 8 o'clock. 6) Characteristics of waterline corrosion: The three-dimensional morphology of waterline corrosion is characterized by a beaded pattern, with multiple defects continuously distributed along the interface. Based on this, the number of pipe segments with a continuous distribution of defects at points 1-4 and 8-11 (defect spacing less than 10cm) and a length greater than 1m (a segment is defined as longer than 1m) is calculated using the following formula: in, express The number of pipe sections with continuous distribution of clockwork defects (defect spacing less than 10cm) and a length greater than 1m. =1,2,3,4,8,9,10,11,12; 7) Weld corrosion characteristics: Calculate the proportion of defects located less than 10cm from the nearest circumferential weld. The calculation formula is as follows: in, This indicates the number of defects that are less than 10cm away from the nearest circumferential weld. 8) Top corrosion characteristics: Calculate the proportion of corrosion pits near 12 o'clock (11:30-12:30), using the formula shown below: 9) CO2 corrosion characteristics: Calculate the proportion of plateau-shaped defects at the 4-8 o'clock position. The calculation formula is as follows: in, The number of plateau-shaped defects between 4 and 8 o'clock. This represents the total number of defects between 4 and 8 o'clock. 10) CO2 / H2S corrosion characteristics: Calculate the proportion of pitting defects at 4-8 o'clock. The calculation formula is as follows: in, This indicates the number of pitting defects between the 4 and 8 o'clock positions. This represents the total number of defects between 4 and 8 o'clock. 11) Bacterial corrosion characteristics: Calculate the probability of having more than 10 equiaxed defects within a 1m interval in the direction from point 4 to point 8. The calculation formula is as follows: in, This indicates the number of pipe segments with more than 10 equiaxed defects within a 1m interval in the direction from 4 to 8 o'clock (1m is 1 segment). The total length of the pipeline, in km; 12) Operating medium characteristics – temperature: Calculate the average of the starting and ending temperatures, using the formula shown below: 13) Operating medium characteristics -- CO2 partial pressure: Convert the CO2 content (in percentage) in the original data to MPa. The calculation formula is as follows: Where A represents the CO2 volume percentage, in % %. This is the partial pressure of CO2, expressed in MPa. , These are the starting and ending pressures of the pipeline operation, respectively. 14) Operating medium characteristics -- H2S partial pressure: Convert the H2S content in the original data (in ppm) to MPa. The calculation formula is as follows: in, for Content, unit MPa; T is the pipeline operating temperature; calculated by formula (14); ppm is the H2S content, unit ppm; , These are the starting and ending pressures of the pipeline operation, respectively. 15) Operating medium characteristics -- CO2 / H2S partial pressure ratio: Calculate the ratio of the partial pressures of CO2 and H2S. The calculation formula is shown below: Where R is the ratio of the partial pressures of CO2 to H2S; 16) Operating medium characteristics - bacterial content: directly extract bacterial content information; 17) Operating medium characteristics – presence or absence of scale / structure factor: Calculate the structure factor for four scale types: calcium carbonate, calcium sulfate, barium sulfate, and strontium sulfate. The calculation formula is shown below: ① Calcium carbonate scaling factors: in, , , , , , , , , , They are respectively , , , , , , , , , The content of is expressed in mg / L; μ is the ionic strength in mol / L; AIK is the total alkalinity in mol / L; K is the correction factor, which is determined by both temperature and ionic strength μ and is dimensionless; SAI is the scaling factor and is dimensionless. ② Calcium sulfate scaling factors: in, , They are respectively , Content, in mg / L; X is and The concentration difference, in mol / L; is the solubility product constant, dimensionless; S is the predicted value of calcium sulfate scaling tendency, in mol / L; P is the calcium sulfate scaling factor. ③ Barium sulfate scaling factor: in, , They are respectively , Content, in mg / L; m and a are the initial conditions in the water. , The concentration is expressed in mol / L. is the solubility product of BaSO4, which is temperature-dependent and dimensionless; B is the scaling factor of barium sulfate in the water after the water quality has stabilized, with units of mol / L. The amount of BaSO4 scale, in mg / L; ④ Strontium sulfate scaling factor: in, , They are respectively , Content, in mg / L; The solubility product of strontium sulfate is taken as 2.8 × 10⁻⁶. -7 W represents the strontium sulfate scaling factor.
[0008] Preferably, the three models in S5 include a purification pipeline failure cause identification model, a mechanism-based gathering and transportation pipeline failure cause identification model, and a ML-based gathering and transportation pipeline failure cause identification model.
[0009] This is a data-driven intelligent identification software for the causes of corrosion failure in steel oil and gas pipelines. The software's functional modules include a homepage, data management module, cause identification module, expert verification module, control strategy module, model management module, and system management module. The functions of each module are described below: Homepage: This module mainly showcases the data visualization of the system's main functions; Data Management Module: This module mainly implements unit management, pipeline management, and data dictionary; Cause identification module: This module mainly realizes data reporting, failure identification, and result viewing; Data reporting module: This module mainly involves inputting various parameters and displaying the input results through data reporting; Expert Verification Module: This module mainly manages verification and allows users to view results. It allows manual verification of the cause identification results to ensure their accuracy and allows users to view the verified results. Control Strategy Module: This module mainly enables the generation of countermeasures, viewing of results, and tracking of effects; Model Management Module: This module mainly implements dataset management, model training, and model application. The dataset management module is responsible for the statistics of all validation datasets and can add confirmed correct data from external sources. Model training uses the extracted feature dataset to train the XGBoost model. Model application loads the relevant model training parameters into the model, enabling it to predict the causes of failure. System Management Module: This module mainly implements failure target management, user management, menu management, role management, and log management; it also mainly implements the collection and display of user permissions and related error messages for the system.
[0010] Compared with the prior art, the present invention has the following beneficial effects. This invention discloses a method for intelligent identification of corrosion failure causes in steel oil and gas pipelines. This method automates the entire process from data upload to data processing, feature extraction, model selection, cause identification, and protective measure recommendation. Those skilled in the art can use the algorithms and processes provided by this invention to automatically and rapidly identify failure causes, thereby achieving similar functionality to this invention.
[0011] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description
[0012] Figure 1 This is a flowchart of the intelligent identification algorithm for the causes of corrosion failure in steel oil and gas pipelines according to the present invention. Figure 2 This is a schematic diagram of data alignment according to the present invention; Figure 3 This is a schematic diagram illustrating the selection of the failure cause identification model for the present invention; Figure 4 This is a schematic diagram of the purification pipeline failure cause identification model of the present invention; Figure 5 This is a schematic diagram of the mechanism-based failure cause identification model for gathering and transportation pipelines of the present invention. Figure 6 This is a flowchart of the training and testing process for the ML-based failure cause identification model for gathering and transportation pipelines of the present invention. Figure 7 This is a schematic diagram illustrating the application of the ML-based failure cause identification model for gathering and transportation pipelines of the present invention. Figure 8 This is a schematic diagram of the functional modules of the intelligent identification software for the causes of corrosion failure in pipelines according to the present invention. Detailed Implementation
[0013] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with specific embodiments.
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] The intelligent identification method for corrosion failure causes in steel oil and gas pipelines provided by this invention can identify the causes of corrosion failure in pipelines and provide corresponding corrosion protection countermeasures. This invention uses oilfield purification and gathering steel pipelines as an example for illustrative purposes, but it is not limited to the identification of corrosion causes in oilfield purification and gathering steel pipelines; the identification of corrosion causes in any pipeline is applicable to this invention.
[0017] The intelligent identification method for the causes of corrosion failure in steel oil and gas pipelines provided by this invention is as follows: Figure 1 As shown, Figure 1 A flowchart illustrating an intelligent identification method for the causes of corrosion failure in steel oil and gas pipelines according to an embodiment of the present invention is shown, specifically including: S1: Acquisition of internal detection data and operating condition media data The internal inspection data collection includes the type of defect, mileage location (km), elevation location, clock position, distance from upstream circumferential weld (m), distance from downstream circumferential weld (m), pipe section length (m), defect length (mm), width (mm), depth (mm), and wall thickness reduction (%).
[0018] The data collected for the operating medium includes: pipeline name, pipeline type, transported medium, pipeline length (km), pipeline wall thickness (mm), starting temperature (°C), ending temperature (°C), starting pressure (MPa), ending pressure (MPa), and daily liquid production (m³). 3 / day), daily gas production (m³) 3 / day), water content (%), gas-liquid ratio (%), flow rate (m / s), CO2 content (mol %), H2S content (ppm), O2 content (mol %), mineralization (mg / L), K + Content (mg / L), Na + Content (mg / L), Ba 2+ Content (mg / L), Sr 2+ Content (mg / L), Ca 2+ Content (mg / L), Mg 2+ Content (mg / L), Fe 2+ Content (mg / L), Cl - Content (mg / L), CO3 2- Content (mg / L), HCO3 - Content (mg / L), SO4 2- Content (mg / L), pH value, bacterial count (cFU / mL), presence or absence of scale.
[0019] S2: Internal detection data processing, working condition media data alignment, and multi-source dataset construction The internal detection data processing involves standardizing and unifying the collected internal detection data, including the following: 1) Filter out data with internal defects based on defect type; 2) Calculate the distance (m) to the nearest circumferential weld based on the distance (m) from the upstream circumferential weld, the distance (m) from the downstream circumferential weld, and the pipe section length (m). There are four possible scenarios: ① Only information on the distance (m) from the upstream circumferential weld and the distance (m) from the downstream circumferential weld is available; Calculation method: Take the minimum value of the two.
[0020] ② Only information on distance from the upstream circumferential weld (m) and pipe section length (m) is available; Calculation method: Subtract the distance from the upstream circumferential weld from the pipe section length to obtain the distance from the downstream circumferential weld, and then take the minimum value between the distances from the upstream and downstream circumferential welds.
[0021] ③ Only information on distance from downstream circumferential weld (m) and pipe section length (m) is available; Calculation method: Subtract the distance to the downstream circumferential weld from the pipe section length to obtain the distance to the upstream circumferential weld, and then take the minimum value between the distances to the upstream and downstream circumferential welds.
[0022] ④ Includes information on the distance (m) from the upstream circumferential weld, the distance (m) from the downstream circumferential weld, and the length (m) of the pipe section.
[0023] Calculation method: The calculation method of case ① shall be adopted.
[0024] 3) For internal inspection data that lacks defect depth information, the defect depth information is roughly estimated using the wall thickness reduction (%) and pipe wall thickness. Defect depth = pipe wall thickness × wall thickness reduction.
[0025] The aforementioned alignment of operating condition media data refers to aligning internal inspection data and operating condition media data belonging to the same pipeline. The alignment method of this invention involves adding a column to the pipeline's operating condition media data table, named "Internal Inspection Data Table Name." Thus, each row in the operating condition media data table contains both internal inspection data and operating condition media data, such as... Figure 2 As shown.
[0026] The aforementioned multi-source dataset construction refers to storing all pipeline data in the same data table using the data alignment method described above. This data table is called the internal corrosion failure multi-source dataset. This dataset serves as the original dataset for subsequent defect feature extraction, model training, and failure cause identification.
[0027] S3: Analysis of the causes and classification of internal corrosion types The aforementioned analysis and classification of internal corrosion causes refers to the process of analyzing and classifying the causes of internal corrosion in oil and gas pipelines to obtain all types of failure causes related to internal corrosion. Specifically, this includes the following: Based on the three-dimensional morphological characteristics, spatial distribution characteristics, and operating conditions of internal corrosion in oil and gas pipelines, the causes and reverse analysis of internal corrosion failure are summarized, and the results are shown in Table 1. Nine typical types of internal corrosion failure are summarized: CO2 corrosion, CO2 / H2S corrosion, erosion corrosion, dissolved oxygen corrosion, top corrosion, bacterial corrosion, weld corrosion, waterline corrosion, and under-deposit corrosion, providing a theoretical basis for the feature extraction module.
[0028] Table 1 Typical Causes of Corrosion Failure in Oil and Gas Pipelines S4: Typical Internal Corrosion Defect Characteristics, Working Condition Medium Feature Extraction, and Feature Extraction Module Construction The extraction of typical internal corrosion defect features and operating medium characteristics includes features such as the number of defects per kilometer, spatial distribution of defects, corrosion severity, defect morphology, erosion corrosion, waterline corrosion, weld corrosion, top corrosion, CO2 corrosion, CO2 / H2S corrosion, bacterial corrosion, and operating medium characteristics. The extraction methods are based on the theoretical framework provided in Table 1. The extraction methods for each feature are described below: 1) Defect quantity characteristics per kilometer (N_L) The number of defects per kilometer is equal to the ratio of the total number of corrosion defects (N) in the pipeline to the total length of the pipeline (L), that is: 2) Spatial distribution characteristics of defects Among all defects, calculate the probability of a defect being located at position 1 to 12 o'clock, denoted as P1 to P12. For example, position 1 o'clock is from 0:30 to 1:30. The calculation formula is as follows: 3) Characteristics of corrosion severity Among all defects, the percentages of defects with wall thickness reduction of less than 10%, 20%, 30%, 40%, and 50% are calculated sequentially and denoted as P_10, P_20, P_30, P_40, and P_50, respectively. The calculation formulas are as follows: 4) Defect morphology characteristics The morphology of the defect is described using three shapes: equiaxed (0.5 <= length / width <= 2), elongated (length / width >= 10), and vertical (length / width <= 0.1). The proportion of each of the three defect shapes is calculated using the following formulas: in, , , These represent the number of equiaxed defects, elongated defects, and vertical defects, respectively.
[0029] 5) Characteristics of erosion corrosion Based on the three-dimensional morphological characteristics of erosion corrosion and the clock position, the proportion of defects with a length / width > 10 in the vicinity of 4-8 o'clock was calculated. , Defined as a characteristic of erosion corrosion, the calculation formula is as follows: in, The number of defects with a length / width greater than 10 around points 4-8. This represents the total number of defects between 4 and 8 o'clock.
[0030] 6) Characteristics of waterline corrosion The three-dimensional morphological characteristics of waterline corrosion are beaded, with multiple defects continuously distributed along the interface. Based on this, the number of pipe segments (more than 1m in length) with continuous defect distribution at points 1-4 and 8-11 (defect spacing less than 10cm) is calculated (a segment is defined as longer than 1m). The calculation formula is as follows: in, express The number of pipe sections with continuous distribution of clockwork defects (defect spacing less than 10cm) and a length greater than 1m. =1, 2, 3, 4, 8, 9, 10, 11, 12.
[0031] 7) Weld corrosion characteristics Calculate the proportion of defects that are less than 10cm away from the nearest circumferential weld. The calculation formula is as follows: in, This indicates the number of defects that are less than 10cm away from the nearest circumferential weld.
[0032] 8) Top corrosion characteristics The following formula is used to calculate the proportion of erosion pits around 12:00 (11:30-12:30): 9) CO2 corrosion characteristics Calculate the proportion of plateau-like defects between 4 and 8 o'clock. The calculation formula is as follows: in, The number of plateau-shaped defects between 4 and 8 o'clock. This represents the total number of defects between 4 and 8 o'clock.
[0033] 10) CO2 / H2S corrosion characteristics Calculate the proportion of pitting defects at 4-8 o'clock. The calculation formula is as follows: in, This indicates the number of pitting defects between the 4 and 8 o'clock positions. This represents the total number of defects between 4 and 8 o'clock.
[0034] 11) Characteristics of bacterial corrosion The probability of having more than 10 equiaxed defects within a 1m interval in the direction from point 4 to point 8 is calculated using the following formula: in, This indicates the number of pipe segments with more than 10 equiaxed defects within a 1m interval in the direction from 4 to 8 o'clock (1m is 1 segment). This represents the total length of the pipeline, in km.
[0035] 12) Operating medium characteristics -- temperature The average of the starting and ending temperatures is calculated using the following formula: 13) Operating medium characteristics -- CO2 partial pressure Convert the CO2 content (in percentage) in the original data to MPa using the following formula: Where A represents the CO2 volume percentage, in % %. This is the partial pressure of CO2, expressed in MPa. , These are the starting and ending pressures of the pipeline operation, respectively.
[0036] 14) Operating medium characteristics -- H2S pressure divider Convert the H2S content in the original data, which is in ppm, to MPa using the following formula: in, for Content, unit MPa; T is the pipeline operating temperature; calculated by formula (14); ppm is the H2S content, unit ppm; , These are the starting and ending pressures of the pipeline operation, respectively.
[0037] 15) Operating medium characteristics -- CO2 / H2S partial pressure ratio The ratio of the partial pressures of CO2 to H2S is calculated using the following formula: Where R is the ratio of the partial pressures of CO2 to H2S.
[0038] 16) Operating medium characteristics - bacterial content Directly extract bacterial content information.
[0039] 17) Operating medium characteristics -- presence or absence of scale / structure factor The structure factors for four scale types—calcium carbonate, calcium sulfate, barium sulfate, and strontium sulfate—are calculated using the following formulas: ① Calcium carbonate scaling factors: in, , , , , , , , , , They are respectively , , , , , , , , , The content of is in mg / L; μ is the ionic strength in mol / L; AIK is the total alkalinity in mol / L; K is the correction factor, which is determined by both temperature and ionic strength μ and is dimensionless; SAI is the scaling factor and is dimensionless.
[0040] ② Calcium sulfate scaling factors: in, , They are respectively , Content, in mg / L; X is and The concentration difference, in mol / L; is the solubility product constant, dimensionless; S is the predicted value of calcium sulfate scaling trend, in mol / L; P is the calcium sulfate scaling factor.
[0041] ③ Barium sulfate scaling factor: in, , They are respectively , Content, in mg / L; m and a are the initial conditions in the water. , The concentration is expressed in mol / L. is the solubility product of BaSO4, which is temperature-dependent and dimensionless; B is the scaling factor of barium sulfate in the water after the water quality has stabilized, with units of mol / L. The value represents the amount of BaSO4 deposits, expressed in mg / L.
[0042] ④ Strontium sulfate scaling agent: in, , They are respectively , Content, in mg / L; The solubility product of strontium sulfate is taken as 2.8 × 10⁻⁶. -7 W represents the strontium sulfate scaling factor.
[0043] 18) Construction of Feature Extraction Module The feature extraction module described herein acquires internal detection data and operating medium data from the multi-source dataset of internal corrosion failures. It then extracts features such as the number of defects per kilometer (N_L), spatial distribution of defects, corrosion severity, morphology, scouring corrosion, waterline corrosion, weld corrosion, top corrosion, CO2 corrosion, CO2 / H2S corrosion, bacterial corrosion, and operating medium features (temperature, CO2 partial pressure, H2S partial pressure, CO2 / H2S partial pressure ratio, bacterial content, presence or absence of scale / structural factors). This provides input parameters for the failure cause identification model. This module is built using the Python language.
[0044] S5: Construction of a Model for Identifying the Causes of Corrosion Failure in Pipelines The construction of the pipeline internal corrosion failure cause identification model includes the construction of an internal corrosion failure cause identification model for purification pipelines and an internal corrosion failure cause identification model for gathering and transportation pipelines. The internal corrosion failure cause identification model for gathering and transportation pipelines further includes a mechanism-based failure cause identification model and a machine learning-based failure cause identification model, for a total of three internal corrosion failure cause identification models. All models were constructed using Python. The model selection method is as follows... Figure 3 As shown, the failure cause identification model based on machine learning (ML) requires training before it can be applied. The construction methods for the three models are explained below: 1) Purification pipeline failure cause identification model The aforementioned failure cause identification model for cleanroom pipelines can identify the causes of failure in cleanroom pipelines. There are four types of causes: weld corrosion, erosion corrosion, CO2 corrosion, and bacterial corrosion. The algorithm principle diagram is shown below. Figure 4 As shown, this algorithm is implemented using Python. The model will identify the cause of failure according to the following steps: Table 2. Pseudocode of the algorithm for identifying the causes of failure in cleanroom pipelines. Where Pipe_low represents the proportion of defects located at the bottom of the pipe in the 4:30~7:30 direction.
[0045] 2) Mechanism-based failure cause identification model for gathering and transportation pipelines The mechanism-based failure cause identification model for gathering and transportation pipelines can identify the causes of failures in these pipelines. There are nine types of causes: weld corrosion, erosion corrosion, top corrosion, waterline corrosion, oxygen corrosion, CO2 corrosion, CO2 / H2S corrosion, bacterial corrosion, and under-deposit corrosion. The algorithm's principle diagram is shown in Figure 5. Implemented in Python, the model will identify the failure causes according to the following steps: Table 3. Pseudocode of Mechanism-Based Failure Cause Identification Algorithm for Gathering and Transportation Pipelines 3) ML-based failure cause identification model for gathering and transportation pipelines The ML-based failure cause identification model for gathering and transportation pipelines can identify the failure causes of gathering and transportation pipelines. There are 9 types of causes: weld corrosion, erosion corrosion, top corrosion, waterline corrosion, oxygen corrosion, CO2 corrosion, CO2 / H2S corrosion, bacterial corrosion, and under-deposit corrosion. Figure 6 The training and testing process of the model is demonstrated, and the specific process is as follows: ① Use the feature extraction module to extract features from the data in the multi-source dataset of internal corrosion defects to form a multi-source feature set; ② Conduct a preliminary screening of causes to remove cause types with a small number of samples (weld corrosion, erosion corrosion, top corrosion, waterline corrosion, oxygen corrosion, and under-deposit corrosion). ③ Data augmentation was performed on CO2 corrosion, CO2 / H2S corrosion, and bacterial corrosion samples. The original dataset before augmentation consisted of 42 data points: 12 for CO2 corrosion, 24 for CO2 / H2S corrosion, and 6 for bacterial corrosion. Augmentation method: The pipeline was randomly divided into sections of 2km, 4km, 8km, 16km, and 32km for data augmentation. The augmented dataset consisted of 151 data points: 55 for CO2 corrosion, 47 for CO2 / H2S corrosion, and 49 for bacterial corrosion.
[0046] ④ XGBoost machine learning model construction: The XGBoost algorithm library in Python is used to build the model; ⑤ Data partitioning method: 5-fold cross-validation; ⑥ Based on the prediction results of the test set, optimize the model parameters to obtain the optimal model parameters; ⑦ Input new data in the data format of the multi-source dataset of internal corrosion defects, and sequentially go through the feature extraction module to extract features, conduct preliminary screening of causes, predict the model, and output the causes to achieve cause identification.
[0047] Figure 7 This demonstrates the overall workflow of applying an ML model, and shows how to implement this functionality using Python based on the defined workflow. The model will identify failure causes by following these steps: Table 4. Pseudocode of the ML-based algorithm for identifying failure causes in gathering and transportation pipelines. S6: Construction of Corrosion Control Strategies for Pipelines The pipeline corrosion control strategy is constructed by building corresponding control measures based on the failure cause type and its severity. When a corresponding indicator is triggered, an automatic recommendation of control measures is implemented. As shown in Table 5, the pipeline corrosion control strategy construction is implemented using Python. Wherein: The direction is from 4 to 8 o'clock. The ratio of the number of corrosion defects to the total number of defects in the 4-8 o'clock direction; The direction is from 4 to 8 o'clock. The ratio of the number of corrosion defects to the total number of defects in the 4-8 o'clock direction; The localized corrosion depth during oxygen corrosion, The localized corrosion depth during under-deposit corrosion. The depth of localized corrosion during erosion corrosion, This refers to the pipe type.
[0048] Table 5 Corrosion Control Strategies for Pipelines Software tool development and application (such as) Figure 8The software tool development and application described above integrates all algorithms and steps in steps S1-S6 to automate the entire process from data upload to data processing, feature extraction, model selection, cause identification, and prevention strategy recommendation, along with visualized operations. The software's functional modules include a homepage, data management module, cause identification module, expert verification module, control strategy module, model management module, and system management module, as shown in Figure 8. The functions of each module are as follows: Homepage: This module mainly showcases the data visualization of the system's main functions.
[0049] Data Management Module: This module mainly implements unit management, pipeline management, and data dictionary.
[0050] Cause identification module: This module mainly realizes data reporting, failure identification, and result viewing.
[0051] Data reporting module: This module mainly involves entering various parameters and displaying the results of data reporting.
[0052] Expert Verification Module: This module primarily manages verification and allows for result viewing. It enables manual verification of the cause identification results to ensure their accuracy and allows users to view the verified results.
[0053] Control Strategy Module: This module mainly enables the generation of countermeasures, viewing of results, and tracking of effects.
[0054] Model Management Module: This module primarily manages datasets, trains models, and applies models. The dataset management module provides statistics on all validation datasets and allows for the addition of confirmed correct data from external sources. Model training uses extracted feature datasets to train the XGBoost model. Model application loads relevant model training parameters into the model, enabling it to predict the causes of failures.
[0055] System Management Module: This module primarily implements fault target management, user management, menu management, role management, and log management. It mainly handles user permissions and the collection and display of related error messages.
[0056] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A data-driven intelligent identification method for the causes of corrosion failure in steel oil and gas pipelines, characterized by: The specific method includes the following steps: S1: Acquisition of internal detection data and operating condition media data; S2: Internal detection data processing, working condition media data alignment, and multi-source dataset construction: The internal detection data processing involves unifying and standardizing the collected internal detection data, including the following:
1. Filter out data with internal defects based on defect type; 2. Calculate the distance m to the nearest circumferential weld based on the information of the distance m from the upstream circumferential weld, the distance m from the downstream circumferential weld, and the length m of the pipe section; 3. For internal inspection data that lacks defect depth information, the defect depth information can be roughly estimated using the wall thickness reduction percentage and pipe wall thickness. Defect depth = pipe wall thickness × wall thickness reduction. S3: Internal Corrosion Cause Analysis and Cause Type Classification: The internal corrosion cause analysis and cause type classification refers to obtaining all failure cause types of internal corrosion by analyzing and classifying the internal corrosion causes of oil and gas pipelines. S4: Typical internal corrosion defect features, working condition medium feature extraction and feature extraction module construction: The typical internal corrosion defect features and working condition medium feature extraction include the feature of the number of defects per kilometer, the feature of defect spatial distribution, the feature of corrosion severity, the feature of defect morphology, the feature of scouring corrosion, the feature of waterline corrosion, the feature of weld corrosion, the feature of top corrosion, the feature of CO2 corrosion, the feature of CO2 / H2 S corrosion, the feature of bacterial corrosion, and the feature of working condition medium; S5: Construction of Pipeline Internal Corrosion Failure Cause Identification Model: The construction of the pipeline internal corrosion failure cause identification model includes the construction of a purification pipeline internal corrosion failure cause identification model and a gathering and transportation pipeline internal corrosion failure cause identification model; among them, the gathering and transportation pipeline internal corrosion failure cause identification model includes a mechanism-based failure cause identification model and a machine learning-based failure cause identification model, for a total of three internal corrosion failure cause identification models are constructed. All models are constructed using Python. The input of the pipeline internal corrosion failure cause identification model is pipeline internal detection data and pipeline operating condition medium data, and the output is the internal corrosion type; S6: Construction of Corrosion Control Strategies in Pipelines: The construction of corrosion control strategies in pipelines involves building corresponding control strategies based on the failure cause type and its severity; when the corresponding indicator is triggered, control strategies are automatically recommended.
2. The data-driven intelligent identification method for the causes of corrosion failure in steel oil and gas pipelines according to claim 1, characterized in that: The internal detection data and operating condition media data acquisition in S1 include the following data: The internal inspection data collection includes: mileage location (km), elevation location, clock position, distance from upstream circumferential weld (m), distance from downstream circumferential weld (m), pipe section length (m), defect length (mm), width (mm), depth (mm), and wall thickness reduction (%). The data collected for the operating medium includes: pipeline name, pipeline type, transported medium, pipeline length (km), pipeline wall thickness (mm), starting temperature (°C), ending temperature (°C), starting pressure (MPa), ending pressure (MPa), and daily liquid production (m³). 3 / day, daily gas production (m³) 3 / day, moisture content %, gas-liquid ratio %, flow rate m / s, CO2 content mol%, H2S content ppm, O2 content mol%, mineralization mg / L, K + Content (mg / L), Na + Content (mg / L), Ba 2+ Content (mg / L), Sr 2+ Content (mg / L), Ca 2+ Content (mg / L), Mg 2+ Content (mg / L), Fe 2+ Content (mg / L), Cl - Content (mg / L), CO3 2- Content (mg / L), HCO 3- Content (mg / L), SO4 2- Content (mg / L), pH value, bacterial count (cFU / mL), presence or absence of scale.
3. The data-driven intelligent identification method for the causes of corrosion failure in steel oil and gas pipelines according to claim 1, characterized in that; The four possible scenarios and corresponding calculation methods in step 2) of S2 are as follows: ① Only information on the distance m from the upstream circumferential weld and the distance m from the downstream circumferential weld is available; Calculation method: Take the minimum value of the two; ②Only information is available regarding the distance (m) from the upstream circumferential weld and the length (m) of the pipe section; Calculation method: Subtract the distance from the upstream circumferential weld from the pipe section length to obtain the distance from the downstream circumferential weld, and then take the minimum value between the distances from the upstream and downstream circumferential welds; ③ Only information on the distance m from the downstream circumferential weld and the length m of the pipe section is available; Calculation method: Subtract the distance to the downstream circumferential weld from the pipe section length to obtain the distance to the upstream circumferential weld, and then take the minimum value between the distances to the upstream and downstream circumferential welds; ④ Includes information on the distance m from the upstream circumferential weld, the distance m from the downstream circumferential weld, and the length m of the pipe section; Calculation method: The calculation method of case ① shall be adopted.
4. The data-driven intelligent identification method for the causes of corrosion failure in steel oil and gas pipelines according to claim 1, characterized in that: The extraction methods for each step in step S4 are as follows: 1) Defect Quantity Characteristics per Kilometer N_L: The defect quantity characteristic per kilometer is equal to the ratio of the total number of corrosion defects N in the pipeline to the total pipeline length L, that is: 2) Spatial distribution characteristics of defects: Among all defects, calculate the probability of a defect being located at the 1 to 12 o'clock positions sequentially, denoted as P1 to P12, using the following formula: 3) Corrosion severity characteristics: Among all defects, the proportions of defects with wall thickness reduction of less than 10%, 20%, 30%, 40%, and 50% are calculated sequentially and denoted as P_10, P_20, P_30, P_40, and P_50, respectively. The calculation formulas are as follows: 4) Defect morphology characteristics: Three morphologies are used to describe the morphology of the defects: equiaxed (0.5 <= length / width <= 2), elongated (length / width >= 10), and vertical (length / width <= 0.1); the proportion of each of the three morphologies is calculated using the following formulas: in, , , These represent the number of equiaxed defects, elongated defects, and vertical defects, respectively. 5) Erosion corrosion characteristics: Based on the three-dimensional morphological characteristics of erosion corrosion and the clock position, calculate the proportion of defects with a length / width > 10 near the 4-8 o'clock position. , Defined as a characteristic of erosion corrosion, the calculation formula is as follows: in, The number of defects with a length / width greater than 10 around points 4-8. This represents the total number of defects between 4 and 8 o'clock. 6) Characteristics of waterline corrosion: The three-dimensional morphology of waterline corrosion is characterized by a beaded pattern, with multiple defects continuously distributed along the interface. Based on this, the number of pipe segments with a continuous distribution length greater than 1m for defects at points 1-4 and 8-11 is calculated. The defect spacing is defined as less than 10cm and greater than 1m, and is considered a single segment. The calculation formula is shown below: in, express The number of pipe sections with a continuous distribution length of more than 1m due to clock defects. =1,2,3,4,8,9,10,11,12; 7) Weld corrosion characteristics: Calculate the proportion of defects located less than 10cm from the nearest circumferential weld. The calculation formula is as follows: in, This indicates the number of defects that are less than 10cm away from the nearest circumferential weld. 8) Top corrosion characteristics: Calculate the proportion of corrosion pits near point 12, using the formula shown below: 9) CO2 corrosion characteristics: Calculate the proportion of plateau-shaped defects at the 4-8 o'clock position. The calculation formula is as follows: in, The number of plateau-shaped defects between 4 and 8 o'clock. This represents the total number of defects between 4 and 8 o'clock. 10) CO2 / H2S corrosion characteristics: Calculate the proportion of pitting defects at 4-8 o'clock. The calculation formula is as follows: in, This indicates the number of pitting defects between the 4 and 8 o'clock positions. This represents the total number of defects between 4 and 8 o'clock. 11) Bacterial corrosion characteristics: Calculate the probability of having more than 10 equiaxed defects within a 1m interval in the direction from point 4 to point 8. The calculation formula is as follows: in, This indicates the number of pipe sections with more than 10 equiaxed defects within a 1m interval in the direction from 4 to 8 o'clock. The total length of the pipeline is expressed in km. 12) Operating medium characteristics – temperature: Calculate the average of the starting and ending temperatures, using the formula shown below: 13) Operating medium characteristics -- CO2 partial pressure: Convert the CO2 content (in percentage) in the original data to MPa. The calculation formula is as follows: Where A represents the CO2 volume percentage, in % %. This is the partial pressure of CO2, expressed in MPa. , These are the starting and ending pressures of the pipeline operation, respectively. 14) Operating medium characteristics -- H2S partial pressure: Convert the H2S content in the original data (in ppm) to MPa. The calculation formula is as follows: in, for Content, unit MPa; T is the pipeline operating temperature; calculated by formula (14); ppm is the H2S content, unit ppm; , These are the starting and ending pressures of the pipeline operation, respectively. 15) Operating medium characteristics -- CO2 / H2S partial pressure ratio: Calculate the ratio of the partial pressures of CO2 and H2S. The calculation formula is shown below: Where R is the ratio of the partial pressures of CO2 to H2S; 16) Operating medium characteristics - bacterial content: directly extract bacterial content information; 17) Operating medium characteristics – presence or absence of scale / structure factor: Calculate the structure factor for four scale types: calcium carbonate, calcium sulfate, barium sulfate, and strontium sulfate. The calculation formula is shown below: ① Calcium carbonate scaling factors: in, , , , , , , , , , They are respectively , , , , , , , , , The content of is expressed in mg / L; μ is the ionic strength in mol / L; AIK is the total alkalinity in mol / L; K is the correction factor, which is determined by both temperature and ionic strength μ and is dimensionless; SAI is the scaling factor and is dimensionless. ② Calcium sulfate scaling factors: in, , They are respectively , Content, in mg / L; X is and The concentration difference, in mol / L; is the solubility product constant, dimensionless; S is the predicted value of calcium sulfate scaling tendency, in mol / L; P is the calcium sulfate scaling factor. ③ Barium sulfate scaling factor: in, , They are respectively , Content, in mg / L; m and a are the initial conditions in the water. , The concentration is expressed in mol / L. is the solubility product of BaSO4, which is temperature-dependent and dimensionless; B is the scaling factor of barium sulfate in the water after the water quality has stabilized, with units of mol / L. The amount of BaSO4 scale, in mg / L; ④ Strontium sulfate scaling factor: in, , They are respectively , Content, in mg / L; The solubility product of strontium sulfate is taken as 2.8 × 10⁻⁶. -7 W represents the strontium sulfate scaling factor.
5. The data-driven intelligent identification method for the causes of corrosion failure in steel oil and gas pipelines according to claim 1, characterized in that: The three models in S5 include a purification pipeline failure cause identification model, a mechanism-based gathering and transportation pipeline failure cause identification model, and a ML-based gathering and transportation pipeline failure cause identification model.
6. A data-driven intelligent identification system for the causes of corrosion failure in steel oil and gas pipelines using the method described in any one of claims 1-5, characterized in that: The software's functional modules include a homepage, data management module, cause identification module, expert verification module, control strategy module, model management module, and system management module. The functions of each module are as follows: Homepage: This module showcases the data visualization of the system's functions; Data Management Module: This module implements unit management, pipeline management, and data dictionary; Cause identification module: This module enables data reporting, failure identification, and result viewing; Data reporting module: This module records various parameters and displays the results of data reporting. Expert Verification Module: This module manages verification and allows users to view results. It also allows for manual verification of the causal identification results to ensure their accuracy. Users can view the results that have already been verified. Control strategy module: This module enables the generation of countermeasures, viewing of results, and tracking of effects; Model Management Module: This module manages datasets, trains models, and applies models. The dataset management module provides statistics on all validation datasets and allows the addition of confirmed correct data from external sources. Model training uses the extracted feature datasets to train the XGBoost model. Model application loads the relevant model training parameters into the model, enabling it to predict the causes of failures. System Management Module: This module implements failure target management, user management, menu management, role management, and log management; it also implements the collection and display of user permissions and related error messages for the system.
Citation Information
Patent Citations
Analyzing method of failure cause of oil and gas field pipeline
CN108318539A
Visual analysis method and device for multiple internal corrosion failure types of pipeline
CN119625363A