Fault diagnosis method suitable for large oilfield water injection system

By constructing standardized data sets and weighted topological diagrams, combined with collaborative analysis of multiple fault trees, accurate fault diagnosis and energy efficiency optimization of oilfield water injection systems are achieved, solving the problems of untimely fault diagnosis and insufficient energy efficiency optimization in existing technologies, and improving diagnostic coverage and analysis efficiency.

CN120804929APending Publication Date: 2025-10-17NORTHEAST GASOLINEEUM UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510886647.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing fault diagnosis technology for oilfield water injection systems relies on a single parameter threshold alarm, ignoring the temporal correlation between flow and energy consumption and the constraints of the pipe network topology. It is difficult to handle fuzzy states, separates fault handling from energy efficiency optimization, lacks a multi-tree collaboration mechanism, and has not established a parameter self-update mechanism, resulting in untimely fault diagnosis and insufficient energy efficiency optimization.

Method used

By collecting injection well flow, pressure, and energy consumption data, a standardized data set with topological identification is generated, a weighted topological map is constructed, and the status score is calculated based on the topological characteristics. Abnormal patterns are identified, multiple fault tree types are activated, fuzzy node evaluation is performed, a fault hypothesis set is generated, and three-level verification is performed. The optimal maintenance strategy is selected and the pipe network topology structure weight is updated.

Benefits of technology

It achieves accurate diagnosis and energy efficiency optimization of complex faults, improves diagnostic coverage and analysis efficiency, reduces uncertainty misjudgment rate, solves the false alarm and diagnostic blind spot problems of traditional methods, and forms a closed-loop optimization system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120804929A_ABST
    Figure CN120804929A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of oil field intelligent operation and maintenance, and particularly discloses a fault diagnosis method suitable for a large-scale oil field water injection system. According to the method, flow, pressure and energy consumption data of a water injection well are collected and subjected to standardization processing, and a pipe network topological structure with feature annotations is constructed; after an abnormal mode is identified based on topological features and operation data, a fault hypothesis set with confidence is generated by adopting a multi-fault tree cooperative activation and hybrid evaluation mechanism; and obtaining a definite diagnosis conclusion through a three-stage verification process, and finally realizing energy efficiency influence evaluation and optimal maintenance strategy selection. According to the method, the defects of a traditional method in the aspects of composite fault diagnosis and energy efficiency optimization are overcome, the diagnosis accuracy is effectively improved, the system false alarm rate is reduced, the maintenance cost is reduced, and an integrated solution of intelligent fault diagnosis and optimization decision is provided for the oilfield water injection system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent operation and maintenance of oil fields, and in particular to a method suitable for diagnosing faults of water injection systems in large oil fields. Background Art

[0002] The oilfield water injection system is a complex water power system with a wide coverage area and a large scale. After decades of development and operation, some operating modules of the water injection system can no longer adapt to the production needs of the new era. Some aging phenomena of the system are serious, and part of the system layout urgently needs to be transformed and optimized. In order to ensure the normal production of the oilfield, engineering and technical personnel often need to conduct fault inspections on the water injection system and need to repair faults in a timely manner when faults occur. However, relying solely on the experience of engineering and technical personnel to inspect and troubleshoot faults often fails to discover and eliminate faults in a timely manner. The current fault diagnosis technology of oilfield water injection systems has three technical bottlenecks:

[0003] 1. The existing system relies on a single parameter threshold alarm (such as pressure exceeding the limit), ignoring the temporal correlation between flow and energy consumption and the topological constraints of the pipeline network.

[0004] 2. It only covers specific fault scenarios (such as pump unit failure), lacks a multi-tree coordination mechanism, and is difficult to handle fuzzy states such as "high pressure".

[0005] 3. Fault handling and energy efficiency optimization are performed separately. The maintenance plan ignores topological conduction effects (such as valve adjustment causing pipeline pressure oscillation). The diagnostic model is statically solidified, and a parameter self-update mechanism is not established. Summary of the Invention

[0006] The present invention provides a method for fault diagnosis of large oilfield water injection systems to solve the problem of how to achieve accurate diagnosis of complex faults and energy efficiency optimization decision-making based on the real-time operating parameters and topological structure characteristics of the oilfield water injection system through the collaborative activation of multiple fault trees and a hybrid evaluation mechanism.

[0007] In order to solve the above technical problems, the present invention provides a method for diagnosing faults in a large oilfield water injection system, comprising:

[0008] Collect injection well flow, pressure, and energy consumption data, clean and standardize them, and generate standardized data sets with topological identifiers;

[0009] Construct the pipe network topology based on the standardized data set, calculate the edge weights and node centrality, and generate a weighted topology graph with feature annotations;

[0010] Based on the data of the standardized data set and weighted topological map, the health score is calculated by combining topological features, identifying underfilling, overfilling and leakage abnormal patterns, and generating a subsystem health status report;

[0011] Based on the abnormal pattern in the subsystem health state report, the corresponding fault tree type is activated, the fuzzy node condition evaluation and potential fault cause combination derivation are performed, and a fault hypothesis set with confidence is generated;

[0012] The expression of the fault hypothesis set with confidence is:

[0013]

[0014] Where CF(H) represents the comprehensive confidence of hypothesis H; w k represents the kth evidence weight; s k represents the kth evidence support degree; m is the total number of possible hypotheses; n is the number of evidence of the current hypothesis;

[0015] is a continuous multiplication operator; w j is the weight coefficient of the jth evidence; s j is the support degree of the jth evidence;

[0016] The fault hypothesis set is verified in three levels, fast screening, fine verification and impact assessment, and a final diagnosis of the fault diagnosis conclusion is generated;

[0017] According to the fault diagnosis conclusion, the influence of the fault on energy efficiency is evaluated, the effects of different maintenance schemes are simulated, the optimal maintenance strategy is selected, and the weight parameters of the pipe network topology structure are updated.

[0018] Further, the step of generating a standardized data set with a topology identifier comprises:

[0019] Obtain the injection well flow data, pipe network pressure data and injection station energy consumption data, perform missing value filling and outlier removal, and obtain the cleaned original data;

[0020] Extract the time series features from the cleaned original data, perform sampling rate alignment and unit standardization, and obtain the time alignment data;

[0021] The time alignment data is attached with a spatial position label to generate a standardized data set with a topology identifier.

[0022] Further, the step of generating a weighted topology graph with feature annotation comprises:

[0023] Based on the spatial position label in the standardized data set, the pipe network topology structure of nodes and edges is constructed;

[0024] Obtain the flow and pressure parameters from the standardized data set, calculate the pipe network edge weight and node centrality;

[0025] According to the pipe network topology structure, the pipe network edge weight and the node centrality, a high sensitivity path is identified, and a weighted topology graph with feature annotation is generated.

[0026] Further, the step of generating the subsystem health status report comprises:

[0027] Selecting pressure and flow indicators from the standardized dataset, performing fuzzy membership conversion to obtain fuzzy language variables;

[0028] Combining path characteristics in the weighted topology graph and fuzzy language variables to calculate the operating state score of the injection well and the pipe network;

[0029] According to the operating state score, identifying under-injection, over-injection, and leakage anomaly patterns, and generating a subsystem health status report.

[0030] Further, the step of activating the corresponding fault tree type based on the anomaly pattern in the subsystem health status report comprises:

[0031] Based on the anomaly pattern in the subsystem health status report, activating the corresponding fault tree type;

[0032] The fault tree type includes an emergency accident fault tree, an injection well injection condition fault tree, an injection station energy consumption fault tree, an injection pipe network energy consumption fault tree, and an injection well energy consumption fault tree.

[0033] Further, the step of performing fuzzy node condition evaluation and potential fault cause combination derivation comprises:

[0034] Combining the fault tree type and the parameters in the standardized dataset to perform fuzzy evaluation of fault tree node conditions;

[0035] Fuzzy evaluation includes converting quantitative parameters into fuzzy language variables and performing condition judgment based on membership functions.

[0036] Further, the step of generating a fault hypothesis set with confidence comprises:

[0037] According to the fuzzy evaluation results, deriving potential fault cause combinations, and generating a fault hypothesis set with confidence.

[0038] Further, the step of generating a confirmed fault diagnosis conclusion comprises:

[0039] Performing first-layer rapid screening on the fault hypothesis set to retain high-probability fault hypotheses;

[0040] Combining real-time data in the standardized dataset to perform second-layer fine verification;

[0041] Based on the propagation path in the weighted topology graph, evaluating the third-layer influence range, and generating a confirmed fault diagnosis conclusion.

[0042] Further, the step of selecting the optimal maintenance strategy and updating the weight parameters of the pipe network topology structure comprises:

[0043] According to the fault diagnosis conclusion, the influence degree on the energy efficiency of the water injection system is calculated.

[0044] Based on the weighted topology graph, the energy efficiency improvement effect of different maintenance schemes is simulated.

[0045] The optimal maintenance strategy is selected, and the weight parameters of the pipe network topology structure are updated.

[0046] Further, the step of selecting the optimal maintenance strategy and updating the weight parameters of the pipe network topology structure further comprises:

[0047] The updated weight parameters of the pipe network topology structure are fed back to the pipe network topology structure construction process to form a closed-loop optimization system.

[0048] The key innovations of the present application include:

[0049] (1) The PageRank algorithm is improved and applied to the collaborative analysis of multiple fault trees, solving the blind area of traditional methods for compound faults and improving the diagnostic completeness.

[0050] (2) The Navier-Stokes equation is introduced into the fault modeling of the water injection system, which can accurately quantify the spatiotemporal propagation law of the fault in the pipe network.

[0051] (3) Zadeh fuzzy set and Dempster-Shafer evidence theory are fused to effectively handle the uncertainty of monitoring data and the conflict of expert knowledge.

[0052] The main beneficial effects are as follows:

[0053] (1) Through standardized data collection and cleaning, the multi-source heterogeneous data such as water injection well flow, pressure, energy consumption, etc. are spatiotemporally aligned and topologically identified, solving the false alarm problem caused by traditional single parameter monitoring.

[0054] (2) Based on the improved PageRank algorithm, the dynamic fault tree activation mechanism realizes intelligent switching and collaborative analysis of five types of fault trees, effectively improving the diagnostic coverage rate of complex faults and improving the analysis efficiency compared with the traditional single fault tree method.

[0055] (3) The fault propagation model accurately simulates the fault diffusion path of the fluid system, combines fuzzy integration and Bayesian-DS evidence theory, effectively reduces the misdiagnosis rate of uncertain diagnosis, improves the comprehensive confidence of the fault hypothesis set, and is significantly better than the traditional method. BRIEF DESCRIPTION OF DRAWINGS

[0056] Figure 1A flowchart of a method suitable for large oilfield water injection system fault diagnosis provided by the embodiment of the application is shown in the figure. DETAILED DESCRIPTION

[0057] Embodiment one: reference Figure 1 A flowchart of a method suitable for large oilfield water injection system fault diagnosis provided by the embodiment of the application is shown in the figure, which can at least include steps S100-S600:

[0058] S100, collect water injection well flow, pressure and energy consumption data, perform cleaning and standardization processing, and generate a standardized data set with topology identification.

[0059] S200, construct a pipe network topology structure based on the standardized data set, calculate edge weights and node centrality, and generate a weighted topology graph with feature annotation.

[0060] S300, calculate state scores based on the data of the standardized data set and the weighted topology graph, identify under-injection, over-injection and leakage abnormal patterns, and generate a subsystem health state report.

[0061] S400, based on the abnormal patterns in the subsystem health state report, activate the corresponding fault tree type, perform fuzzy node condition evaluation and potential fault reason combination derivation, and generate a fault hypothesis set with confidence.

[0062] S500, three-level verification is performed on the fault hypothesis set, rapid screening, fine verification and impact assessment are performed, and a confirmed fault diagnosis conclusion is generated.

[0063] S600, according to the fault diagnosis conclusion, evaluate the impact of the fault on energy efficiency, simulate the effect of different maintenance schemes, select the optimal maintenance strategy and update the weight parameters of the pipe network topology structure.

[0064] Step S100 at least includes steps S110-S130:

[0065] S110, obtain water injection well flow data, pipe network pressure data and water injection station energy consumption data, perform missing value filling and abnormal value elimination, and obtain cleaned original data;

[0066] The instantaneous flow data is collected by the electromagnetic flowmeter deployed at the water injection wellhead, and the sampling frequency is set to 1 time per minute; the pressure data is obtained by installing the pressure transmitter at the key nodes of the pipe network, and the sampling frequency is 1 time per 5 minutes; the energy consumption data is collected by installing the intelligent electric meter at the water injection station power distribution cabinet, and the sampling frequency is 1 time per hour. The data is transmitted to the central server through the industrial Internet of Things gateway to form an initial data set.

[0067] Specifically, the implementation process of the missing value filling and outlier removal includes: for the missing value, a spatiotemporal correlation filling algorithm is adopted, first, the data of the adjacent time points of the same monitoring point is searched, if there is no valid data, the data of the spatial adjacent monitoring points at the same time is searched; for the outliers, a dynamic threshold detection mechanism based on the 3σ criterion is established, and double verification is performed in combination with the rated parameter range of the equipment. The cleaning process further includes eliminating signal jump interference, and a sliding window mean filtering algorithm is adopted, and the window size is set to be different from 5 to 15 sampling points according to the parameter characteristics.

[0068] Understandably, the implementation process of obtaining the cleaned original data is: aligning the various types of data processed by the above processing according to the time stamp to form a structured data table containing the time stamp, the monitoring point ID, the flow value, the pressure value and the energy consumption value. Each record in the data table corresponds to a complete parameter set of a monitoring point at a sampling time, which serves as the basis for subsequent processing.

[0069] S120, extracting time sequence features from the cleaned original data, performing sampling rate alignment and unit standardization to obtain time-aligned data;

[0070] An independent time sequence is constructed for each parameter of each monitoring point, and statistical features including but not limited to: moving average trend, differential fluctuation amplitude and periodic component are calculated. Specifically, the moving average trend is calculated by using a 24-hour sliding window, which reflects the daily variation law of the parameter; the differential fluctuation amplitude calculates the absolute value of the change rate of adjacent sampling points, which reflects the stability of the parameter; and the periodic component extracts the main frequency component by using fast Fourier transform.

[0071] Further, the implementation process of the sampling rate alignment and unit standardization includes: uniformly resampling the data of different sampling frequencies to a common frequency of 1 time per hour, for high-frequency data, using arithmetic average aggregation, and for low-frequency data, using linear interpolation expansion; and the unit standardization uniformly converts the flow to m 3 / h, the pressure to MPa, and the energy consumption to kW·h, so as to eliminate the dimension effect. Specifically, the standardization further includes parameter normalization, which maps each parameter value to the interval [0, 1], and adopts the maximum and minimum value normalization method.

[0072] The implementation of obtaining the time-aligned data is: generating a structured data set containing the standardized time stamp, the monitoring point ID, the normalized flow value, the normalized pressure value, the normalized energy consumption value and various time sequence features. The data set is arranged in chronological order, and the data of each monitoring point forms a continuous time sequence, which facilitates subsequent spatiotemporal analysis.

[0073] S130, adding a spatial position label to the time-aligned data to generate a standardized data set with a topological identifier;

[0074] Based on the injection system layout map provided by the oilfield geographic information system, a unique spatial coordinate and a topological identifier are assigned to each monitoring point. Specifically, the topological identifier includes: the subsystem number (injection well, main pipe network, branch pipe network, injection station, etc.), upstream and downstream connection relationship, and hierarchical position in the pipe network. Understandably, the spatial coordinate uses the oilfield self-defined coordinate system, containing plane position and elevation information.

[0075] Further, the implementation process of generating the standardized dataset with topological identifier is: associating the time-aligned data with spatial topological information to form a comprehensive dataset containing spatio-temporal complete information. The structure of the dataset includes: basic parameter part (timestamp, spatial coordinate, monitoring value), time characteristic part (various statistical characteristics), and spatial topological part (connection relationship, hierarchical information). Specifically, the dataset uses column storage format, which facilitates quick retrieval by time or space dimension.

[0076] The standardized dataset, as the basis for all subsequent analysis steps, its key characteristics are reflected in: consistent sampling frequency and standardized value range in time dimension, clear topological connection relationship in space dimension, and reliable numerical value after cleaning and extracted feature index in parameter dimension. The standardized characteristics of the dataset enable subsequent pipe network modeling, state evaluation and fault diagnosis to be carried out in a unified spatio-temporal framework.

[0077] Step S200 includes at least steps S210-S230:

[0078] S210, based on the spatial position label in the standardized dataset, constructing a pipe network topological structure of nodes and edges;

[0079] First, analyze the spatial position label information contained in the standardized dataset, which includes the geographic coordinates of the monitoring point, the type of the subsystem it belongs to, and the upstream and downstream connection relationship. Specifically, each injection well, injection station and pipe network branch point is abstracted as a topological node, and the pipe connecting these nodes is abstracted as a topological edge. Understandably, the mapping relationship between the nodes and edges is realized by analyzing the connection relationship field in the spatial position label, ensuring one-to-one correspondence between the physical system and the topological structure.

[0080] Further, the construction process of the topological structure includes node attribute labeling and edge attribute initialization. The node attribute labeling includes: assigning a unique identifier to each node, recording the node type (injection well, injection station, pipe network node, etc.), and marking the hierarchical position of the node in the system. The edge attribute initialization includes: assigning a unique identifier to each edge, recording the start node and end node connected by the edge, and initializing the edge length attribute. Specifically, the edge length is obtained by calculating the Euclidean distance of the spatial coordinates of the connected nodes, serving as the basic parameter for topological analysis.

[0081] The storage form of the pipe network topology adopts an adjacency list data structure, each node maintains a list of directly connected edges, and each edge records the references of the two nodes connected. The data structure supports efficient topology traversal and relationship query, providing a basis for subsequent weight calculation and path analysis. Understandably, the topology is constructed in memory as a directed graph model, with the direction consistent with the actual water flow direction, reflecting the physical flow characteristics of the water injection system.

[0082] S220, obtain flow and pressure parameters from the standardized data set, calculate pipe network edge weight and node centrality;

[0083] For each topological edge, its corresponding sensor data on the physical pipe segment is associated. Specifically, the flow and pressure monitoring values of the pipe segment in the last 24 hours are extracted from the standardized data set, and the time-weighted average value is calculated as the basic input of the edge weight. The calculation of the edge weight considers three dimensions: the flow dimension weight reflects the pipe segment's conveying capacity, the pressure dimension weight reflects the pipe segment's resistance characteristics, and the length dimension weight reflects the pipe segment's spatial scale.

[0084] Further, the specific calculation process of the edge weight is as follows: first, normalize the flow value to map it to the 0-1 interval to represent the relative conveying capacity; then, perform inverse proportional conversion on the pressure drop value to represent the degree of pipe patency; finally, calculate the comprehensive weight coefficient in combination with the pipe segment length. The weight coefficient adopts the geometric mean method to integrate the three dimensions, ensuring balanced contribution of each dimension. Specifically, the weight calculation also considers static attributes such as pipe segment material and service life, and adjusts the maximum weight value through a correction coefficient.

[0085] The implementation process of calculating node centrality includes the calculation of three indicators: degree centrality, closeness centrality, and betweenness centrality. The degree centrality counts the number of directly connected edges of each node, reflecting the local importance of the node; the closeness centrality calculates the sum of the reciprocals of the shortest path lengths from the node to all other nodes in the network, reflecting the global accessibility of the node; the betweenness centrality counts the frequency of the node appearing in the shortest paths of other node pairs, reflecting the pivotal role of the node. Specifically, the centrality indicators adopt the Z-score standardization method to eliminate dimensional differences, and the final node centrality is the weighted sum of the three indicators.

[0086] The storage form of the weight and centrality calculation results is that the edge weight is attached to the topological edge object as an attribute, and the node centrality is attached to the topological node object as an attribute. The attribute update mechanism supports dynamic recalculation, and when the monitoring data in the standardized data set is updated, the recalculation of the weight and centrality is automatically triggered. Understandably, the dynamic update feature enables the topology model to reflect real-time state changes of the system, providing the latest structural information for fault diagnosis.

[0087] S230, identifying high sensitivity paths according to the pipe network topology, the pipe network edge weight and the node center degree, and generating a weighted topology graph with feature annotations;

[0088] First, the path sensitivity evaluation index is defined based on the edge weight and the node center degree. Specifically, the path sensitivity is determined by three factors: the coefficient of variation of the edge weight on the path reflects stability, the average center degree of the path node reflects importance, and the path length reflects the influence range. The sensitivity evaluation index determines the weight of each factor by using the analytic hierarchy process, and calculates the comprehensive sensitivity score of the path.

[0089] Further, the high sensitivity path identification algorithm includes the following steps: screening candidate paths with lengths exceeding a threshold from all possible paths; calculating the comprehensive sensitivity score of each candidate path; and selecting the top 10% as high sensitivity paths according to the score. Specifically, the path search uses an improved Dijkstra algorithm, considering the dual constraints of edge weight and node center degree. The algorithm is implemented with a maximum search depth limit to ensure that the calculation efficiency is adapted to the scale of large pipe networks.

[0090] The implementation of generating a weighted topology graph with feature annotations is to attach various types of feature annotations calculated on the basis of the topology structure. Specifically, the feature annotations include: edge weight annotations to show the relative importance of pipe sections, node center degree annotations to show the degree of node hub, and high sensitivity path annotations to show the weak links of the system. The annotation form uses hierarchical color coding for easy visualization and analysis. Understandably, the weighted topology graph is stored in a graph database, supporting complex topology queries and real-time update operations.

[0091] The output format of the weighted topology graph contains two versions: the complete version retains all calculation details for internal use by the system, and the simplified version only retains key features for user interface display. Specifically, the complete version contains the original values of all edge weights and node center degrees, as well as the detailed composition of high sensitivity paths; the simplified version only marks high sensitivity paths and key nodes, highlighting the parts of the system that need to be focused on. The dual version design takes into account the needs of analysis depth and display effect.

[0092] Step S300 includes at least steps S310-S330:

[0093] S310, selecting pressure and flow indicators from the standardized data set, performing fuzzy membership conversion to obtain fuzzy language variables;

[0094] Firstly, the injection well flow rate and pipeline pressure monitoring data of the last 24 hours are extracted from the standardized data set, which has been cleaned and standardized by S100, and has a unified time resolution and dimension. Specifically, the flow rate and pressure data of each monitoring point are constructed into time series, and their statistical characteristics including mean, standard deviation, range and trend slope are calculated. Understandably, the statistical characteristics are used as input parameters for fuzzy processing, reflecting the operating state characteristics of the monitoring point.

[0095] Further, the implementation process of the fuzzy membership conversion includes defining five fuzzy language variables for each parameter: "very low", "low", "normal", "high", and "very high", and designing a trapezoidal membership function for each language variable. Specifically, the boundary values of the membership function are determined according to the operating specifications of the oilfield water injection system and the historical data distribution, and different threshold standards are used for monitoring points at different locations. The fuzzy processing converts the flow rate and pressure parameter values of each monitoring point into the membership degree of the corresponding language variable, forming a fuzzy language description.

[0096] Further, for each parameter of each monitoring point, a five-dimensional vector is generated to represent its membership degree to each language variable. The vector is normalized to ensure that the sum of the dimensions is 1. Specifically, the fuzzy language variables are stored in a specially designed fuzzy fact library, associated with the spatial location information and time stamp of the monitoring point, providing standardized input for subsequent state scoring. Understandably, the fuzzy processing effectively solves the uncertainty judgment problem caused by parameter fluctuations.

[0097] S320, combining the path characteristics in the weighted topology graph with the fuzzy language variables, calculate the operating state score of the injection well and the pipeline network;

[0098] Firstly, three key topological features are extracted from the weighted topology graph generated in S200: node centrality reflects the importance of location, edge weight reflects the operating state of the pipe section, and high sensitivity path marker reflects the vulnerable link of the system. Specifically, the topological features are subjected to multidimensional correlation analysis with the fuzzy language variables obtained in S310 to establish a comprehensive evaluation model.

[0099] Further, the calculation process of the operating state score uses the analytic hierarchy process to build: the first layer weight is determined by the node centrality, reflecting the importance difference of different positions in the system; the second layer weight is determined by the edge weight, reflecting the reliability of the pipe section state; the third layer score is determined by the membership degree of the fuzzy language variable, reflecting the deviation degree of the actual operating parameter from the ideal state. Specifically, the scoring model uses a differentiated evaluation strategy for injection wells and pipeline networks, focusing on flow rate stability evaluation for injection wells and pressure balance evaluation for pipeline networks.

[0100] For each evaluation unit (single well or pipe section), first determine its topological feature weight, then calculate the basic score combined with fuzzy language variables, and finally perform weighted smoothing through the scores of spatially adjacent units. The score results are normalized to the 0-100 interval, with higher scores indicating better status. Specifically, the scoring process uses a sliding time window mechanism to comprehensively consider recent status change trends, avoiding misjudgments caused by transient fluctuations. Understandably, the score results reflect not only the current status but also historical change information.

[0101] The output form of the running status score is that each evaluation unit obtains a comprehensive score and multiple dimension sub-item scores (flow stability, pressure balance, etc.). The score results are associated with spatial location information in the weighted topological graph to form a status score map with spatial distribution characteristics. Specifically, the score map is visualized in the form of a heat map, which facilitates intuitive identification of weak links in the system. The score data is stored in the status assessment database, supporting time series analysis and comparison.

[0102] S330, identify under-injection, over-injection and leakage anomaly patterns according to the running status score, and generate a subsystem health status report;

[0103] First, define the judgment rules for the three types of anomaly patterns: under-injection mode is characterized by low flow score and high pressure score; over-injection mode is characterized by high flow score and low pressure score; leakage mode is characterized by simultaneous rapid decline in flow and pressure scores. Specifically, the judgment rules are formulated in combination with expert experience, with different threshold standards set for nodes at different locations.

[0104] Further, the anomaly pattern recognition algorithm includes the following steps: scan the status scores of all evaluation units to mark potential abnormal points; check the score patterns of adjacent units to confirm the abnormal range; analyze the abnormal propagation path in combination with the topological connection relationship. Specifically, the identification process uses a multi-scale analysis method to focus on both single-point anomalies and regional anomaly patterns. A state transition probability matrix is set in the algorithm implementation to improve the continuity of anomaly judgment.

[0105] The implementation of generating a subsystem health status report is to integrate all anomaly recognition results and classify and summarize them by subsystem (water injection well group, pipe network section, water injection station, etc.). Specifically, the report content includes: comprehensive score of each subsystem, anomaly pattern distribution, detailed analysis of key abnormal points, and anomaly development trend prediction. Understandably, the report is stored in a structured format, including both text description and visual charts. The report version management function supports historical version tracing and comparative analysis.

[0106] The output of the health status report is divided into three levels: the summary layer provides an overview of the overall health of the system; the subsystem layer shows the detailed status of each unit; the exception point layer focuses on problem positioning and analysis. Specifically, the report triggers the fault tree type activation of step S400 after being generated, realizing seamless connection of evaluation and diagnosis. The report also supports custom filtering and sorting functions to facilitate specialized analysis for different concerns.

[0107] Step S400 includes at least steps S410-S430:

[0108] S410, based on the abnormal mode in the subsystem health status report, activate the corresponding fault tree type;

[0109] A mapping relationship matrix M is established between the abnormal mode and the fault tree type, and the matrix element M i,j is calculated by the following formula:

[0110] M i,j =σ(α·S i +β·C j +γ·R i,j )

[0111] Where: σ is the Sigmoid activation function; S i represents the severity score of the i-th abnormal mode; C j represents the confidence weight of the j-th fault tree; R i,j represents the historical association frequency of abnormal mode i and fault tree j; α, β, γ are trainable parameters; i is the abnormal mode index; j is the fault tree type index. Example: when the under- injection (i = 1) and the injection well fault tree (j = 2) historical association frequency R 1,2 = 0.8, α = 0.7, the matching weight is enhanced.

[0112] Specifically, the fault tree type includes five types: emergency accident fault tree T emerg , injection well injection fault tree T well , injection station energy consumption fault tree T station , injection pipe network energy consumption fault tree T pipe and injection well energy consumption fault tree T effic . The activation weight is calculated by an improved version of the PageRank algorithm:

[0113]

[0114] Where: PR(T k ) represents the PageRank value of fault tree T k ; d is the damping factor (0.85); N is the total number of fault trees; In(T k ) represents the pointer to T ka set of fault trees of T j ) represents the matching score of the current abnormal pattern and T j a set of fault trees pointed by T i,k is the matching score of the current abnormal pattern and T k is the cardinality of the set.

[0115] The fault trees whose final activation weights exceed the threshold τ enter the next stage of analysis. This algorithm solves the problem of cooperative activation of multiple fault trees and significantly improves the diagnostic accuracy.

[0116] S420, combined with the fault tree type and the parameters in the standardized data set, the fault tree node condition fuzzy evaluation is carried out;

[0117] The implementation of the combination of the fault tree type and the parameters in the standardized data set for the fault tree node condition fuzzy evaluation is to construct an evaluation model based on the Navier-Stokes equation to describe the fault propagation:

[0118]

[0119] Where: u=(u1,u2,u3) represents the fault propagation velocity field; p represents the node pressure (from the standardized data set); v represents the fault diffusion coefficient (historical data statistics); f represents the external disturbance term; u1 is the axial propagation velocity along the pipeline; u2 is the radial diffusion velocity; u3 is the vortex disturbance velocity; is the Nabla operator, is the Laplace operator.

[0120] Numerical solution is carried out by Monte Carlo method:

[0121] X n+1 =X n +a(X n ,t n )Δt+b(X n ,t n )ΔW n

[0122] Where: X n represents the node state vector; a(X,t) is the drift term; b(X,t) is the diffusion term; ΔW n represents the Wiener process increment.

[0123] Fuzzy evaluation uses Zadeh extension principle:

[0124] μ A (x)=sup{α∈[0,1]∣x∈A α}

[0125] Where: μ A(x) represents the membership degree of node condition x to fuzzy set A; A α represents the α-cut set of the fuzzy set A; sup is the supremum operation.

[0126] The fuzzy truth value of the node condition is calculated by the measure formula:

[0127]

[0128] Where: g(x) represents the weight function (normalized data set flow parameter); represents the Sugeno fuzzy integral operator.

[0129] This evaluation model can effectively handle the uncertainty and fuzziness in node conditions and provide an accurate quantitative basis for fault diagnosis.

[0130] S430, deriving potential fault cause combinations based on the fuzzy evaluation results, and generating a fault hypothesis set with confidence levels;

[0131] The implementation method of generating a confidence level fault hypothesis set by deriving potential fault cause combinations based on fuzzy evaluation results is as follows: constructing a hypothesis space search algorithm using the Hamiltonian principle:

[0132]

[0133] Where: p=(p1,…,p n ) represents the generalized momentum (fault propagation intensity); q=(q1,…,q n )surface

[0134] represents the generalized coordinates (fault location parameters); T(p) represents the kinetic energy term (fault diffusion tendency); V(q) represents the potential energy term (system resistance).

[0135] Solve the optimal fault path by Hamilton's equation:

[0136]

[0137] The confidence calculation adopts the Bayesian-DS evidence fusion framework:

[0138]

[0139] Where: m(A) represents the confidence of hypothesis A; m1, m2 represent the confidence distribution of different evidence sources;

[0140] Represents the conflict factor.

[0141] Final comprehensive confidence calculation:

[0142]

[0143] wherein: CF(H) represents the comprehensive confidence of hypothesis H; w k represents the kth evidence weight (determined by expert knowledge);

[0144] s k represents the kth evidence support degree (fuzzy evaluation result); m is the total number of possible hypotheses; n is the number of evidence is a multiplication operator, representing the product calculation of all items in the sequence k = 1 to n; w j

[0145] is the weight coefficient of the jth evidence; s j is the support degree of the jth evidence.

[0146] Step 500 at least contains steps S510-S530:

[0147] S510, first layer fast screening of the fault hypothesis set, retaining high probability fault hypotheses;

[0148] Firstly, the fault hypothesis set with confidence generated by step S400 is received, which contains multiple potential fault causes and their confidence scores. Specifically, the fast screening adopts a dual filtering mechanism based on confidence threshold and time sensitivity. Understandably, the confidence threshold is dynamically adjusted according to the fault type, setting a higher threshold for emergency accident type faults and a medium threshold for general operation problems.

[0149] Further, the screening process includes three parallel processing channels: the first channel processes high confidence hypotheses with confidence exceeding 0.85, which are directly retained; the second channel processes medium confidence hypotheses with confidence between 0.7 and 0.85, which are subjected to time series trend matching; the third channel processes low confidence hypotheses with confidence between 0.6 and 0.7, which are subjected to relevance verification. Specifically, the time series trend matching extracts the relevant parameter change trend of the last 72 hours from the standardized data set of step S100, and calculates the matching degree with the fault hypothesis.

[0150] The implementation of retaining high probability fault hypotheses is as follows: the output results of each channel are merged, and duplicates are removed to form a preliminary screened fault hypothesis subset. Specifically, each hypothesis retained in the subset satisfies at least one of the following conditions: confidence exceeds the type-related threshold; parameter trend matching degree exceeds 0.75; appears continuously in three consecutive sampling periods. Understandably, this screening mechanism ensures that key fault hypotheses are retained while significantly reducing the computational load of subsequent processing.

[0151] S520, combined with real-time data in the standardized data set, second layer fine verification is performed;

[0152] Real-time monitoring data at the current time is acquired, which is from the latest update of the standardized dataset generated in S100. Specifically, the verification process adopts a multidimensional cross-checking method, including three verification modules of parameter real-time value checking, device state correlation analysis, and operation record comparison.

[0153] Further, the implementation process of the parameter real-time value checking is as follows: for each fault hypothesis, the standard range of its diagnostic parameters is extracted, and deviation analysis is performed with the real-time monitoring value. The device state correlation analysis obtains the device running state code from the SCADA system, and verifies the logical consistency of the fault hypothesis with the device state. The operation record comparison retrieves the recent operation record from the production management database, and confirms whether the fault hypothesis has time correlation with the operation change.

[0154] Specifically, the fine verification adopts a weighted scoring mechanism, and assigns weights to each verification module: the parameter checking weight is 0.5, the device state weight is 0.3, and the operation record weight is 0.2. The verification score of each hypothesis is calculated as the sum of the product of the module score and the weight. Understandably, the hypothesis with a score exceeding 0.8 is marked as verified, the hypothesis with a score of 0.6-0.8 enters the review process, and the hypothesis with a score lower than 0.6 is directly eliminated.

[0155] The output of the fine verification is a set of high-reliability hypotheses that pass the verification, each hypothesis being accompanied by a detailed verification report. The report contains the specific scores of each verification module, the time stamp and value of the key evidence data points, and the final verification conclusion. This output serves as the input of S530, ensuring that the impact assessment is based on fault hypotheses that have undergone strict verification.

[0156] S530, based on the propagation path in the weighted topology graph, assesses the third layer of influence range, and generates a confirmed fault diagnosis conclusion;

[0157] First, the connection relationship and propagation path of the fault point are extracted from the weighted topology graph generated in S200. Specifically, the assessment process includes two stages of local influence range analysis and system-level influence deduction. In the local influence analysis stage, the device nodes directly connected to the fault point are determined; in the system-level influence deduction stage, the propagation process of the fault along the high-sensitivity path is simulated.

[0158] Further, the impact assessment adopts a prediction method based on the fault propagation model, considering three key factors: the propagation speed determined by the fault type, the propagation path determined by the pipe network topology structure, and the resistance determined by the device state. Specifically, the propagation speed is obtained from the expert knowledge base according to the preset value of the fault type; the propagation path is determined by the edge weight and node centrality in the weighted topology graph; and the device resistance is calculated from the device running parameters.

[0159] The implementation of the confirmed fault diagnosis conclusion is to integrate the three-level verification results to form a comprehensive diagnosis report containing fault location, cause analysis, influence range and emergency degree. Specifically, the structure of the report includes: fault point accurate position description (using topology identification), root cause analysis, direct impact equipment list, potential propagation risk prediction, and emergency treatment priority suggestion. Understandably, the report also contains the confidence index of each verification link for decision reference.

[0160] Step S600 includes at least steps S610-S630:

[0161] S610, according to the fault diagnosis conclusion, the influence degree of water injection system energy efficiency is calculated;

[0162] First, the confirmed fault diagnosis conclusion generated in S500 step is analyzed, and the fault type, position coordinate and influence range parameters are extracted. Specifically, the fault type includes nine categories such as equipment performance degradation, pipe network leakage, valve regulation failure, etc., and each category of fault has a preset energy efficiency loss coefficient matrix. Understandably, the position coordinates come from the topology identification in the diagnosis conclusion, which completely corresponds to the weighted topology graph space position in S200 step.

[0163] Further, the calculation process adopts a multi-dimensional influence superposition algorithm: the first dimension calculates the direct energy consumption loss based on the rated power and operating efficiency deviation value of the fault equipment; the second dimension calculates the conduction performance loss according to the upstream and downstream relationship in the pipe network topology structure; the third dimension calculates the system regulation loss according to the pressure regulation amount made by the water injection station to compensate for the fault. Specifically, the calculation of the conduction performance loss needs to combine the high sensitivity path identified in S230 step to determine the main propagation channel of energy loss.

[0164] The quantitative output form of the influence degree is the energy efficiency loss index, including absolute value (kWh) and relative value (percentage). Specifically, the index calculation needs to refer to the historical energy efficiency baseline value in the S100 step standardized data set to obtain the accurate loss amount by comparing the energy efficiency changes before and after the fault. Understandably, the calculation process distinguishes between short-term recoverable loss and long-term structural loss, providing a basis for maintenance scheme selection.

[0165] S620, based on the weighted topology graph, the energy efficiency improvement effect of different maintenance schemes is simulated;

[0166] First, three typical maintenance scheme models are constructed: the emergency repair scheme for key equipment failure, the optimized scheduling scheme for pipe network operating parameters, and the system reconstruction scheme for structural defects. Specifically, the input parameters of the scheme model include: the specific defect description in the fault diagnosis conclusion, the edge weight parameter calculated in S220 step, and the operating state score in S320 step.

[0167] Further, the simulation analysis adopts a space-time deduction mechanism: in the spatial dimension, the spread range of the scheme implementation is predicted according to the connection relationship of the weighted topological graph; in the time dimension, three stages of emergency treatment period (0-6 hours), short-term recovery period (6-72 hours) and long-term optimization period (more than 72 hours) are simulated respectively. Specifically, the simulation process considers the maintenance resource constraint conditions, including the number of repair teams, spare parts inventory, construction window period and other realistic factors.

[0168] The evaluation index of the energy efficiency improvement effect adopts a comprehensive benefit value, which is calculated by weighting the energy efficiency improvement amount, implementation cost and influence duration. Specifically, the energy efficiency improvement amount is obtained by comparing the system energy consumption values before and after maintenance; the implementation cost is obtained by calling the standard material library data according to the type of the maintenance scheme; and the influence duration considers the construction complexity and system recovery time. Understandably, each scheme generates a predicted effect curve graph to show the energy efficiency change trend in different periods.

[0169] S630, selecting an optimal maintenance strategy and updating the weight parameters of the pipe network topological structure;

[0170] First, a multi-objective decision matrix is established, including energy efficiency improvement value, implementation cost, effect speed and risk coefficient. Specifically, the decision adopts the analytic hierarchy process to determine the weight of each dimension: the production tense period focuses on the effect speed, the cost control period focuses on the implementation cost, and the energy efficiency optimization period focuses on the improvement amplitude. Further, the selection of the optimal strategy needs to perform sensitivity analysis to verify the stability of the decision under different working conditions.

[0171] The implementation process of updating the weight parameters of the pipe network topological structure includes: according to the type of the adopted maintenance scheme, determining the set of topological elements that need to be updated. For equipment repair type scheme, updating the equipment state coefficient of the corresponding node; for pipe network adjustment type scheme, updating the resistance coefficient of the related pipe segment; for system reconstruction type scheme, updating the topological connection relationship. Specifically, the updated value is calculated by reverse calculation of the maintenance effect simulation data, and a smooth transition algorithm is adopted to avoid parameter mutation.

[0172] The update operation adopts a version management mechanism: saving the current weight parameters as a historical version, recording the update time and maintenance scheme number; writing new weight parameters and activating them. Understandably, the updated weight parameters will participate in the edge weight calculation in the next round of S220 step, forming a closed loop optimization. Specifically, the update result generates a parameter change report, which is archived synchronously with the maintenance scheme.

[0173] Obviously, the above-described embodiments are only some embodiments but not all the embodiments of the present application, the preferred embodiments of the present application are shown in the drawings, but do not limit the patent scope of the present application. The present application can be implemented in many different forms, and conversely, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent replacements to some technical features therein. Any equivalent structure made by using the content of the specification and drawings, directly or indirectly applied to other related technical fields, is also within the patent protection scope of the present application.

Claims

1. A method for diagnosing faults in large oilfield water injection systems, characterized in that: include: Collect injection well flow, pressure, and energy consumption data, clean and standardize them, and generate standardized data sets with topological identifiers; Construct the pipe network topology based on the standardized data set, calculate the edge weights and node centrality, and generate a weighted topology graph with feature annotations; Based on the data of the standardized data set and weighted topological map, the health score is calculated by combining topological features, identifying underfilling, overfilling and leakage abnormal patterns, and generating a subsystem health status report; Based on the abnormal patterns in the subsystem health status report, the corresponding fault tree type is activated, fuzzy node condition evaluation and potential fault cause combination deduction are performed, and a set of fault hypotheses with confidence levels is generated; The expression of the fault hypothesis set with confidence is: Among them, CF(H) represents the comprehensive confidence of hypothesis H; w k represents the kth evidence weight; s k represents the support of the kth piece of evidence; m is the total number of possible hypotheses; n is the number of evidences for the current hypothesis; is the multiplication operator; w j is the weight coefficient of the jth evidence; s j is the support of the j-th evidence; Perform three-level verification on the fault hypothesis set, performing rapid screening, detailed verification, and impact assessment to generate confirmed fault diagnosis conclusions; Based on the fault diagnosis conclusions, the impact of the fault on energy efficiency is evaluated, the effects of different maintenance plans are simulated, the optimal maintenance strategy is selected, and the weight parameters of the pipeline network topology are updated.

2. The method according to claim 1, characterized in that The step of generating a standardized data set with topology identification includes: Obtain injection well flow data, pipe network pressure data, and injection station energy consumption data, fill in missing values ​​and remove outliers to obtain cleaned raw data; Extract time series features from the cleaned raw data, perform sampling rate alignment and unit standardization to obtain time-aligned data; Spatial location labels are attached to the time-aligned data to generate a standardized dataset with topological identification.

3. The method according to claim 1, characterized in that The step of generating a weighted topological graph with feature annotations includes: Based on the spatial location labels in the standardized dataset, the network topology of nodes and edges is constructed; Obtain flow and pressure parameters from standardized data sets and calculate network edge weights and node centrality; According to the pipeline network topology structure, pipeline network edge weights and node centrality, high-sensitivity paths are identified and a weighted topology graph with feature annotations is generated.

4. The method according to claim 1, wherein The steps of generating a subsystem health status report include: Pressure and flow indicators are selected from the standardized data set, and fuzzy membership conversion is performed to obtain fuzzy linguistic variables; The operation status scores of injection wells and pipe networks are calculated by combining the path characteristics and fuzzy linguistic variables in the weighted topological graph. Identify underfill, overfill, and leakage anomaly patterns based on operational status scores and generate subsystem health status reports.

5. The method according to claim 1, wherein The step of activating the corresponding fault tree type based on the abnormal pattern in the subsystem health status report includes: Based on the abnormal pattern in the subsystem health status report, activate the corresponding fault tree type; Fault tree types include emergency fault tree, water injection condition fault tree of water injection well, energy consumption fault tree of water injection station, energy consumption fault tree of water injection network and energy consumption fault tree of water injection well.

6. The method according to claim 1, characterized in that The steps of performing the combined derivation of fuzzy node condition evaluation and potential fault causes include: Combining the fault tree type with the parameters in the standardized data set, a fuzzy evaluation of the fault tree node conditions is performed; Fuzzy evaluation involves converting quantitative parameters into fuzzy linguistic variables and making conditional judgments based on membership functions.

7. The method according to claim 1, characterized in that The step of generating a fault hypothesis set with confidence levels includes: The potential fault cause combinations are derived based on the fuzzy evaluation results, and a set of fault hypotheses with confidence levels is generated.

8. The method according to claim 1, characterized in that The step of generating a confirmed fault diagnosis conclusion includes: Perform a first-level quick screening of the fault hypothesis set and retain high-probability fault hypotheses; Combined with real-time data from standardized datasets, a second layer of fine-grained verification is performed; Based on the propagation path in the weighted topology graph, the third-layer impact range is evaluated and a confirmed fault diagnosis conclusion is generated.

9. The method according to claim 1, characterized in that The steps of selecting the optimal maintenance strategy and updating the weight parameters of the pipe network topology structure include: Calculate the impact on the energy efficiency of the water injection system based on the fault diagnosis conclusion; Simulate the energy efficiency improvement effects of different maintenance schemes based on weighted topology graphs; Select the optimal maintenance strategy and update the weight parameters of the pipe network topology.

10. The method according to claim 9, characterized in that The step of selecting the optimal maintenance strategy and updating the weight parameters of the pipe network topology structure also includes: The updated weight parameters of the pipe network topology are fed back to the pipe network topology construction process to form a closed-loop optimization system.