Knowledge and data fusion-based interpretable pumping unit pump condition diagnosis method

By expanding the confidence rule base and combining indicator diagram features with expert knowledge, the problem of insufficient accuracy and robustness in pump condition diagnosis of oil pumping units is solved, and efficient fault identification and intelligent management under complex operating conditions are achieved.

CN121407927AActive Publication Date: 2026-01-27SHANDONG JIANZHU UNIV

Patent Information

Application Number
CN202511962436.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-01-27
Estimated Expiration
2045-12-24

AI Technical Summary

Technical Problem

Existing pump condition diagnostic methods for oil pumping units are insufficient in terms of accuracy, interpretability, and robustness. In particular, they are difficult to accurately identify faults such as gas interference, valve leakage, and insufficient fluid supply under complex operating conditions, and they fail to fully integrate expert knowledge and qualitative factors.

Method used

A knowledge- and data-fusion approach is adopted, which combines the extended confidence rule base (EBRB) inference model with dynamometer features, expert knowledge and data-driven information for fault diagnosis, and uses a genetic algorithm to optimize model parameters to build an online identification system for real-time diagnosis.

Benefits of technology

It improves the accuracy and robustness of pump condition diagnosis for oil pumping units, can identify common faults under complex operating conditions, has good interpretability and adaptability, and supports intelligent management and predictive maintenance of oilfields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121407927A_ABST
    Figure CN121407927A_ABST
Patent Text Reader

Abstract

The invention discloses an interpretable oil pumping unit pump condition diagnosis method based on knowledge and data fusion, which relates to the field of agricultural equipment, and is characterized by comprising the following steps: S1, collecting multi-source data of operation of an oil pumping unit, including an indicator diagram, stroke, displacement, load, voltage and current; s2, carrying out preprocessing and feature extraction on the collected data, and constructing a feature set comprising geometric features, derived indexes and electrical parameter features; and S3, inputting the features into an extended belief rule base reasoning model, and fusing expert knowledge and data driving information. The technical problem to be solved by the invention is to provide the interpretable pumping unit pump condition diagnosis method based on knowledge and data fusion, the method has relatively strong generalization, interpretability and robustness, and is beneficial to improving intelligent management and predictive maintenance of oil field operation, so that a high-intelligence and high-efficiency decision-making process is realized, and the working efficiency is improved. And the goal of a knowledge-based intelligent system is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and petroleum machinery diagnostic technology, and more specifically, to an interpretable pump condition diagnostic method for oil pumping units based on knowledge and data fusion. Background Technology

[0002] Currently, pump condition diagnosis for oil pumping units mainly relies on manual judgment based on experience (such as manual interpretation of dynamometer diagrams) or a single data-driven method (such as deep learning classifiers). Manual judgment depends on expert experience, which is inefficient and easily influenced by subjective factors. While pure data-driven methods can improve diagnostic accuracy to some extent, they often lack interpretability and robustness, making it difficult to adapt to complex and changing on-site conditions. Especially under typical fault modes such as gas interference, valve leakage, and insufficient fluid supply, existing methods are prone to inter-class confusion, resulting in insufficient diagnostic accuracy and reliability. Therefore, there is an urgent need for a new diagnostic model that can maintain the high accuracy advantages of data-driven methods while possessing knowledge interpretability and robustness.

[0003] Currently, most oil wells in China and around the world still rely on rod pumps for production. However, rod pumping systems are highly susceptible to malfunctions during operation, which not only affects normal oilfield production and reduces crude oil output but also significantly increases production costs. Therefore, timely and accurate understanding of the operating conditions of rod pumping systems and the identification and diagnosis of problems in the oil production process are crucial for improving oilfield development efficiency and recovery benefits. For a long time, the study of pumping well malfunctions has been a key topic in the field of oil production engineering. Existing domestic and international research largely relies on dynamometer cards for analysis and diagnosis, and accurate identification of dynamometer cards is considered a crucial step in fault diagnosis. By analyzing the load, displacement, and other relevant information contained in the dynamometer card, various operating conditions of the pumping system can be assessed.

[0004] However, existing methods for diagnosing faults in industrial oil well equipment still have the following shortcomings:

[0005] Insufficient handling of qualitative factors: Existing methods often focus on quantitative data while neglecting qualitative factors such as equipment operating status, environmental conditions, and operator experience. Although these factors are difficult to quantify, they have a significant impact on fault diagnosis results. For example, when equipment operates under high temperature, high humidity, or high load conditions, its fault modes may differ from those under normal operating conditions. Ignoring these factors can lead to one-sided diagnostic results that fail to fully reflect the nature of the fault.

[0006] Reliance on complete datasets: During operation, industrial equipment often experiences data loss due to sensor malfunctions, data transmission interruptions, or storage errors. Existing diagnostic methods typically rely on complete datasets for modeling and analysis. When data is incomplete, diagnostic models often fail to function properly or their accuracy decreases significantly, limiting their applicability in complex field environments.

[0007] Lack of Expert Knowledge Integration: Expert experience is a crucial resource in fault diagnosis, providing prior information, identifying rare patterns, and assisting in the analysis of complex faults. However, most existing methods are data-driven and fail to fully utilize and integrate expert knowledge. When faced with complex or rare fault conditions, the lack of experience support often leads to unreliable diagnostic results.

[0008] In summary, existing pump condition diagnostic methods for oil pumping units still have significant shortcomings in terms of accuracy, interpretability, and robustness. There is an urgent need to propose a new diagnostic method that can integrate data-driven approaches and expert knowledge to improve fault identification capabilities and reliability under complex operating conditions. Summary of the Invention

[0009] The technical problem this invention aims to solve is to provide an interpretable method for diagnosing sucker rod pump conditions based on knowledge and data fusion. This method first employs an advanced feature extraction, reconstruction, and fusion approach, combining domain expertise with the structural feature mechanisms supporting the sucker rod system. Subsequently, an Extended Confidence Rule Base (EBRB) inference model is applied for fault classification. In this framework, key information features extracted from the dynamometer card serve as antecedents, while the diagnostic category and decision category serve as consequents. The ERB is established using actual operating data, enabling it to perform effective inference based on complex dynamometer card conditions and real-world scenarios. The model achieves information fusion through Evidence-Based Reasoning (ER) algorithms, balancing accuracy and interpretability in its judgment results. Compared to traditional data-driven methods, this method demonstrates stronger robustness and adaptability under diverse and complex oilfield conditions. At the application level, this method improves the ability to identify common faults such as gas interference, valve leakage, and insufficient fluid supply. By combining expert knowledge with data-driven approaches, the proposed feature extraction and inference framework represents a more comprehensive and effective approach to diagnosing sucker rod pump faults. This method has strong generalization, interpretability, and robustness, which helps to improve the intelligent management and predictive maintenance of oilfield operations, thereby achieving a highly intelligent and efficient decision-making process, which is in line with the goal of knowledge-based intelligent systems.

[0010] The present invention achieves its objective by employing the following technical solution:

[0011] An interpretable pump condition diagnosis method for oil pumping units based on knowledge and data fusion, characterized by the following steps:

[0012] S1: Collect multi-source data on the operation of the pumping unit, including dynamometer diagram, stroke, displacement, load, voltage, and current;

[0013] S2: Preprocess and extract features from the collected data to construct a feature set including geometric features, derived indices, and electrical parameter features;

[0014] S3: Input the features into the Extended Belief Rule-Based (EBRB) inference model and fuse them with expert knowledge and data-driven information;

[0015] S4: Optimize the parameters of the extended confidence rule base inference model using a genetic algorithm, and construct the optimal rule set by combining rule merging and pruning mechanisms;

[0016] S5: Output the classification results of the pump condition of the oil pumping unit and give the diagnostic confidence level.

[0017] As a further limitation of this technical solution, the method is based on an online recognition system, which includes:

[0018] Data acquisition and preprocessing module: acquires multi-source data such as dynamometer diagram, electrical parameters (voltage, current), displacement, stroke, load, and temperature in real time from sensors and monitoring equipment at the pumping unit site; and cleans, normalizes, and extracts features from the acquired data to ensure the quality of the input data;

[0019] Online learning and model training module: The system embeds an extended confidence rule base and a deep learning hybrid inference model, which can dynamically adjust parameters according to new input samples to achieve continuous online training of the model; it introduces rule optimization and pruning mechanisms to avoid rule base redundancy and improve inference efficiency and interpretability; it supports custom optimization functions, taking diagnostic accuracy, robustness and generalization ability as joint optimization objectives;

[0020] Fault Diagnosis and Classification Module: The model can classify and diagnose common operating conditions such as gas interference, valve leakage, insufficient liquid supply, sucker rod breakage, wax deposition, and sand production; while outputting diagnostic results, it also provides confidence scores and possible causes of failure to facilitate decision-making by engineers;

[0021] Real-time monitoring and alarm module: The system supports real-time monitoring of the operating status of oil wells. When the monitoring parameters exceed the threshold or the diagnostic results indicate a potential fault, an alarm is automatically triggered. Alarm information can be pushed to maintenance personnel through a visual interface, large screen display or mobile terminal to achieve remote monitoring and early warning.

[0022] Data storage and management module: The system uses a database to uniformly store and manage historical operating data, training samples, and diagnostic results, supporting data traceability and playback; it provides visualization analysis functions, including trend analysis, anomaly detection, and statistical reports, to assist engineers in decision-making.

[0023] As a further limitation of this technical solution, in S2, the valve working position and seven geometric features are extracted from the indicator diagram;

[0024] The quadrilateral area S_ABCD of the four valve working positions, the area of ​​the upper left region of the indicator diagram S_L-up, the area of ​​the upper right region of the indicator diagram S_R-up, the area of ​​the lower right region of the indicator diagram S_R-down, the area of ​​the lower left region of the indicator diagram S_L-down, the working displacement D_BC of the fixed valve, and the working displacement D_DA of the traveling valve.

[0025] Features 2 through 5 describe the area of ​​the four parts of the indicator diagram from the four positions of valve operation, and the sum of these four feature values ​​describes the area of ​​the indicator diagram as a whole.

[0026] As a further limitation of this technical solution, the derived index with a clear physical meaning is constructed in S2:

[0027] α = S_ABCD: Total area of ​​the quadrilateral, which approximately represents the single-stroke liquid production or work capacity;

[0028] η_up = D_BC, η_down = D_DA: represent the effective stroke lengths of upward inhalation and downward exhalation, respectively;

[0029] κ_SV=(S_L-up – S_R-up) / (S_L-up + S_R-up): The asymmetry between the left and right sides of the upper half of the graph, reflecting the timing offset of the station's on / off adjustment.

[0030] κ_TV=(S_R-down – S_L-down) / (S_R-down + S_L-down): The asymmetry of the lower half of the graph reflects the valve opening / closing timing offset.

[0031] ρ = η_upper / η_lower: Relative balance of the effective strokes of the upper and lower parts;

[0032] Δ_UD_hat = (sum of areas in the upper half – sum of areas in the lower half) / (α): Lower energy / partial balance index.

[0033] Based on experience: low α ⇒ insufficient liquid production / liquid slugging; low η_low and high κ_TV are common in operating valve leakage; low η_up and high κ_SV are common in station valve leakage; when η_up and η_down decrease simultaneously and are symmetrical, it is often related to gas interference; if the geometry is basically symmetrical but ρ deviates from 1, it is often a timing / configuration logic problem.

[0034] As a further limitation of this technical solution, the construction process of the extended confidence library is as follows: determine the reference values ​​of the antecedent attributes and the result attributes, transform the confidence structure of the input and output data, and set the weights of each antecedent attribute and rule.

[0035] The steps for optimizing the parameters of the extended confidence rule base in S4 are as follows:

[0036] S41: Preprocess the raw data, reduce noise and normalize the data of the dynamometer diagram to avoid affecting the calculation due to the different dimensions of load and displacement;

[0037] S42: Perform feature engineering on the preprocessed data. Based on quantitative and mechanistic analysis, extract the working position of the pump valve and seven geometric features from the indicator diagram. Specifically, the valve working position is extracted from the indicator diagram using curvature and centroid decomposition methods. The specific steps are as follows:

[0038] Find the centroid of each indicator diagram. An indicator diagram is composed of discrete points; it can be viewed as a polygon composed of discrete points. Therefore, finding the centroid of an indicator diagram is equivalent to finding the centroid of a polygon.

[0039] Based on the position of the center of gravity, the dynamometer diagram is divided into four areas: upper left, upper right, lower right, and lower left.

[0040] Calculate the curvature of each discrete point in the four regions.

[0041] Extract the valve's operating position from four regions. The point with the greatest curvature in each of the four regions represents the corresponding valve's operating position.

[0042] S43: Construct an extended confidence library for oil pumping unit fault diagnosis and classification;

[0043] An extended confidence rule base typically includes several rules, each consisting of several antecedent attributes (rule antecedents) and several consequent attributes (rule consequents). For oil pumping fault diagnosis and classification, currently, based on the dynamometer card diagnosis method, the numerical information of displacement and load, and the extended confidence rule base that constitutes the dynamometer card, the antecedent is the characteristic factor of the dynamometer card, and the consequent attribute is the result of the diagnosis and classification.

[0044] S44: Perform parameter optimization for the expanded confidence rule base;

[0045] Specifically, joint optimization of the EBRB system is employed. Starting with the dataset used to train the EBRB model, the system is initialized based on this training set, including setting the initial structure and parameters. Subsequently, two optimization steps are performed: structure optimization and parameter optimization. In the structure optimization phase, the Relief F algorithm is used to identify and select the most informative features, thereby optimizing the structure of the EBRB model. Simultaneously, in the parameter optimization phase, the Differential Evolutionary Algorithm is used to adjust the model's parameters to improve its performance. This results in an optimal parameter set, representing the best structure and parameter configuration for the optimized EBRB system. Finally, this optimal parameter set is used to build a higher-performing EBRB model.

[0046] S45: Fault diagnosis based on extended confidence rule base.

[0047] As a further limitation of this technical solution, the specific steps of S41 are as follows:

[0048] Filter the acquired data by field, first filtering out the required fields, such as: current, stroke, displacement, and load physical characteristics.

[0049] Then, missing values ​​and outliers are handled. Missing values ​​are either deleted or filled in, and individual outliers are deleted. Then, the units of the data in each field are standardized.

[0050] Finally, the category information is converted into label encoding.

[0051] To reduce the impact of noise, an averaging filter is used to smooth the indicator curve.

[0052] As a further limitation of this technical solution, the specific steps for constructing the extended confidence database are as follows:

[0053] Determine the reference values ​​for the antecedent and result attributes;

[0054] Establish an identification framework for each attribute, that is, classify each attribute into levels, and set reference values ​​by experts based on the characteristics of each attribute and the range of attribute values;

[0055] Transform the confidence structure of input and output data;

[0056] Set the weights of each antecedent attribute and rule;

[0057] The weights of all antecedent attributes are set to 1, and will be adjusted later using optimization methods. Rule weights are a comprehensive measure of rule consistency, importance, and reliability, but currently, rule consistency is primarily used to reflect their weights.

[0058] As a further limitation of this technical solution, the specific steps of S45 are as follows:

[0059] Collect the antecedent attribute data of the fault to be diagnosed, and convert them into a confidence structure according to the confidence structure transformation formula of the input information. When the data of a certain attribute is missing, the confidence assignment value corresponding to that attribute is set to 0. The confidence structure transformation formula of the input information is as follows: (1)

[0060] in: Indicates the first The matching degree of each attribute at the j-th reference level is used to describe the closeness between the input feature value and the reference value at each level.

[0061] This represents the actual observed value of the i-th diagnostic attribute;

[0062] This represents the feature value or boundary value of the i-th attribute at the j-th reference level (or standard level);

[0063] The total number of attribute levels, that is, the number of levels or intervals into which each attribute is divided (e.g., J=5 for "very low, low, medium, high, very high").

[0064] (2)

[0065] in: It represents a set of matching relationships or feature vectors between attributes and levels, containing membership information of each attribute at different levels;

[0066] The total number of attributes in the M diagnostic system, i.e. the number of diagnostic features considered;

[0067] Calculate the activation weight of each rule based on the similarity between the antecedent attributes of the fault to be diagnosed and the antecedent attributes of the rules. The activation weight formula for the k-th rule is:

[0068] (3)

[0069] in: This represents the set of values ​​or observation data vectors of the input sample on the i-th attribute, that is, the actual measured value of the i-th feature input into the diagnostic model;

[0070] The attribute weight coefficients of the object to be diagnosed are used to reflect the importance of each attribute in the current diagnosis;

[0071] The weight coefficients of each attribute in the rule premise section are used to measure the strength of the rule's dependence on different attributes;

[0072] The set of attribute states in the premise of rule number l;

[0073] The premise of rule k is the set of values ​​or states of the r attributes, that is, the set of attribute values ​​in the actual observed data.

[0074] The activation value of the l-th rule represents the degree of similarity between the rule and the input sample;

[0075] L represents the total number of rules, that is, the number of rules contained in the rule base, which is used for normalization calculations among all rules;

[0076] Calculation result attribute reference value The combined reliability assignment value, where the result attribute value is calculated. Combined reliability allocation The formula is as follows:

[0077] (4)

[0078] in: Rule weights control the importance of each rule;

[0079] The confidence level of the k-th rule for the s-th conclusion;

[0080] The confidence level of the k-th rule in relation to the v-th conclusion during the fusion process;

[0081] N refers to the total number of rules in the reasoning system, i.e., the number of rules available for matching. This formula is a confidence synthesis formula based on the Evidence-Based Reasoning (ER) algorithm, used to weight and fuse the confidence assignments of each rule output into a final diagnostic result's composite confidence level.

[0082] Because this article assumes a reference value for the result attribute. And there is no unassigned reliability. Therefore, By directly substituting the corresponding utility value function, the calculated inference result is the classification value for diagnosing the fault. The formula for calculating the maximum utility value is as follows:

[0083] (5)

[0084] in: This represents the utility function corresponding to diagnostic objective D;

[0085] Indicates utility value;

[0086] This represents the overall confidence level corresponding to the Nth diagnostic conclusion;

[0087] Indicate conclusion Supplemental confidence level;

[0088] This indicates the Nth diagnostic conclusion;

[0089] The formula for calculating the minimum utility value is as follows:

[0090] (6)

[0091] in: Indicates the first diagnostic conclusion The confidence level is used to reflect the impact of lower confidence conclusions when calculating lower bound utility.

[0092] This indicates the diagnosis of the lowest level (or the lowest risk level).

[0093] As a further limitation of this technical solution, the optimization process of the parameters of the extended confidence rule base inference model adopts a custom optimization function composed of cross-entropy loss function and regularization term, so as to improve the robustness and interpretability of the model while ensuring classification accuracy.

[0094] As a further limitation of this technical solution, the classification results include a variety of typical fault types such as gas interference, valve leakage, insufficient liquid supply, sucker rod breakage, wax deposition, and sand production.

[0095] Compared with the prior art, the advantages and positive effects of the present invention are:

[0096] The first aspect of the present invention provides a method for constructing and fusing pump condition characteristics of an oil pumping unit. By extracting geometric features (such as S_ABCD, S_L-up, S_R-up, etc.) and electrical parameter features (current, voltage, power factor, etc.) from the indicator diagram, and constructing them using fusion and derived indices, a quantitative description of the pump condition is achieved.

[0097] A second aspect of this invention provides a method for classifying and reasoning about pumping unit conditions, using an Extended Confidence Rule Base (EBRB) model for diagnosis. This method combines expert knowledge with data-driven features, employs a genetic algorithm to optimize model parameters and rule weights, supports rule merging and pruning mechanisms, and defines a custom optimization function, thereby improving model interpretability and generalization while ensuring diagnostic accuracy.

[0098] A third aspect of this invention provides an online identification system for diagnosing pumping unit conditions. This system supports real-time data acquisition, model training, and inference, avoiding the limitations of traditional offline training. It can dynamically update the model based on new samples from the field, improving the timeliness and reliability of diagnosis. The system includes a data acquisition and preprocessing module, an online learning and training module, a fault classification module, a real-time monitoring and alarm module, and a data management module.

[0099] This invention integrates data-driven approaches with expert knowledge, improving the accuracy and robustness of diagnosis; it proposes a feature construction method based on centroid decomposition and multi-source feature fusion, enhancing the ability to describe complex working conditions; it applies an extended confidence rule base and genetic algorithm optimization to ensure the interpretability and adaptability of the diagnostic model; it introduces an online recognition system to achieve continuous learning and dynamic updating of the model, significantly improving the real-time performance and reliability of diagnosis; and it has good prospects for industrial application, supporting intelligent management and predictive maintenance in oilfields. Attached Figure Description

[0100] Figure 1 This is a schematic diagram illustrating the construction and fusion of pump condition characteristics of the oil pumping unit according to the present invention.

[0101] Figure 2 This is a flowchart of the classification reasoning process based on the extended confidence rule base of the present invention. Detailed Implementation

[0102] The following detailed description of a specific embodiment of the present invention is provided in conjunction with the accompanying drawings. However, it should be understood that the scope of protection of the present invention is not limited to the specific embodiment.

[0103] Example 1: A method for constructing and fusing pump condition characteristics based on geometric and electrical parameters, which describes in detail the calculation methods of geometric quantities and derived indices of the dynamometer diagram.

[0104] Based on quantitative and mechanistic analysis, the valve working position and seven geometric features were extracted from the indicator diagram; the quadrilateral area S_ABCD of the four valve working positions, the area of ​​the upper left region S_L-up of the indicator diagram, the area of ​​the upper right region S_R-up of the indicator diagram, the area of ​​the lower right region S_R-down of the indicator diagram, the area of ​​the lower left region S_L-down of the indicator diagram, the working displacement D_BC of the fixed valve, and the working displacement D_DA of the floating valve.

[0105] Features 2 through 5 describe the area of ​​the four parts of the indicator diagram from the four positions of valve operation, and the sum of these four feature values ​​describes the area of ​​the indicator diagram as a whole.

[0106] Further construct derived indices with clear physical meaning:

[0107] α = S_ABCD: Total area of ​​the quadrilateral, which approximately represents the single-stroke liquid production or work capacity;

[0108] η_up = D_BC, η_down = D_DA: represent the effective stroke lengths for upward inhalation and downward exhalation, respectively;

[0109] κ_SV = (S_L-up – S_R-up) / (S_L-up + S_R-up): The asymmetry of the upper half of the graph reflects the timing offset of the station's on / off adjustment.

[0110] κ_TV = (S_R-down – S_L-down) / (S_R-down + S_L-down): The asymmetry of the lower half of the graph, reflecting the valve opening / closing timing offset;

[0111] ρ = η_upper / η_lower: Relative balance of the effective strokes of the upper and lower parts;

[0112] Δ_UD_hat = (sum of areas in the upper half – sum of areas in the lower half) / (α): Lower energy / partial balance index.

[0113] Based on experience: low α ⇒ insufficient liquid production / liquid slugging; low η_low and high κ_TV are common in operating valve leakage; low η_up and high κ_SV are common in station valve leakage; when η_up and η_down decrease simultaneously and are symmetrical, it is often related to gas interference; if the geometry is basically symmetrical but ρ deviates from 1, it is often a timing / configuration logic problem.

[0114] Finally, electrical parameters and dynamometer diagram geometric features are integrated into a unified framework for fusion. At the implementation level, a feature-level fusion strategy is first adopted, concatenating the geometric features and electrical indicators from each work cycle into a comprehensive feature vector. After standardization, this vector is input into the confidence rule base inference model to support subsequent decision-making and analysis.

[0115] Example 2: Applying the Extended Confidence Rule Base (EBRB) for pump condition classification reasoning, combined with parameter optimization and rule pruning of the genetic algorithm.

[0116] The process of constructing the extended confidence base is as follows: determine the reference values ​​of the antecedent attributes and the result attributes, transform the confidence structure of the input and output data, and set the weights of each antecedent attribute and rule.

[0117] Perform parameter optimization for the expanded confidence rule base.

[0118] Specifically, for ease of understanding, the following description of the scheme described in this application is provided in conjunction with the accompanying drawings:

[0119] S41: Preprocess the raw data, reduce noise and normalize the data of the dynamometer diagram to avoid affecting the calculation due to the different dimensions of load and displacement;

[0120] The specific steps of S41 are as follows:

[0121] The data that has been obtained is filtered by field selection. First, the required fields are selected, such as fields of physical characteristics such as current, stroke, displacement, and load.

[0122] Then, missing values ​​and outliers are handled. Missing values ​​are either deleted or filled in, and individual outliers are deleted. Then, the units of the data in each field are standardized.

[0123] Finally, the category information is converted into label encoding.

[0124] To reduce the impact of noise, an averaging filter is used to smooth the indicator curve.

[0125] S42: Perform feature engineering on the preprocessed data. Based on quantitative and mechanistic analysis, extract the working position of the pump valve and seven geometric features from the indicator diagram. Specifically, the valve working position is extracted from the indicator diagram using curvature and centroid decomposition methods. The specific steps are as follows:

[0126] Find the centroid of each indicator diagram. An indicator diagram is composed of discrete points; it can be viewed as a polygon composed of discrete points. Therefore, finding the centroid of an indicator diagram is equivalent to finding the centroid of a polygon.

[0127] Based on the position of the center of gravity, the dynamometer diagram is divided into four areas: upper left, upper right, lower right, and lower left.

[0128] Calculate the curvature of each discrete point in the four regions.

[0129] Extract the valve's operating position from four regions. The point with the greatest curvature in each of the four regions represents the corresponding valve's operating position.

[0130] S43: Construct an extended confidence library for oil pumping unit fault diagnosis and classification;

[0131] An extended confidence rule base typically includes several rules, each consisting of several antecedent attributes (rule antecedents) and several consequent attributes (rule consequents). For oil pumping fault diagnosis and classification, currently, based on the dynamometer card diagnosis method, the numerical information of displacement and load, and the extended confidence rule base that constitutes the dynamometer card, the antecedent is the characteristic factor of the dynamometer card, and the consequent attribute is the result of the diagnosis and classification.

[0132] The specific steps for building an extended confidence base are as follows:

[0133] Determine the reference values ​​for the antecedent and result attributes;

[0134] Establish an identification framework for each attribute, that is, classify each attribute into levels, and set reference values ​​by experts based on the characteristics of each attribute and the range of attribute values;

[0135] Transform the confidence structure of input and output data;

[0136] Set the weights of each antecedent attribute and rule;

[0137] The weights of all antecedent attributes are set to 1, and will be adjusted later using optimization methods. Rule weights are a comprehensive measure of rule consistency, importance, and reliability, but currently, rule consistency is primarily used to reflect their weights.

[0138] S44: Perform parameter optimization for the expanded confidence rule base;

[0139] Specifically, joint optimization of the EBRB system is employed. Starting with the dataset used to train the EBRB model, the system is initialized based on this training set, including setting the initial structure and parameters. Subsequently, two optimization steps are performed: structure optimization and parameter optimization. In the structure optimization phase, the Relief F algorithm is used to identify and select the most informative features, thereby optimizing the structure of the EBRB model. Simultaneously, in the parameter optimization phase, the Differential Evolutionary Algorithm is used to adjust the model's parameters to improve its performance. This results in an optimal parameter set, representing the best structure and parameter configuration for the optimized EBRB system. Finally, this optimal parameter set is used to build a higher-performing EBRB model.

[0140] S45: Fault diagnosis based on extended confidence rule base.

[0141] The specific steps of S45 are as follows:

[0142] Collect the antecedent attribute data of the fault to be diagnosed, and convert them into a confidence structure according to the confidence structure transformation formula of the input information. When the data of a certain attribute is missing, the confidence assignment value corresponding to that attribute is set to 0. The confidence structure transformation formula of the input information is as follows: (1)

[0143] in: This represents the matching degree of the i-th attribute at the j-th reference level, used to describe the closeness between the input feature value and the reference value at each level;

[0144] This represents the actual observed value of the i-th diagnostic attribute;

[0145] This represents the feature value or boundary value of the i-th attribute at the j-th reference level (or standard level);

[0146] The total number of attribute levels, that is, the number of levels or intervals into which each attribute is divided (e.g., J=5 for "very low, low, medium, high, very high").

[0147] (2)

[0148] in: It represents a set of matching relationships or feature vectors between attributes and levels, containing membership information of each attribute at different levels;

[0149] The total number of attributes in the M diagnostic system, i.e. the number of diagnostic features considered;

[0150] Calculate the activation weight of each rule based on the similarity between the antecedent attributes of the fault to be diagnosed and the antecedent attributes of the rules. The activation weight formula for the k-th rule is:

[0151] (3)

[0152] in: This represents the set of values ​​or observation data vectors of the input sample on the i-th attribute, that is, the actual measured value of the i-th feature input into the diagnostic model;

[0153] The attribute weight coefficients of the object to be diagnosed are used to reflect the importance of each attribute in the current diagnosis;

[0154] The weight coefficients of each attribute in the rule premise section are used to measure the strength of the rule's dependence on different attributes;

[0155] The set of attribute states in the premise of rule number l;

[0156] The premise of rule k is the set of values ​​or states of the r attributes, that is, the set of attribute values ​​in the actual observed data.

[0157] The activation value of the l-th rule represents the degree of similarity between the rule and the input sample;

[0158] L represents the total number of rules, that is, the number of rules contained in the rule base, which is used for normalization calculations among all rules;

[0159] Calculation result attribute reference value The combined reliability assignment value, where the result attribute value is calculated. Combined reliability allocation The formula is as follows:

[0160] (4)

[0161] in: Rule weights control the importance of each rule;

[0162] The confidence level of the k-th rule for the s-th conclusion;

[0163] The confidence level of the k-th rule in relation to the v-th conclusion during the fusion process;

[0164] N refers to the total number of rules in the inference system, that is, the number of rules available for matching.

[0165] This formula refers to the confidence synthesis formula based on the Evidence Reasoning (ER) algorithm, which is used to fuse the confidence assignments of each rule output into a composite confidence score for the final diagnostic result according to weights.

[0166] Because this article assumes a reference value for the result attribute. And there is no unassigned reliability. Therefore, By directly substituting the corresponding utility value function, the calculated inference result is the classification value for diagnosing the fault. The formula for calculating the maximum utility value is as follows:

[0167] (5)

[0168] in: The utility function corresponding to the diagnostic target D is used to calculate the utility value of the input feature x under different diagnostic conclusions. It is the inference result function output by the fault diagnosis system.

[0169] It represents the utility value, which reflects the quality or importance of a diagnosis or conclusion in a specific situation;

[0170] This represents the overall confidence level (Belief Degree) corresponding to the Nth diagnostic conclusion, i.e., the system's confidence in the conclusion. The degree of credibility;

[0171] Indicate conclusion The supplementary confidence (or residual confidence) is used to reflect the amount of confidence that was not fully allocated and is used for normalization calculations;

[0172] This represents the Nth diagnostic conclusion, which is the last fault type or state category that the inference system may output.

[0173] The formula for calculating the minimum utility value is as follows:

[0174] (6)

[0175] in: Indicates the first diagnostic conclusion The confidence level is used to reflect the impact of lower confidence conclusions when calculating lower bound utility.

[0176] This indicates the diagnosis of the lowest level (or the lowest risk level).

[0177] The solution described in this application provides a learning method for pumping unit condition identification and classification (reasoning). Addressing the problems in current oil well equipment fault diagnosis, such as insufficient handling of qualitative factors of faults, reliance on complete datasets, and a data-driven approach that fails to fully integrate expert knowledge, this invention proposes a learning method for pumping unit fault diagnosis identification and classification (reasoning) based on the Confidence Rule Base (EBRB). Compared to previous pumping unit fault diagnosis identification and classification methods, this invention has the following improvements:

[0178] (1) Selection of geometric features of the indicator diagram: Since the distance between the working positions of the valve can objectively describe the working condition of the rod-operated plunger, based on the geometric features of the indicator diagram based on the working position of the valve, seven geometric features of the indicator diagram are further proposed to better characterize the indicator diagram.

[0179] (2) The proposed center-of-gravity decomposition method: Based on quantitative and mechanistic analysis, the valve working position and seven geometric features are extracted from the dynamometer card. However, in oilfields, the dynamometer diagram is often irregular due to various constraints (such as gas influence and feed fluid failure). The distribution of valve working positions is also often affected. Therefore, a method called center-of-gravity decomposition is proposed to provide a reasonable segmentation strategy for the dynamometer card, thereby enabling the computer to automatically identify the potential region of each valve working position under different operating conditions.

[0180] (3) Classification method for oil pumping unit fault diagnosis based on extended confidence rule base reasoning: Extended confidence rule base reasoning is an uncertain reasoning technique developed on the basis of DS evidence theory. It adopts the IFTHEN knowledge representation structure based on trust degree, which can well represent the fuzzy, uncertain and incomplete information in the system, and is not limited by the number of rule antecedents. The reasoning mechanism is clear and it is currently widely used in pattern recognition, risk assessment and consumer behavior prediction.

[0181] In order to improve the adaptability and accuracy of the model, this invention introduces mechanisms such as parameter optimization, rule merging and rule pruning: (1) Model parameter optimization: Genetic algorithm (GA) is used to optimize the parameters in the rule base globally to avoid getting trapped in local optima;

[0182] (2) Rule merging: Merge rules with similar semantics or low contribution to reduce redundancy and improve model simplicity;

[0183] (3) Rule pruning: Invalid or interfering rules are deleted to ensure the effectiveness of the rule base and the efficiency of inference. In addition, this invention defines an optimization function that combines the cross-entropy loss function with the regularization term, which not only ensures classification accuracy but also avoids model overfitting. The goal of the optimization function is to minimize classification error while maintaining the rationality and interpretability of the model structure. In the implementation process, this method adopts an iterative optimization strategy based on genetic algorithms. First, an initial population containing expert priors is generated. After decoding, fitness calculation, selection, crossover and mutation, the optimal rule set is gradually evolved. This optimization process introduces an elite retention mechanism to ensure that the optimal solution is not lost during the iteration process and terminates when the population fitness remains unchanged for a long time, outputting the final optimization parameters and rule set. Finally, the method of this invention can achieve efficient classification of common fault types of pumping units (such as gas interference, valve leakage, insufficient liquid supply, etc.) and shows high accuracy, recall and F1 score in performance evaluation. This method combines the high accuracy of data-driven and the interpretability and robustness of knowledge-driven methods, and is suitable for intelligent pump condition diagnosis under complex working conditions.

[0184] Example 3: A real-time diagnostic method based on an online identification system demonstrates the complete process of data acquisition, model updating, and real-time alarm.

[0185] This system overcomes the limitations of traditional machine learning models that rely on offline training and cannot adapt to dynamic working conditions. It enables real-time data collection, training, and application on-site, thereby significantly improving the timeliness and reliability of diagnosis.

[0186] The overall architecture of this system includes the following core modules:

[0187] Data acquisition and preprocessing module: Real-time acquisition of multi-source data such as dynamometer diagrams, electrical parameters (voltage, current), displacement, stroke, load, and temperature from on-site sensors and monitoring equipment of the pumping unit; and cleaning, normalization, and feature extraction of the acquired data to ensure the quality of input data.

[0188] Online learning and model training module: The system embeds an extended confidence rule base (EBRB) and a deep learning hybrid inference model, which can dynamically adjust parameters based on new input samples to achieve continuous online training of the model; it introduces rule optimization and pruning mechanisms to avoid rule base redundancy and improve inference efficiency and interpretability; it supports custom optimization functions, taking diagnostic accuracy, robustness and generalization ability as joint optimization objectives.

[0189] Fault Diagnosis and Classification Module: The model can classify and diagnose common operating conditions such as gas interference, valve leakage, insufficient fluid supply, sucker rod breakage, wax deposition, and sand production; while outputting diagnostic results, it also provides confidence scores and possible causes of failure, which facilitates decision-making by engineers.

[0190] Real-time monitoring and alarm module: The system supports real-time monitoring of the operating status of oil wells. When the monitoring parameters exceed the threshold or the diagnostic results indicate a potential fault, an alarm is automatically triggered. Alarm information can be pushed to maintenance personnel through a visual interface, large screen display or mobile terminal to achieve remote monitoring and early warning.

[0191] Data storage and management module: The system uses a database to uniformly store and manage historical operating data, training samples, and diagnostic results, supporting data traceability and playback; it provides visualization analysis functions, including trend analysis, anomaly detection, statistical reports, etc., to assist engineers in decision-making.

[0192] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for interpretable pump condition diagnosis of oil pumping units based on knowledge and data fusion, characterized in that, Includes the following steps: S1: Collect multi-source data on the operation of the pumping unit, including dynamometer diagram, stroke, displacement, load, voltage, and current; S2: Preprocess and extract features from the collected data to construct a feature set including geometric features, derived indices, and electrical parameter features; S3: Input features into the extended confidence rule base inference model and fuse them with expert knowledge and data-driven information; S4: Optimize the parameters of the extended confidence rule base inference model using a genetic algorithm, and construct the optimal rule set by combining rule merging and pruning mechanisms; S5: Output the classification results of the pump condition of the oil pumping unit and give the diagnostic confidence level.

2. The method for interpretable pumping unit condition diagnosis based on knowledge and data fusion according to claim 1, characterized in that: This method is implemented based on an online recognition system, which includes: Data acquisition and preprocessing module: acquires dynamometer diagrams, electrical parameters, displacement, stroke, load, and temperature data in real time from sensors and monitoring equipment at the pumping unit site; and cleans, normalizes, and extracts features from the acquired data to ensure the quality of the input data; Online learning and model training module: The system embeds an extended confidence rule base and a deep learning hybrid inference model, dynamically adjusting parameters based on new input samples to achieve continuous online training of the model; it introduces rule optimization and pruning mechanisms to avoid rule base redundancy and improve inference efficiency and interpretability; it supports custom optimization functions, using diagnostic accuracy, robustness and generalization ability as joint optimization objectives; Fault Diagnosis and Classification Module: The model can classify and diagnose common operating conditions such as gas interference, valve leakage, insufficient liquid supply, sucker rod breakage, wax deposition, and sand production; while outputting diagnostic results, it also provides confidence scores and possible causes of failure to facilitate decision-making by engineers; Real-time monitoring and alarm module: The system supports real-time monitoring of the operating status of oil wells. When the monitoring parameters exceed the threshold or the diagnostic results indicate a potential fault, an alarm is automatically triggered. Alarm information is pushed to maintenance personnel through a visual interface, large screen display or mobile terminal to achieve remote monitoring and early warning. Data storage and management module: The system uses a database to uniformly store and manage historical operating data, training samples, and diagnostic results, supporting data traceability and playback; and providing visualization analysis functions.

3. The method for interpretable pumping unit condition diagnosis based on knowledge and data fusion according to claim 2, characterized in that: In S2, the valve working position and seven geometric features are extracted from the indicator diagram; The quadrilateral area S_ABCD of the four valve working positions, the area of ​​the upper left region of the indicator diagram S_L-up, the area of ​​the upper right region of the indicator diagram S_R-up, the area of ​​the lower right region of the indicator diagram S_R-down, the area of ​​the lower left region of the indicator diagram S_L-down, the working displacement D_BC of the fixed valve, and the working displacement D_DA of the traveling valve.

4. The method for interpretable pumping unit condition diagnosis based on knowledge and data fusion according to claim 3, characterized in that: The derived index with a clear physical meaning is constructed in S2: α = S_ABCD: Total area of ​​the quadrilateral, which approximately represents the single-stroke liquid production or work capacity; η_up = D_BC, η_down = D_DA: represent the effective stroke lengths for upward inhalation and downward exhalation, respectively; κ_SV = (S_L-up – S_R-up) / (S_L-up + S_R-up): The asymmetry of the upper half of the graph reflects the timing offset of the station's on / off adjustment. κ_TV = (S_R-down – S_L-down) / (S_R-down + S_L-down): The asymmetry of the lower half of the graph reflects the valve opening / closing timing offset. ρ = η_upper / η_lower: Relative balance of the effective strokes of the upper and lower parts; Δ_UD_hat = (sum of areas in the upper half – sum of areas in the lower half) / (α): Lower energy / partial balance index.

5. The method for interpretable pumping unit condition diagnosis based on knowledge and data fusion according to claim 1, characterized in that: The steps for optimizing the parameters of the extended confidence rule base in S4 are as follows: S41: Preprocess the raw data, reduce noise and normalize the data of the dynamometer diagram to avoid affecting the calculation due to the different dimensions of load and displacement; S42: Perform feature engineering on the preprocessed data, and extract the working position of the pump valve of the oil pump and seven geometric features from the indicator diagram based on quantitative and mechanistic analysis; among them, the working position of the valve is extracted from the indicator diagram by combining curvature and centroid decomposition methods. S43: Construct an extended confidence library for oil pumping unit fault diagnosis and classification; S44: Perform parameter optimization for the expanded confidence rule base; S45: Fault diagnosis based on extended confidence rule base.

6. The method for interpretable pumping unit condition diagnosis based on knowledge and data fusion according to claim 5, characterized in that: The specific steps of S41 are as follows: Filter the acquired data by field, first selecting the required fields; Then, missing values ​​and outliers are handled. Missing values ​​are either deleted or filled in, and individual outliers are deleted. Then, the units of the data in each field are standardized. Finally, the category information is converted into label encoding.

7. The method for interpretable pumping unit condition diagnosis based on knowledge and data fusion according to claim 5, characterized in that: The specific steps for building an extended confidence base are as follows: Determine the reference values ​​for the antecedent and result attributes; Establish an identification framework for each attribute, that is, classify each attribute into levels, and set reference values ​​by experts based on the characteristics of each attribute and the range of attribute values; Transform the confidence structure of input and output data; Set the weights of each antecedent attribute and rule; The weights of each antecedent attribute are all set to 1, and will be adjusted later through optimization. Rule weight is a comprehensive measure of rule consistency, importance and reliability, but currently it is mainly reflected by the consistency of the rule.

8. The method for interpretable pumping unit condition diagnosis based on knowledge and data fusion according to claim 5, characterized in that: The specific steps of S45 are as follows: Collect the antecedent attribute data of the fault to be diagnosed, and convert them into a confidence structure according to the confidence structure transformation formula of the input information. When the data of a certain attribute is missing, the confidence assignment value corresponding to that attribute is set to 0. The confidence structure transformation formula of the input information is as follows: (1) in: This represents the matching degree of the i-th attribute at the j-th reference level; This represents the actual observed value of the i-th diagnostic attribute; This represents the feature value or boundary value of the i-th attribute at the j-th reference level; The total number of attribute levels, that is, the number of levels or intervals into which each attribute is divided; (2) in: It represents a set of matching relationships or feature vectors between attributes and levels, containing membership information of each attribute at different levels; The total number of attributes in the M diagnostic system, i.e. the number of diagnostic features considered; Calculate the activation weight of each rule based on the similarity between the antecedent attributes of the fault to be diagnosed and the antecedent attributes of the rules. The activation weight formula for the k-th rule is: (3) in: This represents the set of values ​​for the i-th attribute of the input sample or the vector of observed data. The attribute weight coefficients of the object to be diagnosed; The weighting coefficients of each attribute in the rule premise section; The set of attribute states in the premise of rule number l; The k-th rule presupposes a set of values ​​or states for r attributes; The activation level of rule l; L represents the total number of rules; Calculation result attribute reference value The combined reliability assignment value, where the result attribute value is calculated. Combined reliability allocation The formula is as follows: (4) in: Rule weights control the importance of each rule; The confidence level of the k-th rule for the s-th conclusion; The confidence level of the k-th rule in relation to the v-th conclusion during the fusion process; N refers to the total number of rules in the reasoning system; Will By directly substituting the corresponding utility value function, the calculated inference result is the classification value for diagnosing the fault. The formula for calculating the maximum utility value is as follows: (5) in: This represents the utility function corresponding to diagnostic objective D; Indicates utility value; This represents the overall confidence level corresponding to the Nth diagnostic conclusion; Indicate conclusion Supplemental confidence level; This indicates the Nth diagnostic conclusion; The formula for calculating the minimum utility value is as follows: (6) in: Indicates the first diagnostic conclusion Confidence level; This indicates the lowest level of diagnostic conclusion.

9. The method for interpretable pumping unit condition diagnosis based on knowledge and data fusion according to claim 1, characterized in that: The optimization process of the extended confidence rule base inference model parameters adopts a custom optimization function composed of cross-entropy loss function and regularization term to improve the robustness and interpretability of the model while ensuring classification accuracy.

10. The method for interpretable pumping unit condition diagnosis based on knowledge and data fusion according to claim 1, characterized in that: The classification results include gas interference, valve leakage, insufficient fluid supply, sucker rod breakage, wax deposition, and sand production.

Citation Information

Patent Citations

  • Method for constructing expert system based on knowledge discovery

    CN101093559A

  • Pumping well fault diagnosis expert system based on generation type rule

    CN109236277A

  • Large-scale motor fault diagnosis method based on combination of acoustic vibration signals and 1D-CNN

    CN112326210A

  • Unmanned aerial vehicle cluster system contribution degree evaluation method based on expansion belief rule reasoning

    CN117391494A

  • Extended rule inference network for generating extended belief rule library

    CN117540769A

Cited By

  • Physical knowledge guided model interpretability analysis method and diagnosis system

    CN121705661A