A Monitoring and Analysis Method for Tower Wells Based on Pearson Correlation Coefficient
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2026-08-14
AI Technical Summary
然而塔架机井的运行环境复杂,引起故障的因素较多,且引起故障的影响因素之间也存在较为复杂的相关关系,导致基于pearson相关系数的监测分析效果不佳、实时发出的报警信号不准确或延迟等问题,从而影响油井现场和集中监控管理人员及时查询故障信息,延误故障处理
[0034]本申请通过每种故障类型的故障数据与每个故障影响参数的监测数据之间的相关系数,确定每种故障类型对应的各主要影响参数,其有益效果在于考虑了影响故障产生的主要影响参数,以便提高后续故障预测的准确性;基于每种故障类型在各时刻的故障数据、以及每种故障类型对应的所有主要影响参数在各时刻的监测数据,构建多元线性回归模型,得到每种故障类型对应的回归方程及其影响图;分析所述影响图中数据点的分布情况以及异常数据点的占比程度,确定每种故障类型对应的回归方程的异常比重,其有益效果在于考虑了模型拟合过程中数据的异常情况,以反映异常数据对模型拟合的影响情况;通过所述影响图中数据点对应的偏离程度及其拟合误差情况,确定每种故障类型对应的回归方程的拟合影响权重,其有益效果在于考虑了每个数据点对应时刻采集的监测数据的异常可能性,以反映其对拟合回归的影响情况,进而说明回归模型整体的拟合质量不稳定性情况;融合所述异常比重与所述拟合影响权重,确定每种故障类型对应的回归方程的外部影响系数,其有益效果在于反映了回归模型的泛化能力,说明异常数据点对回归模型的扭曲预测结果的影响,从而明确回归模型对故障预测准确性的影响程度;根据每种故障类型对应不同主要影响参数之间的相关情况及其多重共线情况,结合所述外部影响系数,确定每种故障类型对应的回归方程的拟合不准确度,其有益效果在于不同主要影响参数之间的关联性以及多重共线性情况,以说明共线性对回归模型参数估计稳定性的影响;基于所述拟合不准确度对回归方程进行修正,得到每种故障类型对应的调整后的回归方程,实时对采油机井的故障类型进行预测预警,其有益效果在于通过对回归模型进行修正调整,提高回归模型对故障预测的准确性和稳定性,使得模型更加稳健,减少误报和漏报现象,从而降低维护成本,减少停机时间,提高塔架式抽油机的运行效率和安全性。
Smart Images

Figure CN121229067B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of tower well monitoring technology, specifically to a tower well monitoring and analysis method based on Pearson correlation coefficient. Background Technology
[0002] Tower-type pumping units are a new type of artificial lift oil production equipment. They utilize electro-mechanical reversing to drive the sucker rod and pump to reciprocate up and down, thus achieving oil extraction. Tower-type pumping units employ a direct balancing method, are suitable for long-stroke, low-stroke operation, and offer advantages such as high pump efficiency, easy balancing adjustment, stepless stroke and stroke adjustment, easy digital control, and significant energy savings. However, because the core components of the tower-type pumping unit, such as the motor and transmission system, are located on the top platform of the tower, maintenance and management personnel must climb to a height to conduct daily inspections, posing a significant safety hazard.
[0003] Currently, real-time monitoring of the core components of tower pumping units is used to establish a monitoring and analysis method based on Pearson correlation coefficients. Correlation analysis determines the relationship and characteristics between various fault types and the monitored data, predicts the operating trend of the pumping unit, diagnoses the location and potential faults, and issues real-time early warning signals for well site and centralized monitoring personnel to query fault information, thus preventing serious equipment damage and safety accidents. However, the operating environment of tower pumping units is complex, with numerous factors causing faults, and complex correlations exist between these factors. This leads to problems such as poor monitoring and analysis effectiveness based on Pearson correlation coefficients, inaccurate or delayed real-time alarm signals, and consequently, delays in timely fault information retrieval by well site and centralized monitoring personnel, hindering fault handling. Summary of the Invention
[0004] To address the aforementioned technical issues, a monitoring and analysis method for tower wells based on Pearson correlation coefficient is provided to resolve existing problems.
[0005] The solution to the technical problem in this application is to provide a monitoring and analysis method for tower wells based on Pearson correlation coefficient, including the following steps:
[0006] Real-time acquisition of fault data for each type of fault at each moment during the operation of the tower oil well, as well as monitoring data for each fault-affecting parameter at each moment;
[0007] By using the correlation coefficient between the fault data of each fault type and the monitoring data of each fault impact parameter, the main impact parameters corresponding to each fault type are determined.
[0008] Based on the fault data of each fault type at each time point, and the monitoring data of all major influencing parameters corresponding to each fault type at each time point, a multiple linear regression model is constructed to obtain the regression equation and its influence diagram for each fault type.
[0009] Analyze the distribution of data points in the influence graph and the proportion of abnormal data points to determine the abnormal weight of the regression equation corresponding to each fault type; determine the fitting influence weight of the regression equation corresponding to each fault type by the deviation of the data points in the influence graph and their fitting error; combine the abnormal weight and the fitting influence weight to determine the external influence coefficient of the regression equation corresponding to each fault type.
[0010] Based on the correlation and multicollinearity between the different main influencing parameters corresponding to each fault type, and in conjunction with the external influence coefficient, the fitting inaccuracy of the regression equation corresponding to each fault type is determined.
[0011] The regression equation is corrected based on the fitting inaccuracy to obtain the adjusted regression equation corresponding to each fault type, and the fault type of the oil well is predicted and warned in real time.
[0012] Preferably, each type of fault includes: motor fault, loose screw fault, and misalignment of the guide rod fault.
[0013] Preferably, each fault-affecting parameter includes: vibration of the reducer bearing, vibration of the reducer end, vibration of the motor bearing, motor temperature, temperature of the reducer drum lubricating oil, and displacement of the cantilever bracket.
[0014] Preferably, the correlation coefficient is the Pearson correlation coefficient calculated between the fault data of each fault type at all times and the monitoring data of each fault influence parameter at all times.
[0015] Preferably, determining the main influencing parameters corresponding to each fault type includes:
[0016] All fault-affecting parameters whose absolute values of the correlation coefficients are greater than a preset threshold are taken as all major-affecting parameters corresponding to each fault type.
[0017] Preferably, determining the anomaly weight of the regression equation corresponding to each fault type includes:
[0018] The number of outliers and high-leverage points in the influence graph is counted and recorded as the number of outliers; the total number of all data points in the influence graph is counted and recorded as the total number.
[0019] The ratio of the number of outliers to the total number is used as the outlier weight in the regression equation corresponding to each fault type.
[0020] Preferably, determining the fitting influence weights of the regression equation corresponding to each fault type includes:
[0021] The mean of the product of the normalized leverage value and the normalized studentized residual of all data points in the influence diagram is used as the fitting influence weight of the regression equation corresponding to each fault type.
[0022] Preferably, the external influence coefficient of the regression equation corresponding to each fault type is the product of the abnormality proportion and the fitting influence weight.
[0023] Preferably, determining the fitting inaccuracy of the regression equation corresponding to each fault type includes:
[0024] Analyze the average level of correlation among all major influencing parameters for each type of fault, and calculate the average correlation.
[0025] Analyze the average level of collinearity among all major influencing parameters corresponding to each fault type, and calculate the average collinearity.
[0026] The fitting inaccuracy B of the regression equation corresponding to the q-th fault type q The calculation formula is: Among them, A q Let μR be the external influence coefficient of the regression equation corresponding to the q-th fault type. q Let be the average correlation coefficient corresponding to the q-th fault type. Let q represent the average collinearity corresponding to the q-th fault type, max[] is the maximum value function, and ε is a preset value greater than 0.
[0027] Preferably, the average correlation is the mean of the absolute values of the Pearson correlation coefficients between the monitoring data of all two main influencing parameters corresponding to each fault type at all times.
[0028] Preferably, the average collinearity is the mean of the variance inflation factor of all major influencing parameters corresponding to each fault type.
[0029] Preferably, the adjusted regression equation corresponding to each fault type is: Y q ′ =exp(-B q )×Y q , where Y q ′ Let B be the adjusted regression equation corresponding to the q-th fault type. q Y represents the fitting inaccuracy of the regression equation corresponding to the q-th fault type. qLet q be the regression equation corresponding to the q-th fault type, and exp() be an exponential function with the natural constant as the base.
[0030] Preferably, the real-time prediction and early warning of fault types in oil wells includes:
[0031] Based on the monitoring data of all major influencing parameters corresponding to each fault type at the current moment, and combined with the adjusted regression equation, the predicted fault data for each fault type at the current moment is calculated.
[0032] If the predicted fault data is greater than or equal to the preset safety threshold for the corresponding fault type, the oil well has a fault; otherwise, the oil well has no fault.
[0033] This application has at least the following beneficial effects:
[0034] This application determines the main influencing parameters corresponding to each fault type by using the correlation coefficient between fault data and monitoring data of each fault influencing parameter. This has the advantage of considering the main influencing parameters affecting fault occurrence, thus improving the accuracy of subsequent fault prediction. Based on the fault data of each fault type at each time point and the monitoring data of all main influencing parameters corresponding to each fault type at each time point, a multiple linear regression model is constructed to obtain the regression equation and its influence map for each fault type. The distribution of data points and the proportion of abnormal data points in the influence map are analyzed to determine the abnormal weight of the regression equation corresponding to each fault type. This has the advantage of considering data anomalies during model fitting, reflecting the impact of abnormal data on model fitting. By analyzing the deviation and fitting error of the data points in the influence map, the fitting influence weight of the regression equation corresponding to each fault type is determined. This has the advantage of considering the possibility of anomalies in the monitoring data collected at each time point, reflecting its impact on the fitted regression, and thus indicating that the overall fitting quality of the regression model is unstable. The system considers various factors, including: 1) the external influence coefficient of the regression equation for each fault type, which is determined by integrating the anomaly proportion with the fitting influence weight. This coefficient reflects the generalization ability of the regression model, explains the impact of abnormal data points on the distorted prediction results of the regression model, and clarifies the degree of influence of the regression model on the accuracy of fault prediction. 2) the fitting inaccuracy of the regression equation for each fault type is determined based on the correlation and multicollinearity between different main influencing parameters, combined with the external influence coefficient. This inaccuracy reflects the correlation between different main influencing parameters and multicollinearity, illustrating the impact of collinearity on the stability of regression model parameter estimation. 3) the adjusted regression equation is obtained based on the fitting inaccuracy, providing real-time prediction and early warning of fault types in oil wells. This improves the accuracy and stability of the regression model in fault prediction by adjusting the model, making it more robust, reducing false alarms and false negatives, thereby reducing maintenance costs, downtime, and improving the operating efficiency and safety of tower pumping units. Attached Figure Description
[0035] The following section provides a more detailed description of the tower well monitoring and analysis method based on Pearson correlation coefficient proposed in this application, with reference to the accompanying drawings.
[0036] Figure 1 A flowchart illustrating the steps of the tower well monitoring and analysis method based on Pearson correlation coefficient provided in this application embodiment;
[0037] Figure 2A flowchart illustrating the steps of the method for obtaining the external influence coefficient of the regression equation corresponding to each fault type provided in the embodiments of this application. Detailed Implementation
[0038] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description of the tower well monitoring and analysis method based on Pearson correlation coefficient proposed in this application, in conjunction with the accompanying drawings and implementation examples, provides further elaboration. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit its scope.
[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0040] Please see Figure 1 The diagram illustrates a flowchart of a tower well monitoring and analysis method based on Pearson correlation coefficient provided in an embodiment of this application. The method includes the following steps:
[0041] Step 1: Real-time acquisition of fault data for each fault type at each moment during the operation of the tower oil well, as well as monitoring data for each fault-affecting parameter at each moment.
[0042] The system collects real-time monitoring data of core components during the operation of the pumping unit using sensors, such as bearing vibration data and motor or lubricating oil temperature data. The monitored data is displayed and stored in real time. Cloud computing is used to perform time-domain and frequency-domain analysis of the stored data. Based on pre-set safety thresholds and edge processing, the system makes a preliminary judgment on the real-time operating status of the pumping unit to diagnose core components and provide early warnings of potential faults. Edge warnings are provided through color, sound, and images and displayed on an LCD screen. For example, faults are diagnosed based on vibration data, temperature changes, and displacement deviations. The diagnostic results and fault handling suggestions are sent to the mobile terminals of pre-designated maintenance personnel. On-site maintenance personnel handle the faults based on the warning information on the mobile terminals. Each fault data point is managed and statistically analyzed to create an online fault early warning analysis database for the tower-type pumping unit, enabling it to autonomously learn new fault types and optimize fault identification.
[0043] New-type tower oil wells may experience various malfunctions during operation. For example, loose bolts can cause bearing vibration, increased oil temperature, increased motor temperature, and cantilever support misalignment. Malfunctions in oil wells can affect the oil production of the pumping unit. Therefore, it is necessary to monitor the core components of the pumping unit.
[0044] Preferably, in this embodiment, six core components are selected from the pumping unit for monitoring. In other implementation methods, the implementer can set them according to the actual situation.
[0045] In this embodiment, the six core components are: reducer bearing, reducer end, motor bearing, motor, reducer drum lubricating oil tank, and cantilever bracket. Therefore, vibration sensors are installed at the reducer bearing, reducer end, and motor bearing to monitor the vibration data of the reducer bearing, reducer end, and motor bearing in real time. Secondly, temperature sensors are installed near the oil inlet of the motor and reducer drum lubricating oil tank to monitor the oil temperature data of the reducer drum and the motor in real time. A displacement sensor is installed at the cantilever bracket to monitor the displacement data of the cantilever bracket in real time.
[0046] The system collects fault data when the mechanical equipment or structure containing the core components malfunctions, that is, it acquires fault data when each type of fault occurs. In this embodiment, the possible fault types of the six core components are: motor failure, loose screw failure, and misalignment of the guide rod.
[0047] In this embodiment, an active power measuring instrument is installed at the motor to measure the motor's output power in real time, and the ratio of the output power to the rated power is recorded as the degree of motor failure.
[0048] It should be noted that the rated power of a motor refers to the maximum power that the motor can output under standard conditions during long-term continuous operation. The rated power is a fixed data known at the time the motor leaves the factory. The smaller the degree of fault, the more likely the motor is to fail, resulting in a reduction in the motor's output power.
[0049] A magnetic induction switch probe is installed at the nut to collect data on the degree of screw looseness; an angle sensor is installed at the rod position to collect data on the tilt angle of the rod, in order to determine whether the rod is off-center, thus indicating the degree of misalignment.
[0050] It should be noted that all data are collected at the same frequency and are collected synchronously. In this embodiment, the collection frequency is 5Hz and the collection period is 1 year. As for other implementation methods, implementers can set it according to their actual situation.
[0051] Thus, fault data for each fault type at each time point, as well as monitoring data for each fault-affecting parameter at each time point, are obtained;
[0052] In this embodiment, the real-time monitoring data for each fault-affecting parameter includes: vibration data of the reducer bearing, vibration data of the reducer end, vibration data of the motor bearing, temperature data of the motor, temperature data of the reducer roller lubricating oil, and displacement data of the cantilever bracket; the real-time fault data for the corresponding fault type includes: the degree of motor fault, the degree of screw loosening, and the degree of misalignment of the guide rod.
[0053] Thus, we obtain the fault data for each fault type at each time point, as well as the monitoring data for each fault-affecting parameter at each time point.
[0054] Step 2: Determine the main influencing parameters corresponding to each fault type by using the correlation coefficient between the fault data of each fault type and the monitoring data of each fault influencing parameter; construct a multiple linear regression model based on the fault data of each fault type at each time point and the monitoring data of all main influencing parameters corresponding to each fault type at each time point to obtain the regression equation and its influence map corresponding to each fault type; analyze the distribution of data points and the proportion of abnormal data points in the influence map to determine the abnormal proportion of the regression equation corresponding to each fault type; determine the fitting influence weight of the regression equation corresponding to each fault type by using the deviation degree and fitting error of the data points in the influence map; and combine the abnormal proportion and the fitting influence weight to determine the external influence coefficient of the regression equation corresponding to each fault type.
[0055] Based on the above analysis, each fault-affecting parameter has a different degree of influence on each fault type. To clarify the impact of different fault-affecting parameters on each fault type, correlation analysis is conducted to determine the main influencing factors corresponding to each fault type, specifically:
[0056] Calculate the Pearson correlation coefficient between the fault data for each fault type at all times and the monitoring data for each fault impact parameter at all times;
[0057] All fault-related parameters whose absolute values of the Pearson correlation coefficients are greater than a preset threshold are taken as all major-related parameters for each fault type.
[0058] Preferably, in this embodiment, the preset threshold value is 0.5. As for other implementation methods, the implementer can set it according to the actual situation.
[0059] Preferably, in this embodiment, the Pearson correlation coefficients between three types of faults, namely motor faults, loose screw faults, and misaligned guide rod faults, and the vibration data of the reducer bearing, the vibration data of the reducer end, the vibration data of the motor bearing, the temperature data of the motor, the temperature data of the reducer roller lubricating oil, and the displacement data of the cantilever bracket are calculated respectively. The calculation of the Pearson correlation coefficient is a well-known technique and will not be described in detail here.
[0060] In this embodiment, the main fault-affecting parameters are the vibration V1 of the motor bearing, the temperature T1 of the motor, the temperature T2 of the lubricating oil, the vibration V2 of the reducer bearing, the vibration V3 of the reducer end, and the displacement W1 of the cantilever seat; the correlation analysis between the three fault types and the fault-affecting parameters is shown in Table 1.
[0061] Table 1
[0062]
[0063] Table 1 shows that the absolute values of the Pearson correlation coefficients between motor faults and motor bearing vibration V1, motor temperature T1, and lubricating oil temperature T2 are greater than 0.5, indicating that motor faults are correlated with these three factors, meaning that motor faults are mainly affected by them. Similarly, the absolute values of the Pearson correlation coefficients between loose screws and motor bearing vibration V1, reducer bearing vibration V2, and reducer end vibration V3 are also greater than 0.5, indicating that loose screws are not directly related to these factors. The vibrations V2 and V3 at the reducer end are correlated, meaning that loose screws are mainly affected by the vibrations V1 of the motor bearing, V2 of the reducer bearing, and V3 of the reducer end. The absolute value of the Pearson correlation coefficient between the misalignment fault of the polished rod and the vibrations V2 and V3 of the reducer bearing, as well as the displacement W1 of the cantilever bracket, is greater than 0.5, indicating that the misalignment fault of the polished rod is correlated with these vibrations. Therefore, the vibrations V1 of the motor bearing, the motor temperature T1, and the lubricating oil temperature T2 are the main influencing parameters for motor faults; the vibrations V1 of the motor bearing, V2 of the reducer bearing, and V3 of the reducer end are the main influencing parameters for loose screws; and the vibrations V2 of the reducer bearing, V3 of the reducer end, and the displacement W1 of the cantilever bracket are the main influencing parameters for misalignment of the polished rod.
[0064] Furthermore, based on each fault type and its corresponding main influencing parameters, a multiple linear regression model is established, specifically as follows:
[0065] Based on the fault data of each fault type at each time point and the monitoring data of each major influencing parameter at each time point, a multiple linear regression model is constructed to obtain the regression equation and influence map of each fault type.
[0066] Preferably, in this embodiment, the least squares method is used to construct a multiple linear regression model. The least squares method for constructing a multiple linear regression model is a well-known technique and will not be described in detail here. The influence plot of the regression equation for each fault type is output through R language. The acquisition of the influence plot is a well-known technique and will not be described in detail here. It should be noted that the horizontal axis of the influence plot is the leverage value, which reflects the degree of influence of each data point on the regression model, and the vertical axis is the studentized residual, which measures the degree of fit of each data point.
[0067] In this embodiment, the regression equation for the motor fault is: Y D =0.00023V1+0.00014T1+0.00031T2+0.2896; The regression equation for the screw loosening fault is: Y L =0.00048V1+0.00025V2+0.00039V3+0.1433; The regression equation for the misalignment fault of the polished rod is: Y G =0.00063V2+0.00034V3+0.00089W1-0.8617.
[0068] In the above regression equation, the model is mainly fitted by selecting one year of monitoring data as the sample. The fitting result cannot reflect the actual overall change, which limits the generalization ability of the model. Secondly, the model assumes that there is a linear relationship between the dependent and independent variables, but the actual situation may be more complex. This assumption bias will affect the accuracy of the model. Outliers or outliers will have a significant impact on the fitted regression, making the model oversensitive to these points. In addition, the multicollinearity problem among independent variables may also lead to unstable model parameter estimation, affecting the explanatory power of the model.
[0069] In fitting each fault type and its highly correlated main influencing parameters as described above, low-correlation influencing factors were ignored. Although low-correlation influencing factors contribute little to the model, they still contain information about the equipment status. These influencing factors may affect the fault under specific conditions or when interacting with other influencing factors. Therefore, completely ignoring them in the model may lead to a decrease in the model's generalization ability, especially affecting the accuracy of fault judgment under boundary conditions or extreme cases. Therefore, in order to reflect the model's generalization ability and clarify the sensitivity of the regression model to changes in monitoring data of other external influencing factors, the external influence coefficients are determined. The flowchart of the method for obtaining the external influence coefficients of the regression equation corresponding to each fault type provided in this application embodiment is as follows: Figure 2 As shown, it specifically includes:
[0070] The number of outliers and high-leverage points in the influence graph is counted and recorded as the number of outliers.
[0071] Preferably, in this embodiment, observation points with a studentized residual greater than 1.5 or less than -1.5 in the influence graph of each fault type are identified as outliers, and observation points with a leverage value greater than 0.2 are identified as high leverage value points.
[0072] The total number of all data points in the influence diagram is recorded as the total number; the ratio of the number of outliers to the total number is used as the outlier weight of the regression equation corresponding to each fault type.
[0073] It should be noted that one data point in the influence diagram corresponds to the monitoring data of all major influencing parameters collected at one time point; therefore, each time point corresponds to one data point in the influence diagram.
[0074] The leverage values and studentized residuals of all data points in the influence diagram are normalized respectively. The mean of the product of the normalized leverage values and the normalized studentized residuals of all data points is used as the fitting influence weight of the regression equation corresponding to each fault type.
[0075] Preferably, in this embodiment, the softmax function is used for normalization. The softmax function is a well-known technique and will not be described in detail here. As other implementation methods, implementers can use other methods of the prior art, such as the sigmoid function, etc. This embodiment does not impose any special restrictions on this. Through normalization, the leverage values and studentized residuals affecting all data points in the graph are converted into probabilities, and then the probability of anomalies in the corresponding data points is reflected by fitting the influence weights.
[0076] The product of the abnormality proportion and the fitted influence weight is used as the external influence coefficient of the regression equation corresponding to each fault type.
[0077] It should be noted that the aforementioned anomaly proportion indicates that the more outlier the data points in the influence graph, the greater the threat these outliers pose to the accuracy and stability of the regression model. A higher anomaly proportion means a higher percentage of outlier data points. The normalized leverage value reflects the relative importance of the monitoring data of the main influencing parameters collected at the corresponding time point in the regression model. A higher normalized leverage value indicates a greater impact of the monitoring data of the main influencing parameters collected at the corresponding time point on the model's regression fit, resulting in a larger external influence coefficient. This reflects a higher dependence of the regression model on that data point. The studentized residual reflects a standardized measure of the difference between the predicted and actual values of the regression model. The larger the normalized studentized residual, the worse the fit of the corresponding data point. The larger the external influence coefficient, the more unstable the regression model is at that data point. Therefore, the fit influence weight reflects the measure of the overall fit quality instability of the regression model. The external influence coefficient reflects the anomalies of the data points, i.e., the impact of outliers and high leverage points on the accuracy of model fault prediction. The larger the external influence coefficient, the worse the model's generalization ability and the stronger its sensitivity to changes in monitoring data of other external influencing factors.
[0078] Thus, the external influence coefficients of the regression equations corresponding to each fault type are obtained.
[0079] Step 3: Based on the correlation between different main influencing parameters corresponding to each fault type and their degree of multicollinearity, and in conjunction with the external influence coefficient, determine the fitting inaccuracy of the regression equation corresponding to each fault type.
[0080] Furthermore, in the tower well operation monitoring and early warning system, the collinearity problem among the main influencing parameters can also affect the fitting effect. The collinearity problem is mainly caused by the inherent interrelationship between the monitored variables. That is, the vibration of the motor bearing V1, the temperature of the motor T1, the temperature of the lubricating oil T2, the vibration of the reducer bearing V2, the vibration of the reducer end V3, and the displacement of the cantilever seat W1 may affect each other due to the physical structure and working principle of the mechanical equipment. For example, the increase in motor temperature may simultaneously lead to the increase in lubricating oil temperature, and the vibration of the motor bearing may be related to the vibration of the reducer bearing. This high correlation makes the independent variables in the model not completely independent, thus causing the collinearity problem.
[0081] Because the correlation between the main influencing parameters increases the risk of collinearity among independent variables in the regression model, it affects the stability and predictive accuracy of the regression model. Collinearity leads to unstable parameter estimates in the regression model; that is, small changes in data can cause large fluctuations in parameter estimates, which reduces the interpretability of the regression equation. Collinearity may also mask the true impact of some independent variables. Therefore, this study analyzes the collinearity among the different main influencing parameters corresponding to each fault type to determine the fit inaccuracy, reflecting the predictive performance of the regression equation. Specifically:
[0082] The mean of the absolute values of the Pearson correlation coefficients between the monitoring data of any two major influencing parameters for each fault type at all times is taken as the average correlation for each fault type.
[0083] The mean of the variance inflation factor of all major influencing parameters corresponding to each fault type is used as the average collinearity of each fault type.
[0084] The method for calculating the fitting inaccuracy of the regression equation corresponding to each fault type is as follows: Among them, B q A represents the fitting inaccuracy of the regression equation corresponding to the q-th fault type. q Let μR be the external influence coefficient of the regression equation corresponding to the q-th fault type. q Let be the average correlation coefficient corresponding to the q-th fault type. Let q be the average collinearity corresponding to the qth fault type, max[] be the maximum value function, and ε be a preset value greater than 0 to avoid the result after mapping by the max[] function being 0. In this embodiment, ε is 0.01. As for other implementation methods, the implementer can set it according to the actual situation.
[0085] It should be noted that the calculation of the variance inflation factor is a well-known technique and will not be elaborated upon here. The μR mentioned... q The larger the value, the more correlated the different main influencing parameters corresponding to each fault type are; that is, there is collinearity among different independent variables in the corresponding regression equation, which has a greater impact on the stability of the fitted regression model. The larger the variance inflation factor (VIF), the more significant the multicollinearity among the independent variables in the regression equation. Generally, a VIF greater than 10 indicates a serious multicollinearity problem, which has a greater impact on the stability of parameter estimation in the fitted model. The VIF function can be used to further investigate this. Make corrections so that when When the value exceeds 10, a larger weight is assigned to the regression equation corresponding to the fault type. When the value is less than 10, a smaller weight is assigned to the regression equation corresponding to the fault type; the external influence coefficient A q The larger the value, the more abnormal data there are in the fitted monitoring data, which has a greater impact on the model's prediction accuracy. These abnormal data will distort the model's prediction results and reduce the accuracy and stability of the fitted model. The larger the obtained fitting inaccuracy, the less reliable the corresponding regression equation is for fault prediction, and the lower the prediction accuracy and stability of the regression equation. Conversely, the smaller the fitting inaccuracy, the better the fitting effect of the regression equation, and the higher the prediction accuracy and stability.
[0086] Thus, the fitting inaccuracy of the regression equation corresponding to each fault type is obtained.
[0087] Step 4: Based on the fitting inaccuracy, the regression equation is corrected to obtain the adjusted regression equation corresponding to each fault type, and the fault type of the oil well is predicted and warned in real time.
[0088] Furthermore, based on the aforementioned fitting inaccuracy, the regression equation is corrected to obtain the adjusted regression equation, specifically as follows:
[0089] The adjusted regression equation for each fault type is: Y q ′ =exp(-B q )×Y q , where Y q ′ Let B be the adjusted regression equation corresponding to the q-th fault type. q Y represents the fitting inaccuracy of the regression equation corresponding to the q-th fault type. q Let q be the regression equation corresponding to the q-th fault type, and exp() be an exponential function with the natural constant as the base.
[0090] It should be noted that the aforementioned fitting inaccuracy B q The larger the value, the worse the fit. In this case, the prediction result needs to be given less weight to reduce the prediction error and make the predicted value closer to the true value.
[0091] Furthermore, based on the adjusted regression equation, the severity of each fault type is predicted in real time, and real-time warnings are issued according to pre-set safety thresholds, specifically:
[0092] Based on the monitoring data of each fault impact parameter at the current moment, and combined with the adjusted regression equation corresponding to each fault type, the predicted fault data of each fault type at the current moment is calculated.
[0093] If the predicted fault data is greater than or equal to the preset safety threshold for the corresponding fault type, the oil well has a fault, and a warning is issued via a fault indicator light; otherwise, the oil well has no fault.
[0094] Preferably, in this embodiment, the predicted fault data Y of the motor fault at the current moment is... D ′ If the value is greater than or equal to the preset safety threshold of 1.3, a motor failure occurs in the oil well, and the corresponding motor failure indicator light will display a red warning light to issue a warning. D ′ The value is less than the preset safety threshold of 1.3, indicating no motor failure in the oil well, and the fault indicator light shows a green safety light; the predicted fault data Y for the loose screw fault at the current moment is... L ′ If the value is greater than or equal to the preset safety threshold of 1.8, a loose screw fault is detected in the oil well. The corresponding fault indicator light will display a red warning light to issue a warning. L ′ The value is less than the preset safety threshold of 1.8, indicating no loose screw fault in the oil well, and the corresponding fault indicator light displays a green safety light; the predicted fault data Y for the polished rod misalignment fault at the current moment is... G ′ If the oil well experiences a polished rod misalignment fault when the value is greater than or equal to the preset safety threshold of 2, the corresponding fault indicator light will display a red warning light to issue a warning. G ′ If the value is less than the preset safety threshold of 2, and the oil well does not have a polished rod misalignment fault, the corresponding fault indicator light will display a green safety light. As another implementation method, the implementer can set it according to the actual situation.
[0095] Therefore, during normal operation, the green safety lights for the three fault types remain on, indicating that the system is in normal working condition. When the predicted fault data for a certain fault type exceeds the preset safety threshold for the corresponding fault type, its fault indicator light changes from green to red. At this time, the on-site inspection personnel will conduct a shutdown inspection of the fault point for the corresponding fault type.
[0096] It should be understood that, although Figure 1 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 1At least some of the steps in the process may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least some of the sub-steps or stages of other steps.
[0097] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0098] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application. Therefore, any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of this application, without departing from the content of the technical solution of this application, shall fall within the protection scope of the technical solution of this application.
Claims
1. A monitoring and analysis method for tower machine wells based on Pearson correlation coefficient, characterized in that, The method includes the following steps: Real-time acquisition of fault data for each type of fault at each moment during the operation of the tower oil well, as well as monitoring data for each fault-affecting parameter at each moment; By using the correlation coefficient between the fault data of each fault type and the monitoring data of each fault impact parameter, the main impact parameters corresponding to each fault type are determined. Based on the fault data of each fault type at each time point, and the monitoring data of all major influencing parameters corresponding to each fault type at each time point, a multiple linear regression model is constructed to obtain the regression equation and its influence diagram for each fault type. Analyze the distribution of data points in the influence graph and the proportion of abnormal data points to determine the abnormal weight of the regression equation corresponding to each fault type; determine the fitting influence weight of the regression equation corresponding to each fault type by the deviation of the data points in the influence graph and their fitting error; combine the abnormal weight and the fitting influence weight to determine the external influence coefficient of the regression equation corresponding to each fault type. Based on the correlation and multicollinearity between the different main influencing parameters corresponding to each fault type, and in conjunction with the external influence coefficient, the fitting inaccuracy of the regression equation corresponding to each fault type is determined. The regression equation is corrected based on the aforementioned fitting inaccuracy to obtain the adjusted regression equation corresponding to each type of failure, and the failure type of the oil well is predicted and warned in real time. The determination of the fitting inaccuracy of the regression equation corresponding to each fault type includes: Analyze the average level of correlation among all major influencing parameters for each type of fault, and calculate the average correlation. Analyze the average level of collinearity among all major influencing parameters corresponding to each fault type, and calculate the average collinearity. No. The fitting inaccuracy of the regression equation corresponding to each fault type The calculation formula is: ,in, For the first The external influence coefficients of the regression equations corresponding to the various fault types. For the first The average correlation of the various fault types For the first The average collinearity corresponding to each fault type To find the maximum value function, The default value is greater than 0.
2. The tower well monitoring and analysis method based on Pearson correlation coefficient as described in claim 1, characterized in that, Each of the fault types includes: motor fault, loose screw fault, and misalignment of the guide rod fault.
3. The tower well monitoring and analysis method based on Pearson correlation coefficient as described in claim 1, characterized in that, Each fault-affecting parameter includes: vibration of the reducer bearing, vibration of the reducer end, vibration of the motor bearing, motor temperature, temperature of the reducer drum lubricating oil, and displacement of the cantilever support.
4. The tower well monitoring and analysis method based on Pearson correlation coefficient as described in claim 1, characterized in that, The correlation coefficient is the Pearson correlation coefficient calculated between the fault data of each fault type at all times and the monitoring data of each fault influence parameter at all times.
5. The tower well monitoring and analysis method based on Pearson correlation coefficient as described in claim 1, characterized in that, The determination of the main influencing parameters corresponding to each fault type includes: All fault-affecting parameters whose absolute values of the correlation coefficients are greater than a preset threshold are taken as all major-affecting parameters corresponding to each fault type.
6. The tower well monitoring and analysis method based on Pearson correlation coefficient as described in claim 1, characterized in that, The determination of the anomaly weight of the regression equation corresponding to each fault type includes: The number of outliers and high-leverage points in the influence graph is counted and recorded as the number of outliers; the total number of all data points in the influence graph is counted and recorded as the total number. The ratio of the number of outliers to the total number is used as the outlier weight in the regression equation corresponding to each fault type.
7. The tower well monitoring and analysis method based on Pearson correlation coefficient as described in claim 1, characterized in that, The determination of the fitting influence weights of the regression equation corresponding to each fault type includes: The mean of the product of the normalized leverage value and the normalized studentized residual of all data points in the influence diagram is used as the fitting influence weight of the regression equation corresponding to each fault type.
8. The tower well monitoring and analysis method based on Pearson correlation coefficient as described in claim 1, characterized in that, The external influence coefficient of the regression equation corresponding to each fault type is the product of the anomaly proportion and the fitting influence weight.
9. The tower well monitoring and analysis method based on Pearson correlation coefficient as described in claim 1, characterized in that, The average correlation is the mean of the absolute values of the Pearson correlation coefficients between the monitoring data of all two major influencing parameters corresponding to each fault type at all times.
10. The tower well monitoring and analysis method based on Pearson correlation coefficient as described in claim 1, characterized in that, The average collinearity is the mean of the variance inflation factor of all major influencing parameters corresponding to each fault type.
11. The method for monitoring and analyzing tower machine wells based on Pearson correlation coefficient as described in claim 1, characterized in that, The adjusted regression equation for each fault type is as follows: ,in, For the first The adjusted regression equations corresponding to the various fault types are as follows: For the first The fitting inaccuracy of the regression equations corresponding to different fault types. For the first The regression equations corresponding to the various fault types, It is an exponential function with the natural constant as the base.
12. The tower well monitoring and analysis method based on Pearson correlation coefficient as described in claim 1, characterized in that, The real-time prediction and early warning of fault types in oil wells includes: Based on the monitoring data of all major influencing parameters corresponding to each fault type at the current moment, and combined with the adjusted regression equation, the predicted fault data for each fault type at the current moment is calculated. If the predicted fault data is greater than or equal to the preset safety threshold for the corresponding fault type, the oil well has a fault; otherwise, the oil well has no fault.
Citation Information
Patent Citations
Oil-water well casing damage early warning method and device and storage medium
CN111476406A