A pipeline status prediction method based on big data
By acquiring multi-source factor information of buried natural gas pipelines and constructing a decision tree to optimize information entropy, the problem of noise data influence is solved, and more accurate pipeline status prediction and timely abnormality monitoring are achieved.
Patent Information
- Application Number
- CN202510201913.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-02-24
AI Technical Summary
Existing technologies cannot effectively eliminate the impact of noise data in the real-time dynamic data of buried natural gas pipelines, resulting in distortion of pipeline status prediction models, affecting the effectiveness of predictions and the timeliness of abnormality monitoring.
By obtaining multi-source factor information of the transmission pipeline, determining the credibility of the data source and the information consistency, building a decision tree, optimizing the information entropy, and predicting the pipeline pressure range based on the decision tree, and comparing it with the actual value to determine whether the pipeline status is abnormal, the influence of noise data is eliminated.
It improves the effectiveness of pipeline status prediction and the timeliness of abnormality monitoring, avoids data distortion of the prediction model, and enhances the accuracy of pipeline status monitoring.
Smart Images

Figure CN119989235B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing and analysis, and in particular to a pipeline state prediction method based on big data. Background Art
[0002] In the field of modern industry and energy transportation, pipeline transportation plays a vital role and is widely used in the transportation of fluids such as oil and natural gas. Traditional pipeline status monitoring methods mainly rely on manual inspections and simple sensor monitoring. This method has many limitations. On the one hand, manual inspections are inefficient and costly, and it is difficult to monitor the operating status of pipelines in real time and comprehensively. On the other hand, simple sensor monitoring can only obtain limited single data and cannot comprehensively analyze the impact of multiple factors on pipeline status. With the increasing complexity and expansion of pipeline systems, these traditional methods have become difficult to meet the needs of accurate prediction and timely warning of pipeline status.
[0003] With the rapid development of big data technology, big data technology can efficiently collect, store, process and analyze massive, multi-source data. By obtaining multi-source factor information in the pipeline transportation process and using big data analysis technology to deeply mine and analyze this data, it can more accurately predict the pipeline status and detect potential faults in a timely manner, thereby ensuring the safe and reliable operation of the pipeline. The use of big data for predictive analysis of pipeline status has important practical significance and application value.
[0004] For example, Chinese patent publication number: CN114722662A, the invention discloses a method for online monitoring of foundation settlement and safety research of buried natural gas pipelines, including: establishing a three-dimensional pipe-soil model and determining key points for settlement monitoring; a signal relay system transmits data collected by the settlement measurement sensor system to a remote terminal server to achieve long-term online settlement monitoring; based on the measured big data of pipeline settlement, the settlement status of the next stage is predicted based on a prediction model and a BP neural network model; by performing Fourier series expansion on the settlement data, the pipeline harmonic settlement curve is obtained, which is input into the three-dimensional model as a loading condition, and combined with the pipeline force, the impact of the current settlement on the pipeline structure is analyzed.
[0005] The following problems also exist in the prior art:
[0006] Existing technologies cannot eliminate the influence of noise data from the real-time dynamic data of buried natural gas pipelines, resulting in data distortion in the constructed prediction model, affecting the effectiveness of pipeline status prediction and the timeliness of abnormality monitoring. Summary of the Invention
[0007] To this end, the present invention provides a pipeline status prediction method based on big data to overcome the problem that the existing technology cannot eliminate the influence of noise data from the acquired real-time dynamic data of the buried natural gas pipeline.
[0008] To achieve the above objectives, the present invention provides a pipeline status prediction method based on big data, comprising:
[0009] Obtaining multi-source factor information of the monitored point of the transmission pipeline at each monitoring time, the multi-source factor information including gas input amount, gas flow rate, branch pipeline flow rate and branch pipeline flow rate of the associated branch area, and exogenous environmental parameters;
[0010] At each monitoring moment, the first data source credibility is determined based on the fluctuation of the external environmental parameters corresponding to the current monitoring point in the time dimension. The second data source credibility is determined based on the first correlation between the gas input volume and the gas flow rate at the current monitoring point in the time dimension and the second correlation between the branch pipeline flow rate and the flow rate in the associated branch area in the order of the branch pipeline diameter arrangement;
[0011] Determining the information consistency at each monitoring moment based on the first data source credibility and the second data source credibility;
[0012] Based on the numerical values of each information dimension, the multi-source factor information is divided into several subcategories. According to the information fit of the multi-source factor information of each subcategory, the initial information entropy of the multi-source factor information of each subcategory is optimized, and a decision tree is constructed based on the information gain of the optimized information entropy.
[0013] Wherein, the leaf node of the decision tree is the pipeline pressure of the current point to be monitored;
[0014] The pipeline pressure prediction interval of the current monitoring point is determined according to the decision tree, and whether the pipeline state is abnormal is determined according to the comparison between the pipeline pressure prediction interval and the actual value of the pipeline pressure.
[0015] Furthermore, the associated branch area is an area in the upstream pipeline area of the point to be monitored where a pipeline branch node closest to the point to be monitored is located.
[0016] Furthermore, the process of determining the fluctuation of the exogenous environmental parameters corresponding to the current monitoring point in the time dimension includes:
[0017] Based on the time dimension sequence, obtain the exogenous environmental parameters at the current monitoring time and other monitoring times within the preset monitoring neighborhood;
[0018] Calculate the difference between the current monitoring time and the exogenous environmental parameters at other monitoring times within the preset monitoring neighborhood.
[0019] Furthermore, the credibility of the first number source is determined according to the difference, and the credibility of the first number source is negatively correlated with the difference.
[0020] Furthermore, the process of determining the first change relevance includes:
[0021] Obtain the gas flow rate of the current monitoring point at each monitoring moment and the gas input volume of the pipeline where the current monitoring point is located;
[0022] Determine the gas flow rate and gas flow rate at several monitoring moments within a preset monitoring neighborhood with the current monitoring moment as the time center;
[0023] The Pearson correlation coefficient between the gas flow rate and the gas input amount at a plurality of monitoring moments within a preset monitoring neighborhood is determined as the first change correlation.
[0024] Furthermore, the process of determining the second change relevance includes:
[0025] Obtaining the associated branch area corresponding to the current monitoring point, determining the diameters of each branch pipeline in the associated branch area, and sorting the diameters of each branch pipeline;
[0026] The gas flow rate and gas flow rate of each branch pipe are determined, the Pearson correlation coefficient of the gas flow rate and gas flow rate corresponding to branch pipes of several branch pipe diameters is calculated, and the Pearson correlation coefficient is determined as the second change correlation.
[0027] Furthermore, the second data source credibility is determined based on the first change correlation and the second change correlation;
[0028] The second data source credibility is positively correlated with the first change correlation and the second change correlation respectively.
[0029] Furthermore, the information consistency is determined based on the credibility of the first data source and the credibility of the second data source;
[0030] The information consistency is obtained by normalizing the product of the first data source credibility and the second data source credibility.
[0031] Furthermore, the process of optimizing the initial information entropy includes:
[0032] The multi-source factor information is divided into several subcategories according to the numerical value of each information dimension;
[0033] Calculate the ratio of the sum of the information fit corresponding to each multi-source factor information in each subcategory to the sum of the information fit of all subcategories;
[0034] The ratio is substituted into the information entropy calculation formula to optimize the initial information entropy.
[0035] Furthermore, the process of determining whether the pipeline status is abnormal includes:
[0036] Comparing the pipeline pressure prediction interval with the pipeline pressure actual value;
[0037] If the actual value of the pipeline pressure does not meet the normal standard conditions, it is determined that the pipeline state is abnormal;
[0038] The normal standard condition is that the actual value of the pipeline pressure falls within the pipeline pressure prediction interval.
[0039] Compared with the prior art, the beneficial effect of the present invention lies in that, by obtaining multi-source factor information of the monitored point of the transmission pipeline at each monitoring moment, the present invention determines the first source credibility according to the external environmental parameters corresponding to the current monitoring point at each monitoring moment, calculates the second source credibility according to the first change correlation determined by the gas flow rate and the gas flow of the current monitoring point and the second change correlation determined by the branch pipeline flow and the flow rate of the associated branch area, determines the information fit at each monitoring moment according to the first source credibility and the second source credibility, optimizes the initial information entropy, constructs a decision tree according to the information gain of the optimized information entropy, determines the pipeline pressure prediction interval of the current monitored point according to the decision tree, and determines whether the pipeline state is abnormal, thereby eliminating the influence of noise data, avoiding data distortion of the constructed prediction model, and improving the effectiveness of pipeline state prediction and the timeliness of abnormality monitoring.
[0040] Furthermore, the present invention determines the credibility of the first data source through the exogenous environmental parameters at the current monitoring moment and other monitoring moments within the preset monitoring neighborhood, obtains the exogenous environmental parameters at multiple monitoring moments and calculates the difference. This method can capture the changes in parameters over time. It can be understood that if the exogenous environmental parameters fluctuate greatly in a short period of time, it may interfere with the monitoring data of the transmission pipeline. The degree of interference is determined by analyzing the size of the difference. The larger the difference, the lower the data credibility. In turn, more reliable data sources are screened out, the influence of noise data is eliminated, and data distortion of the constructed prediction model is avoided.
[0041] Furthermore, the present invention determines the first change correlation by calculating the Pearson correlation coefficient between the gas input and the gas flow rate within a preset monitoring neighborhood. It can be understood that the Pearson correlation coefficient is an effective indicator for measuring the degree of linear correlation between two variables. Calculating the Pearson correlation coefficient can help determine whether the data change trends of the gas flow rate and the gas input at multiple monitoring moments are consistent. Under normal circumstances, the correlation coefficient between the gas flow rate and the gas input should be within a relatively stable range. When the pipeline is in an abnormal state, this correlation may undergo unstable fluctuations, thereby eliminating the influence of noise data, avoiding data distortion of the constructed prediction model, and improving the effectiveness of pipeline status prediction.
[0042] Furthermore, the present invention determines the second change correlation through the Pearson correlation coefficient of the gas flow rate and gas flow corresponding to branch pipes of several branch pipe diameters, quantifies the linear relationship between the gas flow rate and flow of the branch pipes, and sorts the branch pipe diameters in the associated branch area. On this orderly basis, the correlation between the flow rate and flow is studied to reflect the gas flow synergy of the entire pipeline network under branch pipes of different diameters. It can be understood that pipeline state fluctuations may cause the flow of some branch pipes to decrease, thereby affecting the flow rate, causing the Pearson correlation coefficient of the flow rate and flow to deviate from the normal range. Furthermore, the influence of noise data is eliminated, data distortion of the constructed prediction model is avoided, and the effectiveness of pipeline state prediction is improved.
[0043] Furthermore, the present invention determines whether the pipeline state is abnormal by comparing the pipeline pressure prediction interval and the actual value. The pipeline pressure prediction interval determined by the decision tree is obtained by integrating multi-source factor information. The comparison of the predicted value and the actual value after comprehensive consideration of multiple factors can effectively avoid misjudgment caused by single factor judgment. During the construction process of the decision tree, the initial information entropy of the multi-source factor information is optimized, and factors such as the information fit between each factor are considered, so that the predicted value can better reflect the actual situation. When the difference between the predicted value and the actual value is large, it indicates that the pipeline pressure does not follow the natural gas transmission characteristic data relationship under normal conditions, which may be due to the data relationship abnormality caused by state fluctuations. In addition, the influence of noise data is eliminated, data distortion of the constructed prediction model is avoided, and the effectiveness of pipeline state prediction and the timeliness of abnormality monitoring are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 This is a step diagram of a pipeline status prediction method based on big data according to an embodiment of the present invention;
[0045] Figure 2 A diagram showing the steps for determining a first change relevance according to an embodiment of the present invention;
[0046] Figure 3A diagram showing the steps of determining the second change relevance according to an embodiment of the present invention;
[0047] Figure 4 A diagram showing the steps for optimizing initial information entropy according to an embodiment of the present invention;
[0048] Figure 5 This is a logic flow chart for determining whether a pipeline status is abnormal according to an embodiment of the present invention. DETAILED DESCRIPTION
[0049] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.
[0050] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0051] It should be noted that, in the description of the present invention, terms such as "upper", "lower", "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.
[0052] See also Figure 1 As shown, it is a step diagram of the pipeline state prediction method based on big data according to an embodiment of the present invention. The pipeline state prediction method based on big data according to this embodiment includes:
[0053] Step S100, obtaining multi-source factor information of a monitored point of a transmission pipeline at each monitoring moment, wherein the multi-source factor information includes gas input amount, gas flow rate, branch pipeline flow rate and branch pipeline flow rate of an associated branch area, and external environmental parameters;
[0054] In implementation, the time interval of each monitoring moment is 3s-10s, preferably, it can be set to 5s, the gas input amount of the monitored point is the gas input amount of the pipeline where the current monitoring point is located within a time period of 3s, the gas flow rate is the gas flow rate at the location of the monitored point at the monitoring moment, and the branch pipeline flow rate is the gas flow rate of the branch pipeline within a preset time length with the monitoring moment as the starting moment, wherein the preset time length can be set by technical personnel in this field according to the monitoring accuracy requirements, and can be set to 3s, and the branch pipeline flow rate is the gas flow rate at the location of the branch pipeline node at the monitoring moment.
[0055] Among them, the gas input amount, gas flow rate, branch pipe flow rate in the associated branch area, flow rate, and exogenous environmental parameters are dimensionless when participating in the calculation.
[0056] Step S200: At each monitoring moment, a first data source credibility is determined based on the fluctuation of the exogenous environmental parameter corresponding to the current monitoring point in the time dimension, and a second data source credibility is determined based on a first correlation between the gas input volume and the gas flow rate at the current monitoring point in the time dimension, and a second correlation between the branch pipeline flow rate and the flow rate in the associated branch region in the order of the branch pipeline diameters.
[0057] Step S300: determining the information consistency at each monitoring moment according to the first data source credibility and the second data source credibility;
[0058] Step S400: Divide the multi-source factor information into several subcategories based on the numerical values of each information dimension, optimize the initial information entropy of the multi-source factor information of each subcategory according to the information fit of the multi-source factor information of each subcategory, and construct a decision tree based on the information gain of the optimized information entropy;
[0059] Wherein, the leaf node of the decision tree is the pipeline pressure of the current point to be monitored;
[0060] Step S500: determining a pipeline pressure prediction interval of the current monitoring point according to the decision tree, and determining whether the pipeline state is abnormal according to a comparison between the pipeline pressure prediction interval and the actual pipeline pressure value.
[0061] Specifically, the decision tree reflects the dynamic characteristics of the buried natural gas pipeline during gas transmission through the path relationship between nodes. In the process of predicting the pressure characteristics of the real-time transmission process, representative data collected during the gas transmission process can be substituted into the decision tree. By determining the path of the decision tree, the pipeline pressure prediction interval of the monitoring point can be obtained according to the leaf node corresponding to the path. Technical personnel in this field can judge whether the pipeline pressure at the current location of the monitoring point is normal and whether there is any abnormal pipeline pressure based on the pipeline pressure prediction interval.
[0062] Specifically, during the transportation of natural gas, there are many factors that affect the pipeline pressure. In an embodiment of the present invention, the gas input amount, gas flow rate, branch pipeline flow and flow rate of the associated branch area, and exogenous environmental parameters are obtained. Of course, for those skilled in the art, any other environmental parameters suitable for the judgment method of this embodiment can be used as the basis for the parameters selected in this embodiment.
[0063] Specifically, the associated branch area is an area in the upstream pipeline area of the point to be monitored where a pipeline branch node closest to the point to be monitored is located.
[0064] Specifically, the process of determining the fluctuation of the exogenous environmental parameters corresponding to the current monitoring point in the time dimension includes:
[0065] Based on the time dimension sequence, obtain the exogenous environmental parameters at the current monitoring time and other monitoring times within the preset monitoring neighborhood;
[0066] Calculate the difference between the current monitoring time and the exogenous environmental parameters at other monitoring times within the preset monitoring neighborhood.
[0067] During implementation, the preset monitoring neighborhood range can be set to 6, with the current monitoring moment as the moment center and the surrounding 6 other monitoring moments constituting the preset monitoring neighborhood range of the monitoring moment.
[0068] Specifically, exogenous environmental parameters include surface temperature. For the environmental parameters of buried natural gas pipelines, when the surface temperature changes, the natural gas temperature changes accordingly, the volume shrinks or expands, and the pipeline pressure fluctuates. If the surface temperature changes significantly in the time dimension, indicating that the pressure in the buried natural gas pipeline is significantly disturbed at this time, the credibility of the first source determined based on the exogenous environmental parameters is low.
[0069] Specifically, the first number source credibility is determined according to the difference, and the first number source credibility is negatively correlated with the difference.
[0070] Specifically, the present invention determines the credibility of the first data source through the exogenous environmental parameters at the current monitoring moment and other monitoring moments within a preset monitoring neighborhood, obtains the exogenous environmental parameters at multiple monitoring moments and calculates the difference. This method can capture the changes in parameters over time. It can be understood that if the exogenous environmental parameters fluctuate greatly in a short period of time, it may interfere with the monitoring data of the transmission pipeline. The degree of interference is determined by analyzing the size of the difference. The larger the difference, the lower the data credibility. In turn, more reliable data sources are screened out, the influence of noise data is eliminated, and data distortion of the constructed prediction model is avoided.
[0071] Specifically, see Figure 2 As shown in FIG. , which is a diagram showing the steps of determining the first change correlation according to an embodiment of the present invention, the process of determining the first change correlation includes:
[0072] Step S201, obtaining the gas flow rate of the current monitoring point at each monitoring moment and the gas input volume of the pipeline where the current monitoring point is located;
[0073] Step S202, determining the gas flow rate and gas input amount at a plurality of monitoring moments within a preset monitoring neighborhood with the current monitoring moment as the time center;
[0074] Step S203 : determining the Pearson correlation coefficient between the gas flow rate and the gas input amount at a plurality of monitoring moments within a preset monitoring neighborhood as the first change correlation.
[0075] Specifically, the present invention determines the first change correlation by calculating the Pearson correlation coefficient between the gas input amount and the gas flow rate within a preset monitoring neighborhood. It can be understood that calculating the Pearson correlation coefficient can help determine whether the data change trends of the gas flow rate and the gas input amount at multiple monitoring moments are consistent. Under normal circumstances, the correlation coefficient between the gas flow rate and the gas input amount should be in a relatively stable range. When the pipeline state fluctuates, this correlation relationship may undergo unstable fluctuations, thereby eliminating the influence of noise data, avoiding data distortion of the constructed prediction model, and improving the effectiveness of pipeline state prediction.
[0076] Specifically, the Pearson correlation coefficient is an effective indicator for measuring the degree of linear correlation between two variables. Its value range is between -1 and 1. The larger the Pearson correlation coefficient, the higher the correlation between the changes in the two variables. The calculation method of the Pearson correlation coefficient is a technical means well known to those skilled in the art and will not be elaborated here.
[0077] Specifically, see Figure 3 As shown in FIG. , which is a diagram showing the steps of determining the second change correlation according to an embodiment of the present invention, the process of determining the second change correlation includes:
[0078] Step S211: obtaining the associated branch area corresponding to the current monitoring point, determining the diameters of each branch pipeline in the associated branch area, and sorting the diameters of each branch pipeline;
[0079] Step S212, determining the gas flow velocity and gas flow rate of each branch pipe, calculating the Pearson correlation coefficient of the gas flow velocity and gas flow rate corresponding to branch pipes of several branch pipe diameters, and determining the Pearson correlation coefficient as the second change correlation.
[0080] Specifically, the present invention determines the second change correlation through the Pearson correlation coefficient of the gas flow rate and gas flow corresponding to branch pipes of several branch pipe diameters, quantifies the linear relationship between the gas flow rate and flow of the branch pipes, and sorts the diameters of the branch pipes in the associated branch area. On this orderly basis, the correlation between the flow rate and flow is studied to reflect the gas flow synergy of the entire pipeline network under branch pipes of different diameters. It can be understood that unstable pipeline status may lead to a reduction in the flow of some branch pipes, thereby affecting the flow rate, causing the Pearson correlation coefficient of the flow rate and flow to deviate from the normal range. Furthermore, the influence of noise data is eliminated, data distortion of the constructed prediction model is avoided, and the effectiveness of pipeline status prediction is improved.
[0081] Specifically, the change in the gas input amount corresponding to the monitored point will be accompanied by the change in the gas flow rate. Under normal operating conditions, the change in the gas input amount has a certain correlation with the gas flow rate. The larger the gas input amount, the greater the gas flow rate. Similarly, the branch pipeline flow rate and the branch pipeline flow rate in the associated branch area corresponding to the monitored point are also correlated. The larger the branch pipeline flow rate, the greater the branch pipeline flow rate. Therefore, the first change correlation can be determined by calculating the Pearson correlation coefficient between the gas input amount and the gas flow rate within the preset monitoring neighborhood, and the second change correlation can also be determined by the Pearson correlation coefficient of the gas flow rate and the gas flow rate corresponding to the branch pipelines of several branch pipeline diameters.
[0082] Specifically, the second data source credibility is determined according to the first change correlation and the second change correlation;
[0083] The second data source credibility is positively correlated with the first change correlation and the second change correlation respectively.
[0084] Specifically, the information consistency is determined based on the credibility of the first data source and the credibility of the second data source;
[0085] The information consistency is obtained by normalizing the product of the first data source credibility and the second data source credibility.
[0086] In implementation, range normalization can be used for normalization processing. Range normalization is an existing technology and will not be described here in detail. Of course, for those skilled in the art, other normalization processing methods suitable for the normalization processing of this embodiment can be selected as the processing method of this embodiment.
[0087] Specifically, see Figure 4 As shown in FIG, which is a step diagram of optimizing the initial information entropy according to an embodiment of the present invention, the process of optimizing the initial information entropy includes:
[0088] Step S401, classifying the multi-source factor information into several subcategories according to the value of each information dimension;
[0089] Step S402, calculating the ratio of the sum of the information fit corresponding to each multi-source factor information in each subcategory to the sum of the information fit of all subcategories;
[0090] Step S403: Substitute the ratio into the information entropy calculation formula to optimize the initial information entropy.
[0091] In implementation, in the process of dividing the multi-source factor information into several subcategories based on the numerical size of each information dimension, a matrix can be constructed in advance based on the gas input amount, gas flow rate, branch pipeline flow rate of the associated branch area, flow rate, and external environmental parameters in the multi-source factor information. Each column of the matrix corresponds to the data of an information dimension, and each row corresponds to the data of all information dimensions at a monitoring moment. The multi-source factor information is divided into several subcategories according to the numerical interval of the numerical size of a certain information dimension in the multi-source factor information. The initial information entropy is constructed according to the amount of information contained in each subcategory. The information fit is used for optimization based on the initial information entropy. The information gain of the optimized information entropy is used to construct a decision tree. The leaf node of the decision tree is the pipeline pressure of the current monitoring point, and each pipeline pressure node corresponds to a data range.
[0092] For example, the dimensionless values of the gas flow rate in the multi-source factor information at the current monitoring time and the seven monitoring times within the preset monitoring neighborhood are 6.95, 6.81, 6.72, 6.93, 7.07, 6.91, and 7.01, respectively. The above values are divided into four subcategories according to the numerical intervals (6.7, 6.8], (6.8, 6.9], (6.9, 7.0], and (7.0, 7.1]);
[0093] There is 1 data in the numerical interval (6.7, 6.8], 1 data in the numerical interval (6.8, 6.9], 3 data in the numerical interval (6.9, 7.0], and 2 data in the numerical interval (7.0, 7.1].
[0094] The process of constructing the initial information entropy is as follows: let P1 be the probability of the numerical interval (6.7, 6.8], P1 = 1 / 7; let P2 be the probability of the numerical interval (6.8, 6.9], P2 = 1 / 7; let P3 be the probability of the numerical interval (6.9, 7.0], P3 = 3 / 7; let P4 be the probability of the numerical interval (7.0, 7.1], P4 = 2 / 7.
[0095] According to the calculation formula of information entropy Where n = 4;
[0096] Then H(X)=-[1 / 7×log2(1 / 7)+1 / 7×log2(1 / 7)+3 / 7×log2(3 / 7)+2 / 7×log2(2 / 7)];
[0097] Then H(x)≈0.49+0.49+0.53+0.52=2.03;
[0098] The process of optimizing the initial information entropy is:
[0099] Let S i is the sum of the information fit corresponding to the multi-source factor information in the i-th subcategory, and S is the sum of the information fit of all subcategories;
[0100] For the numerical interval (6.7, 6.8]: let the sum of the information fit within the subcategory be S1, and the sum of the information fit within the subcategory be y1.
[0101] For the numerical interval (6.8, 6.9]: let the sum of the information fit within the subcategory be S2, and the sum of the information fit within the subcategory be y2.
[0102] For the numerical interval (6.9, 7.0]: let the sum of the information fit within the subcategory be S3, and the sum of the information fit within the subcategory be y3.
[0103] For the numerical interval (7.0, 7.1]: let the sum of the information fit within the subcategory be S4, and the sum of the information fit within the subcategory be y4.
[0104] The ratios within each subcategory are calculated as follows:
[0105] For the numerical interval (6.7, 6.8]: T1 = S1 / S = y1 / (y1+y2+y3+y4);
[0106] For the numerical interval (6.8, 6.9]: T2 = S2 / S = y2 / (y1+y2+y3+y4);
[0107] For the numerical interval (6.9, 7.0]: T3 = S3 / S = y3 / (y1+y2+y3+y4);
[0108] For the numerical interval [7.0, 7.1]: T4 = S4 / S = y4 / (y1+y2+y3+y4);
[0109] Substitute the T1, T2, T3, and T4 calculated above into the optimized information entropy formula:
[0110] H(X)=-[T1log2(T1)+T2log2(T2)+T3log2(T3)+T4log2(T4)]
[0111] The initial information entropy is optimized according to the above formula, and a decision tree is constructed according to the information gain ID3 algorithm. The decision tree is constructed according to the traditional information gain ID3 algorithm. The influencing factor data with the largest information gain is selected as the root node, and the root node is split according to the category. The calculation of information entropy and weighted optimization are common means of decision tree construction. The method of constructing a decision tree according to the traditional information gain ID3 algorithm is a technical means well known to those skilled in the art and will not be elaborated here.
[0112] Specifically, see Figure 5 As shown in FIG, it is a logic flow chart for determining whether the pipeline status is abnormal according to an embodiment of the present invention. The process of determining whether the pipeline status is abnormal includes:
[0113] Comparing the pipeline pressure prediction interval with the pipeline pressure actual value;
[0114] If the actual value of the pipeline pressure does not meet the normal standard conditions, it is determined that the pipeline state is abnormal;
[0115] If the actual value of the pipeline pressure meets the normal standard condition, it is determined that the pipeline state is normal;
[0116] The normal standard condition is that the actual value of the pipeline pressure falls within the pipeline pressure prediction interval.
[0117] Specifically, the present invention determines whether the pipeline is abnormal by comparing the pipeline pressure prediction interval and the actual value. The pipeline pressure prediction interval determined by the decision tree is obtained by integrating multi-source factor information. The comparison of the predicted value and the actual value after comprehensive consideration of multiple factors can effectively avoid misjudgment caused by single factor judgment. During the construction process of the decision tree, the initial information entropy of the multi-source factor information is optimized, and factors such as the information fit between each factor are considered, so that the predicted value can better reflect the actual situation. When the difference between the predicted value and the actual value is large, it indicates that the pipeline pressure does not follow the natural gas transmission characteristic data relationship under normal conditions, and the data relationship may be abnormal due to pipeline state fluctuations. In addition, the influence of noise data is eliminated, data distortion of the constructed prediction model is avoided, and the effectiveness of pipeline state prediction and the timeliness of abnormality monitoring are improved.
[0118] The specific implementation object of the method of the present invention can be a chip, component or module, and the chip may include a connected processor and memory; wherein the memory is used to store instructions, and when the processor calls and executes the instructions, the chip can execute the pipeline state prediction method based on big data.
[0119] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.
[0120] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A pipeline status prediction method based on big data, characterized in that: include: Obtaining multi-source factor information of the monitored point of the transmission pipeline at each monitoring time, the multi-source factor information including gas input amount, gas flow rate, branch pipeline flow rate and branch pipeline flow rate of the associated branch area, and exogenous environmental parameters; At each monitoring moment, the first data source credibility is determined based on the fluctuation of the external environmental parameters corresponding to the current monitoring point in the time dimension. The second data source credibility is determined based on the first correlation between the gas input volume and the gas flow rate at the current monitoring point in the time dimension and the second correlation between the branch pipeline flow rate and the flow rate in the associated branch area in the order of the branch pipeline diameter arrangement; Determining the information consistency at each monitoring moment based on the first data source credibility and the second data source credibility; Based on the numerical values of each information dimension, the multi-source factor information is divided into several subcategories. According to the information fit of the multi-source factor information of each subcategory, the initial information entropy of the multi-source factor information of each subcategory is optimized, and a decision tree is constructed based on the information gain of the optimized information entropy. Wherein, the leaf node of the decision tree is the pipeline pressure of the current point to be monitored; Determine the pipeline pressure prediction interval of the current monitoring point according to the decision tree, and determine whether the pipeline state is abnormal according to the comparison between the pipeline pressure prediction interval and the actual value of the pipeline pressure; The process of determining the first change relevance includes: Obtain the gas flow rate of the current monitoring point at each monitoring moment and the gas input volume of the pipeline where the current monitoring point is located; Determine the gas flow rate and gas flow rate at several monitoring moments within a preset monitoring neighborhood with the current monitoring moment as the time center; Determine the Pearson correlation coefficient between the gas flow rate and the gas input amount at a plurality of monitoring moments within a preset monitoring neighborhood as the first change correlation; The process of determining the relevance of the second change includes: Obtaining the associated branch area corresponding to the current monitoring point, determining the diameters of each branch pipeline in the associated branch area, and sorting the diameters of each branch pipeline; The gas flow rate and gas flow rate of each branch pipe are determined, the Pearson correlation coefficient of the gas flow rate and gas flow rate corresponding to branch pipes of several branch pipe diameters is calculated, and the Pearson correlation coefficient is determined as the second change correlation.
2. The pipeline state prediction method based on big data according to claim 1 is characterized in that: The associated branch area is an area in the upstream pipeline area of the point to be monitored where a pipeline branch node closest to the point to be monitored is located.
3. The pipeline state prediction method based on big data according to claim 1, characterized in that: The process of determining the fluctuation of the exogenous environmental parameters corresponding to the current monitoring point in the time dimension includes: Based on the time dimension sequence, obtain the exogenous environmental parameters at the current monitoring time and other monitoring times within the preset monitoring neighborhood; Calculate the difference between the current monitoring time and the exogenous environmental parameters at other monitoring times within the preset monitoring neighborhood.
4. The pipeline state prediction method based on big data according to claim 3 is characterized in that: The first number source credibility is determined according to the difference, and the first number source credibility is negatively correlated with the difference.
5. The pipeline state prediction method based on big data according to claim 1 is characterized in that: The second data source credibility is determined according to the first change correlation and the second change correlation; The second data source credibility is positively correlated with the first change correlation and the second change correlation respectively.
6. The pipeline state prediction method based on big data according to claim 5 is characterized in that: The information consistency is determined based on the credibility of the first data source and the credibility of the second data source; The information consistency is obtained by normalizing the product of the first data source credibility and the second data source credibility.
7. The pipeline state prediction method based on big data according to claim 6, characterized in that: The process of optimizing the initial information entropy includes: The multi-source factor information is divided into several subcategories according to the numerical value of each information dimension; Calculate the ratio of the sum of the information fit corresponding to each multi-source factor information in each subcategory to the sum of the information fit of all subcategories; The ratio is substituted into the information entropy calculation formula to optimize the initial information entropy.
8. The pipeline state prediction method based on big data according to claim 1, characterized in that: The process of determining whether the pipeline status is abnormal includes: Comparing the pipeline pressure prediction interval with the pipeline pressure actual value; If the actual value of the pipeline pressure does not meet the normal standard conditions, it is determined that the pipeline state is abnormal; The normal standard condition is that the actual value of the pipeline pressure falls within the pipeline pressure prediction interval.
Citation Information
Patent Citations
Online monitoring and safety research method for foundation settlement of buried natural gas pipeline
CN114722662A
Data monitoring method for steady-state operation of natural gas pipeline
CN116717734A
Fluid monitoring module arrangements
US20200088558A1