A natural gas data management method and system
The method integrates and optimizes natural gas data management through standardization, semantic tagging, and deep learning to enhance predictive scheduling and risk alerting, addressing inefficiencies in existing systems and improving adaptive control.
Patent Information
- Application Number
- CN202510437238.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-09
AI Technical Summary
In the prior art, natural gas data management lacks multi-source data fusion capabilities and semantic modeling capabilities, and cannot achieve real-time perception and dynamic response, resulting in low scheduling efficiency, delayed risk response and lag in safety control.
By obtaining data from SCADA systems, sensor equipment and historical databases for missing value filling and semantic label annotation, a natural gas system ontology network is built, entity relationship maps and feature vector sets are generated, multi-task deep neural network model is trained, gas volume prediction, scheduling and risk warning are carried out, and adaptive optimization model is established.
It realizes the unified fusion and structural semantic expression of multi-source data, improves the prediction and scheduling synergy performance, enhances the system's robustness and dynamic regulation capabilities, and builds an intelligent scheduling system with sustainable iterative optimization.
Smart Images

Figure CN119962975B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing and management, and particularly to a natural gas data management method and system. Background Art
[0002] With the continuous expansion of the urban gas scale and the promotion of the natural gas clean energy strategy, the natural gas transmission and distribution system gradually presents characteristics such as multi-source heterogeneous device access, complex data structure, dynamic change of operation status, and rapid spread of safety risks. During the whole process of natural gas transmission and distribution, a large number of operation parameters and business data from SCADA systems, edge sensors, component analysis devices, and historical databases are involved, such as indicators like pressure, flow rate, temperature, and gas composition, and are distributed among multiple nodes, pipeline segments, and end-users. There are problems such as large structural differences, high timeliness requirements, complex business semantics, and weak correlation relationships among these data, making it difficult to directly use them for intelligent analysis and scheduling optimization.
[0003] In the prior art, natural gas data management mainly focuses on single-point collection and offline processing, lacking the semantic modeling ability and task collaboration mechanism for the whole system, and unable to support real-time perception and dynamic response to the operation situation. At the same time, the prediction and scheduling models are usually designed in isolation and cannot form a closed-loop optimization path with risk early warning or historical strategies, resulting in low scheduling efficiency, delayed risk response, and lagged safety control. Therefore, there is an urgent need for a data management method and system that supports multi-source data fusion, has semantic expression ability, and integrates prediction, scheduling, and risk early warning functions to achieve the intelligent and adaptive operation control of the natural gas transmission and distribution system. Summary of the Invention
[0004] The present invention provides a natural gas data management method and system to solve the problem of how to construct an integrated intelligent management model with prediction, scheduling, early warning, and adaptive optimization capabilities based on multi-source monitoring data, operation status parameters, and semantic ontology relationships in the natural gas transmission and distribution system, and realize the dynamic perception, intelligent scheduling, and risk closed-loop control of the entire natural gas process.
[0005] To solve the above technical problems, the present invention provides a natural gas data management method, and the method includes:
[0006] Obtain the original data from the SCADA system, sensor devices, and historical database, perform missing value filling, field standardization, and semantic label annotation to generate a structured semantic data set;
[0007] Extract the flow node, pipeline segment attribute, and gas quality index information from the structured semantic data set, and combine with the preset ontology rules to construct a natural gas system ontology network, generating an entity relationship map and a set of feature vectors;
[0008] Train a multi - task deep neural network model based on the graph spectrum and feature set to obtain a jointly optimized model;
[0009] Input the natural gas status data into the model to generate a gas volume prediction sequence and a preliminary scheduling vector, and adjust the path constraints according to the system boundary conditions to generate a task execution vector;
[0010] Execute the vector to generate distribution and transmission instructions, collect pressure offsets and flow rate fluctuations, extract risk factors, and then combine with the knowledge rule base to identify abnormal events and generate risk warning parameters;
[0011] Feed the execution results, error sequence, and risk identification back to the model to complete parameter update and weight correction, and output a new round of optimized model parameters.
[0012] Furthermore, the field standardization includes time alignment and unit format unification for SCADA data, sensor data, and historical data.
[0013] Furthermore, the semantic label annotation includes adding node labels and functional categories to each sampled field for ontology network construction.
[0014] Furthermore, the constructed ontology network forms a graph spectrum structure through triple extraction, where the nodes contain natural gas equipment entities and the edges represent operation dependency relationships.
[0015] Furthermore, the Based on the softmax ranking loss function calculation, it is used to optimize the scoring difference of the scheduling path.
[0016] Furthermore, the system boundary conditions include the maximum allowable flow rate of nodes, pressure difference limits, and network connectivity rules.
[0017] Furthermore, the task execution vector is transformed into equipment - layer scheduling instructions through a mapping relationship, including target pressure, valve opening state, and compression ratio setting.
[0018] Furthermore, the pressure offset is the pressure difference between adjacent time steps, and the flow rate volatility is the ratio of flow rate change per unit time.
[0019] Furthermore, the risk knowledge rule base contains multiple abnormal type mapping rules, and the risk level is generated according to the fluctuation characteristics.
[0020] Furthermore, a natural gas data management system. The system is used to implement the above - mentioned method, and the system includes:
[0021] The data acquisition and preprocessing module is used to collect raw data from the SCADA system, sensor monitoring devices and historical databases, and perform missing value filling, field standardization and semantic tag annotation processing;
[0022] The semantic modeling and feature extraction module is used to extract natural gas feature information from structured semantic datasets, construct an ontology network and generate an entity relationship graph and a set of feature vectors;
[0023] The multi-task modeling module is used to train a multi-task deep neural network model based on the graph and feature set to generate a jointly optimized model;
[0024] The scheduling generation and instruction issuing module is used to adjust the scheduling vector according to the prediction results and boundary constraints output by the model and generate device control instructions;
[0025] The status monitoring module is used to collect operation status data during the transmission and distribution execution process and generate pressure offset and flow fluctuation sequences;
[0026] The anomaly recognition and feedback module is used to identify the levels of abnormal events based on a risk rule library and feedback the risk results to the model to complete parameter updates, forming a closed-loop optimization mechanism;
[0027] The policy feedback and model update module is used to feedback the task execution results, prediction error sequences and risk identifications to the jointly optimized model, update the policy parameters of the model and correct the training weights, and generate an updated set of model parameters.
[0028] The following are its main beneficial effects:
[0029] (1) Realize the unified fusion and structural semantic expression of heterogeneous data: By performing field standardization, missing value filling and semantic tag annotation on the operation data in the SCADA system, sensor devices and historical databases, a structured semantic dataset is generated, solving the problems of inconsistent multi-source data formats and unclear semantics in the natural gas industry, and improving the subsequent modeling efficiency and system adaptation ability.
[0030] (2) Construct a multi-task joint optimization model to improve the collaborative performance of prediction and scheduling: The present invention extracts feature information based on the graph structure, trains a deep neural network model with the functions of flow prediction and path scheduling, and introduces a joint optimization mechanism of scheduling score and prediction error. Compared with traditional single-objective models, it has a higher overall performance in terms of prediction accuracy and scheduling rationality.
[0031] (3) Establish a model adaptive update mechanism to enhance the system's robustness and dynamic adjustment ability: During the scheduling execution process, the system collects operation status data, identifies abnormal fluctuations, and combines with the risk rule library to judge the event level. It feeds back the abnormal features and scheduling deviations into the model for parameter correction and policy fine-tuning, constructs an intelligent scheduling system that can be continuously iteratively optimized, and enhances the system's adaptability to complex operation states. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a schematic flowchart of a natural gas data management method provided by an embodiment of the present application;
[0033] Figure 2 It is a structural block diagram of a natural gas data management system provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above BRIEF DESCRIPTION OF THE DRAWINGS are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.
[0035] Referring to "embodiment" herein means that the specific features, structures, or characteristics described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0036] Embodiment 1: Refer to Figure 1 It is a schematic flowchart of a natural gas data management method provided by an embodiment of the present invention, and this process can at least include steps S100 - S700:
[0037] S100. Obtain the original data from the natural gas SCADA system, sensor monitoring devices, and historical database, fill in the missing values, standardize the fields, and label the semantic tags for the original data to generate a structured semantic dataset;
[0038] S200. Extract the natural gas flow nodes, pipe section attributes, and gas quality index information from the structured semantic dataset, and construct the ontology network of the natural gas system in combination with the preset ontology rules to generate an entity relationship graph and a feature vector set;
[0039] S300. Based on the entity relationship graph and the feature vector set, train the multi-task deep neural network model to obtain a joint optimization model with the ability of natural gas flow prediction and path scheduling;
[0040] S400. Input the natural gas status data to be processed into the joint optimization model to generate a gas volume prediction sequence and a preliminary scheduling vector, and adjust the path constraints according to the system boundary conditions to generate a task execution vector;
[0041] S500. Based on the task execution vector and the real-time operation monitoring data, perform command scheduling execution on the natural gas transmission and distribution process, and collect the pressure offset and flow fluctuation sequence during the scheduling process;
[0042] S600. Extract the features of the pressure offset and flow fluctuation sequence, combine with the risk knowledge rule base to perform anomaly identification and event level determination, and generate dynamic risk identification and warning parameters;
[0043] S700. Feed the task execution result, prediction error sequence, and risk identification back to the joint optimization model to update the strategy parameters and correct the training weights, and generate a new round of iteratively optimized model parameters;
[0044] Step S100 includes at least steps S110 - S130:
[0045] S110. Obtain the operation parameter data in the natural gas SCADA system, the on-site sampling data in the distributed sensor monitoring device, and the structured data in the historical business database, and perform unified processing on the field formats of the data to obtain a standardized dataset;
[0046] Specifically, the data input of the natural gas data management system includes three key sources: the operation parameter data of the natural gas pipeline network in the SCADA system , the high-frequency gas physical detection data collected by on-site sensors , and the structured operation record data in the historical business management system , where represents the sampling time, represents the SCADA system acquisition source, represents the monitoring device node, represents the historical database index.
[0047] Input the above three types of heterogeneous data into the data management engine, and for the , and perform field standard mapping processing. Use a unified field template , where represents the th unified field (such as pressure , temperature , flow rate , component concentration , etc.), align the fields of each data source to obtain a standardized data set .
[0048] The field standardization function can be expressed as:
[0049]
[0050] where represents the field format conversion and standard mapping operation, is the unified field set, and the output represents the standardized data sample set.
[0051] S120. Perform missing value repair processing and outlier removal operations on the standardized data set to obtain a cleaned data set after integrity verification.
[0052] On the basis of the , jointly process the possible data missing values and observed outliers therein to construct a complete data set ;
[0053] First, fill in the missing fields with the moving window mean and perform time series interpolation repair. Let represent the value of the th field at time . Use the following repair function:
[0054]
[0055] where is the moving window length, and the value can be 5 or 10, specifically depending on the historical data sampling frequency. j is an index variable of the historical time step under a moving window, representing the jth moment before the current time point t. The value range of j is from 1 to W, that is, j = 1, 2,..., W.
[0056] Secondly, perform 3σ rule rejection on the value. Let its mean be , and the standard deviation be . If it satisfies:
[0057]
[0058] It is determined as an outlier and excluded, and finally a cleaned data set after integrity verification is formed:
[0059]
[0060] Among them, represents the combined missing value repair and outlier exclusion operation.
[0061] S130. Perform entity annotation and semantic label generation on each data record in the cleaned data set, construct a semantic annotation mapping table, and obtain a structured semantic data set.
[0062] Furthermore, for each data record in perform entity annotation and semantic label generation processing, extract structured semantic units, and map information such as device entities, flow channels, geographical locations, and transmission behaviors. Construct a semantic label set , and each corresponds to a semantic label (such as "inlet pressure node", "key gas transmission pipeline segment", "fluctuation point", "risk source", etc.).
[0063] Define the entity annotation function as follows:
[0064]
[0065] Among them, is the semantic mapping function, represents the ontology model (preset ontology structure), represents the rule set, represents at time for data record generated semantic label.
[0066] Finally, establish a semantic annotation mapping table:
[0067]
[0068] and output a structured semantic data set:
[0069]
[0070] The will be directly used as the input data set of the S210 feature extraction and ontology construction module. In particular, in S210, natural gas flow nodes, pipeline segment attributes, and gas quality indicators will be extracted from this data set for feature vector construction and ontology network generation to ensure the context traceability and structural consistency of semantic data.
[0071] Step connection description:
[0072] Connection description between S110 and S120: In step S110, the field standardization of multi-source raw data from the SCADA system, sensor devices, and historical database is completed to form a dataset with a unified structure. This standardized dataset is used as the input of S120 for further performing missing value repair and outlier removal operations.
[0073] Connection description between S120 and S130: The cleaned dataset after integrity verification processed in step S120 provides a stable data foundation without missing values and outliers for S130. S130 uses this as the input to generate semantic labels and perform entity annotation, and outputs a structured semantic dataset. .
[0074] Connection description between S130 and subsequent S200: The structured semantic dataset constructed in S130 contains the natural gas operation data after semantic annotation and its corresponding entity label information, serving as the basic input data source for feature extraction and ontology relationship construction in S200.
[0075] Step S200 includes at least steps S210 - S230:
[0076] S210: Extract flow node features, pipe section attribute features, and gas quality index features related to natural gas flow, pressure, temperature, etc. from the structured semantic dataset to generate an initial set of feature vectors.
[0077] Specifically, receive the structured semantic dataset output from S130 , where represents the th cleaned data record, is its corresponding semantic label, represents the sampling moment. Perform semantic label screening and field parsing on the dataset to extract the core indicators of natural gas operation, including but not limited to:
[0078] Natural gas flow characteristics: , unit m³ / h;
[0079] Gas transmission pressure characteristics: , unit MPa;
[0080] Pipeline temperature characteristics: , unit °C;
[0081] Gas quality component concentration characteristics: , where represents the component category, such as methane, ethane, etc.
[0082] On this basis, an initial set of feature vectors is constructed , which is defined as:
[0083]
[0084] where represents the total number of component types. The will serve as the basic input for the ontology classification and semantic relationship modeling of S220.
[0085] S220, based on the preset ontology rules in the natural gas field, performs semantic parsing and concept classification on the entity attributes and relationship dimensions in the initial set of feature vectors to construct a natural gas entity hierarchical structure model;
[0086] Furthermore, based on the ontology rule library constructed from the expert knowledge in the natural gas field , classify and hierarchically map various attributes in :
[0087] represents the category set, such as "main pipeline section", "pressure regulating node", "user terminal", etc.;
[0088] represents the attribute set, such as "maximum allowable pressure", "normal operating temperature range";
[0089] represents the relationship set, such as semantic connection relationships like "connection", "transportation", "dependency", etc.
[0090] According to the ontology mapping function , perform attribute parsing operations:
[0091]
[0092] where represents the set of natural gas field entities constructed at the sampling time , and each entity contains attribute information and semantic category labels. Combining with the semantic relationship set , construct an entity hierarchical structure model :
[0093]
[0094] This model describes the upstream and downstream dependencies, transmission and distribution logic, and functional classification among entities.
[0095] S230, perform triple extraction and graph structure encoding on the natural gas entity hierarchical structure model to generate an entity relationship graph and a set of trainable feature vectors for the natural gas system;
[0096] After the construction is completed On this basis, further extract triples from the entity attribute relationships to form a knowledge graph structure in standard form. Define the extraction function :
[0097]
[0098] Among them, and represent two semantic entities, represents the semantic relationship connecting them in the triple (such as "downstream connection", "controlled by", "belongs to"); is the entity relationship graph of the natural gas system at time .
[0099] Construct input vectors for the graph neural network (GNN) of the said graph to generate a set of trainable feature vectors , for use in the subsequent model training module:
[0100]
[0101] Among them is the graph structure vector encoding function, which combines the graph structure adjacency matrix and node features to generate an embedded representation vector. The will be used as the input in the next module (S300) for training the multi-task optimization model.
[0102] Description of step connection:
[0103] Description of the connection between S210 and S220: The initial set of feature vectors constructed in S210 is used as the input data for the ontology classification and semantic mapping in S220 to ensure the semantic interpretability of the data fields.
[0104] Description of the connection between S220 and S230: The natural gas entity hierarchical structure model constructed in S220 serves as the structural basis for performing triple extraction and graph encoding in S230, and is used to generate the graph structure and the vector set .
[0105] Step S300 includes at least steps S310 - S330:
[0106] S310. Based on the entity relationship graph and the set of feature vectors, construct the input layer and feature embedding structure of the deep neural network, and embed the semantic vectors into the input layer;
[0107] Specifically, receive the entity relationship graph output from step S230 and its corresponding set of graph structure encoding vectors , where represents the th entity node's semantic feature embedding vector at time .
[0108] Based on the Graph Neural Network (GNN) structure, construct an embedding propagation function and adopt the following information aggregation mechanism for input layer vector fusion:
[0109]
[0110] Among them, represents the th node's embedding representation vector, represents the neighbor set connected to node , represents the attention weight of the neighbor node, represents the learnable transformation weight matrix, represents the non-linear activation function (such as ReLU).
[0111] Furthermore, take the as a component of the input layer of the joint model and construct a unified input tensor for the multi-task network structure:
[0112]
[0113] S320. Based on the embedding structure and the scheduling objective function, jointly model the gas volume prediction task and the path scheduling task, train the multi-task deep neural network, and obtain the pre-trained joint optimization model;
[0114] Based on the input tensor , construct a two-branch multi-task deep neural network structure to separately execute the two sub-tasks of gas volume prediction and path scheduling. Let:
[0115] represent the predicted gas volume of the th node at the future time ; represent the scheduling priority score from node to node .
[0116] Define the joint loss function as:
[0117]
[0118] Among them, is the gas volume prediction error loss term:
[0119]
[0120] It is the scoring ranking error loss (such as pairwise ranking) in the path scheduling task:
[0121]
[0122] Among them, Sum over all training sample pairs (i, j), where i represents the task number and j represents the number of the target path in the current task; It represents the softmax normalization of the predicted score values of all candidate paths k for the i-th task, constituting the normalization term of the probability distribution; The predicted score value of the j-th path of the i-th task by the joint optimization model at time step t.
[0123] By minimizing the above joint loss function , perform backpropagation and weight update, iteratively train the multi-task model, and generate a set of pre-trained parameters:
[0124]
[0125] Among them, , are the network weights of the prediction branch and the scheduling branch respectively; , are the multi-task loss balance weight parameters, and the values are such as , , which are obtained by tuning with the historical validation set.
[0126] S330. Fine-tune and train the joint optimization model, and solidify the model structure and parameter weights to obtain an optimized model parameter set with the functions of natural gas flow prediction and path scheduling;
[0127] Furthermore, based on the actual operation feedback data set for the perform fine-tuning training, specifically including scheduling execution deviation, prediction error and system feedback indicators. Introduce the scheduling deviation function and the prediction residual :
[0128]
[0129] According to and adjust the loss function weighting strategy according to the statistical distribution, and continue to iteratively train the model, and finally solidify the optimized parameter set:
[0130]
[0131] Among them, represents the optimized model parameters after fine-tuning, including the final network structure, number of layers, weight matrix, and scheduling scoring mechanism, which will be used as the basis for the S400 module to generate scheduling task vectors.
[0132] Explanation of step connection:
[0133] Explanation of the connection between S310 and S320: The embedding representation constructed in S310 , as the input data structure of the multi-task neural network in S320, supports simultaneous gas volume prediction and scheduling modeling.
[0134] Explanation of the connection between S320 and S330: The pre-trained parameters generated by training in S320 serve as the initial model for further fine-tuning and optimization in S330, and are used to generate a refined set of final optimized parameters in combination with subsequent feedback data .
[0135] Step S400 includes at least steps S410 - S430:
[0136] S410: Input the natural gas state data to be processed into the joint optimization model to generate a natural gas volume prediction sequence and a preliminary path scheduling vector;
[0137] Specifically, obtain the real-time natural gas state data at the moment to be processed , where the includes pressure distribution , flow distribution , temperature distribution , and station load requests . .
[0138] Input the into the jointly optimized model trained in step S300 to generate a sequence of gas volume prediction values at future times and a path scheduling vector :
[0139]
[0140] Among them, represents the future gas volume prediction sequence for each flow node; represents the scheduling priority scoring matrix from node to node .
[0141] S420. Perform system boundary condition detection and constraint rule matching on the preliminary path scheduling vector, adjust the path plan, and generate a task execution vector that satisfies the constraints;
[0142] Understandably, perform system-level boundary condition constraint matching on the scheduling path scores in . Let the set of natural gas system constraint conditions be , including but not limited to:
[0143] Node maximum allowable flow constraint ; Pipe section maximum pressure difference constraint ; Pipe network connectivity and energy consumption constraints
[0144] According to the scheduling score vector and the above constraints construct a scheduling feasibility function , reorder and adjust the score vector, and generate a task execution vector that meets the conditions:
[0145]
[0146] Among them, each scheduling relationship in represents an executable path for transporting natural gas from node to node at time and satisfies all physical and business constraints.
[0147] S430. Perform node scheduling mapping and device instruction conversion on the task execution vector to generate an executable transmission and distribution scheduling instruction set;
[0148] Furthermore, map the scheduling path relationships in the task execution vector to the device layer nodes of the natural gas scheduling control system to form a node-device mapping table , where represents the controllable valve, compressor or pressure regulating station equipment under the node.
[0149] Based on perform device layer translation on to generate a transmission and distribution scheduling instruction set :
[0150]
[0151] Among them, is a device mapping function, and the output represents at time Executable control instructions, including opening and closing valve operations, setting pressure differences, adjusting compression ratios, etc., and the instruction content is recognized and issued by the SCADA control interface.
[0152] Description of step connection:
[0153] Description of the connection between S410 and S420: Based on the optimization model in S410 Predicted And , as the input vector for S420 to perform path feasibility judgment and scheduling constraint matching, is used to generate .
[0154] Description of the connection between S420 and S430: The task execution vector output by S420 As the basic input for instruction mapping in S430, through cooperation with the device mapping relationship , it is transformed into a scheduling instruction set with device layer execution capabilities .
[0155] Step S500 includes at least steps S510 - S530:
[0156] S510. Send the distribution and scheduling instruction set to the natural gas distribution system execution module to trigger on-site equipment to perform corresponding distribution operations;
[0157] Specifically, receive the distribution and scheduling instruction set generated by step S430 , the Indicates the th control instruction, including equipment operation parameters , such as valve opening and closing status, compressor speed setting, pressure regulation target value, etc.
[0158] Send to the edge execution module of the natural gas distribution system , which accesses the on-site control interface SCADA / PLC layer and executes specific actions. For each instruction , the execution process is as follows:
[0159]
[0160] Among them, represents the on-site actual response actions, including gas flow path switching, pressure setting update, etc.
[0161] S520. During the scheduling execution process, collect the operation status parameters of each flow node and pressure monitoring point to generate an operation status data sequence;
[0162] After the execution module is activated, start the distributed monitoring device to sample the natural gas operation status parameters at key nodes. The acquisition metrics include:
[0163] Pressure monitoring data: ; Flow monitoring data: ; Temperature and other optional environmental parameters: .
[0164] The acquisition period is denoted as , and a continuous operation status data sequence is constructed:
[0165]
[0166] where represents the number of time steps in the observation window. The will be used as the input for subsequent fluctuation detection and anomaly recognition.
[0167] S530. Sample and calculate the pressure offset and flow volatility in the operation status data sequence to generate a status fluctuation sequence;
[0168] Furthermore, perform differential operation and normalization on the continuous status data in to extract the pressure offset and flow fluctuation factors, and construct a status fluctuation sequence .
[0169] First, calculate the pressure offset at node and and the flow change rate at adjacent time steps :
[0170]
[0171] where is a very small positive number to avoid a zero denominator; and are the pressure and flow values of node at time respectively.
[0172] Concatenate the above metrics along the time axis to form a status fluctuation sequence:
[0173]
[0174] The serves as the input variable basis for the subsequent S600 risk identification module, and is used to quantify potential operation anomalies in the natural gas system.
[0175] Step connection description:
[0176] Connection description between S510 and S520: The control instructions issued in S510 directly trigger the on-site equipment to execute the scheduling operation, and the resulting flow rate changes and pressure changes are the operation status data sequences collected in S520 which are the sampling sources
[0177] Connection description between S520 and S530: The operation status sequences collected in S520 provide data support for the differential analysis and volatility calculation in S530, and finally generate the status volatility sequence as the input for the subsequent module
[0178] Step S600 includes at least steps S610 - S630:
[0179] S610. Extract features and variability analysis from the status volatility sequence to extract the characteristics of potential abnormal risk factors;
[0180] Specifically, receive the status volatility sequence output from S530:
[0181]
[0182] where represents the pressure offset of node at time , and represents the flow rate volatility of node . Construct a risk factor feature vector for each node:
[0183]
[0184] where is the mean of the pressure offsets of node within the window, is its standard deviation, and are the mean and standard deviation of the flow rate volatility, used to characterize the abnormal volatility trend
[0185] Furthermore, to enhance the ability to identify mutation risk factors, the coefficient of variation is introduced to characterize the volatility intensity:
[0186]
[0187] where is a small constant used to prevent the denominator from being zero; when is significantly higher than the historical stability threshold of the node , it is marked as a high volatility risk point Derived from historical data distribution, such as empirical values 。
[0188] S620. Match the risk factor characteristics with a preset risk knowledge rule base, identify the abnormal type, and determine the event level;
[0189] Based on the risk factor characteristics extracted above and the coefficient of variation , enter the abnormal identification process. Call the preset natural gas system risk knowledge rule base:
[0190]
[0191] Among them, each rule represents the matching conditions for a type of risk scenario (such as "sudden increase in high pressure", "continuous fluctuation", "sharp drop in flow rate", etc.), in the following form:
[0192]
[0193] Among them, is the rule trigger threshold, is the risk type label. By matching with in the rules, identify the risk type label of each node at time .
[0194] Furthermore, according to the event occurrence frequency, risk intensity, and system impact degree, set the event level function , and output the risk level :
[0195]
[0196] Among them, represents the pipeline network importance weight of node , which can be defined by domain experts or obtained by calculating the topological position (such as the main pipe section , the end user ), is the risk level scoring function, and the output level division is such as "low", "medium", "high", "severe", etc.
[0197] S630. Perform multi-dimensional risk identification coding on the determination result to generate the dynamic identification and warning parameters of the risk event;
[0198] Furthermore, encode the risk type and the level result into a multi-dimensional dynamic risk identification , including the following components:
[0199]
[0200] The is a structured dynamic risk event object, supporting multi-source system identification, storage, and visualization.
[0201] Furthermore, integrate the risk event identifiers of all nodes to construct a system-level risk early warning parameter set at a specific moment:
[0202]
[0203] The will be used as one of the inputs for the subsequent policy feedback module (S700) to trigger model update and parameter correction.
[0204] Explanation of step connection:
[0205] Explanation of the connection between S610 and S620: The risk factor vector extracted in S610 and the mutation index serve as the core inputs for rule matching and event identification in S620, driving the determination of risk type and level.
[0206] Explanation of the connection between S620 and S630: The event type and level results output by S620 serve as the dynamic coding basis in S630. Through uniformly describe the risk events and integrate them into system-level early warning parameters for use in the next module.
[0207] Step S700 at least includes steps S710 - S730:
[0208] S710: Use the scheduling execution result, prediction error sequence, and risk identifier as feedback inputs to generate a joint feedback data set;
[0209] Specifically, obtain the following feedback inputs:
[0210] Actual operation status data from S530 ; Prediction output from S410 for calculating the prediction error sequence; Risk event identifier set from S630 , containing dynamic risk levels and impact factors.
[0211] Based on the above data, construct an error metric function:
[0212]
[0213] Among them, represents the gas volume prediction error, represents the path scheduling deviation, is the true scheduling path score after constraint adjustment (from S420).
[0214] Combine the error sequence and risk events to construct a joint feedback dataset:
[0215]
[0216] The is used as the online training correction input for the joint model.
[0217] S720. Based on the joint feedback dataset, perform gradient correction and policy fine-tuning on the parameter weights in the joint optimization model;
[0218] Furthermore, define a correction loss function:
[0219]
[0220] where: is the feedback loss weighting parameter, such as , , , which is determined by tuning the historical validation set;
[0221] represents a function that maps the risk level and volatility factor to a penalty value, for example, a high-risk event is given a greater update amplitude;
[0222] The third term introduces the influence of risk identification on gradient adjustment, making the model more sensitive to the risk event area.
[0223] Based on the above loss function, use the backpropagation algorithm to perform gradient update on the parameter set in the joint optimization model :
[0224]
[0225] where, is the learning rate (such as ), represents the gradient with respect to the parameter.
[0226] S730. Perform iterative training on the adjusted joint optimization model to generate an updated model parameter set for the next round of inference;
[0227] Based on the obtained weight parameters as the initial model structure, perform iterative fine-tuning training on the data within the subsequent observation window to construct a dynamically evolving version of the optimization model 。
[0228] Let the new input sequence be , and perform the incremental training process:
[0229]
[0230] Among them, represents the incremental training function. The input is the current model and its new data samples, and the output is the new round of model parameters after fine-tuning. After the training is completed, solidify , and replace the original optimized model , which is used for the next round of inference prediction task in step S410.
[0231] Explanation of the connection between steps: Explanation of the connection between S710 and S720:
[0232] The joint feedback data set generated in S710 provides the prediction error and risk event coding, and provides the complete input for the calibration loss function constructed in S720 to achieve gradient update.
[0233] Explanation of the connection between S720 and S730: The parameters updated in S720 serve as the basic model for iterative fine-tuning in S730. After performing the incremental training process, they are solidified into the new inference model parameters to achieve parameter adaptive closed-loop.
[0234] The key innovation points of the present invention include:
[0235] (1) A semantic structure-driven multi-source natural gas data integration method.
[0236] (2) A joint optimization model design for both prediction and scheduling.
[0237] (3) A risk identification and model adaptive closed-loop mechanism based on operation feedback. This method is applicable to data management and intelligent control scenarios of various natural gas transmission and distribution networks, and has good promotion value and application prospects.
[0238] The following are its main beneficial effects:
[0239] (1) Achieve unified fusion and structural semantic expression of heterogeneous data: By performing field standardization, missing value filling, and semantic label annotation on the operation data in the SCADA system, sensor devices, and historical database, a structured semantic data set is generated, solving the problems of inconsistent multi-source data formats and unclear semantics in the natural gas industry, and improving the subsequent modeling efficiency and system adaptation ability.
[0240] (2) Build a multi-task joint optimization model to improve the collaborative performance of prediction and scheduling: Based on the graph structure, the present invention extracts feature information, trains a deep neural network model with traffic prediction and path scheduling functions, and introduces a joint optimization mechanism for scheduling scores and prediction errors. Compared with traditional single-object models, it has a higher overall performance in terms of prediction accuracy and scheduling rationality.
[0241] (3) Establish a model adaptive update mechanism to enhance the system's robustness and dynamic adjustment ability: During the scheduling execution process, the system collects operation status data, identifies abnormal fluctuations, and combines with the risk rule library to judge the event level. It feeds back the abnormal features and scheduling deviations into the model for parameter correction and policy fine-tuning, constructing an intelligent scheduling system that can be continuously iteratively optimized, and enhancing the system's adaptability to complex operation states.
[0242] Embodiment 2: Figure 2 Show a structural block diagram of a natural gas data management system according to an embodiment of the present invention. As Figure 2 shown, the system may include:
[0243] A data collection and preprocessing module 10, configured to obtain raw data related to the operation of the transmission and distribution system from a multi-source natural gas data entry, including pressure, flow, and temperature parameters in the SCADA system, gas component concentrations, valve states, and node operation information collected by edge sensor devices, as well as structured service data in the historical database. This module generates a structured data set with unified semantics and integrity through field standardization, missing value filling, and outlier removal, providing a high-quality data source for downstream semantic modeling and analysis.
[0244] A semantic modeling and ontology construction module 20, configured to perform semantic enhancement and ontology abstraction on the structured data set based on the natural gas industry knowledge system. This module constructs an entity label system and a semantic label mapping relationship, extracts key features such as pipe section attributes, node features, and gas quality indicators, forming a semantic graph and ontology network structure of the natural gas transmission and distribution system, providing a high-dimensional and interpretable structural representation for subsequent task scheduling and model training.
[0245] A multi-task prediction and scheduling optimization module 30, configured to construct a multi-task neural network model based on the constructed semantic graph and feature vectors, and jointly model the natural gas flow trend and transmission and distribution path. This module integrates the gas volume prediction loss function and the path scheduling score function, and obtains pre-trained model parameters with linkage optimization ability through historical data training, realizing the flow prediction of the target node and the generation of the path scheduling vector.
[0246] The boundary constraint and task generation module 40 is used to perform system physical constraint and business rule verification on the predicted and generated scheduling vectors, and match boundary conditions including maximum pressure limit, node load capacity, and pipe network connectivity. This module adjusts the scheduling plan according to the constraint detection results, generates a task execution vector that meets the constraint conditions, and maps the tasks to specific node devices to form a scheduling execution instruction set.
[0247] The scheduling execution and status monitoring module 50 is used to send the scheduling instructions to the natural gas transmission and distribution on-site execution system, and operate devices such as valves, compressors, and pressure regulators through the SCADA control port. At the same time, during the execution process, it collects pressure, flow rate, and temperature parameters of each node in real time. This module records the operation status data through a sliding window mechanism and constructs a time series status monitoring sequence for subsequent volatility evaluation and anomaly identification.
[0248] The risk identification and early warning module 60 is used to extract features and perform mutation analysis on the status fluctuations in the monitoring data, and combine with a preset risk rule library to identify the anomaly type and determine the risk level. This module performs intelligent matching on scenarios such as high-frequency fluctuations, high-pressure mutations, or sharp drops in flow rate, forms a multi-dimensional risk event identification structure, generates risk early warning parameters, and provides a quantitative basis for policy feedback.
[0249] The policy feedback and model update module 70 is used to use feedback information such as scheduling errors, risk events, and prediction deviations as joint inputs to perform policy fine-tuning and parameter adaptive update on the multi-task optimization model. This module constructs an online feedback loss function to realize incremental training of the neural network model, generates updated optimized model parameters, and uses them to replace the original model to continuously improve the prediction accuracy and scheduling robustness of the system under different working conditions.
[0250] A natural gas data management method and system provided by the present invention can realize the efficient fusion, semantic structure modeling, flow trend prediction, scheduling task generation, risk perception early warning, and adaptive closed-loop optimization of multi-source heterogeneous data in the natural gas transmission and distribution process, and has the following remarkable advantages:
[0251] Support the semantic modeling and graph construction of natural gas full-process business data, and improve data interpretability and task reasoning ability;
[0252] Establish a multi-task deep learning model to realize the collaborative optimization of gas volume prediction and path scheduling, and improve the prediction and scheduling decision-making efficiency;
[0253] Based on the feedback-driven mechanism, realize the closed-loop control logic of scheduling instruction - execution monitoring - anomaly identification - policy update, and enhance the intelligence and adaptive ability of the system;
[0254] The modular architecture has good scalability and deployment flexibility, and is suitable for various scenarios of gas transmission and distribution pipe networks in the natural gas industry.
[0255] Obviously, the embodiments described above are only a part of the embodiments of this application, rather than all of them. The preferred embodiments of this application are given in the accompanying drawings, but they do not limit the patent scope of this application. This application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure using the content of the specification and drawings of this application, directly or indirectly applied in other related technical fields, is similarly within the scope of patent protection of this application.
Claims
1. A natural gas data management method, characterized in that, The method includes: Obtaining the original data in the SCADA system, sensor devices, and historical database, performing missing value filling, field standardization, and semantic label annotation to generate a structured semantic dataset; Extracting the flow nodes, pipe section attributes, and gas quality index information in the structured semantic dataset, and constructing a natural gas system ontology network in combination with preset ontology rules to generate an entity relationship graph and a feature vector set; Training a multi-task deep neural network model based on the graph and the feature vector set to obtain a joint optimization model; the training of the joint optimization model takes the joint loss function as the optimization target: ; Among them, is the combined loss function, is the loss term of gas volume prediction error, is the scoring and sorting error loss in the path scheduling task, and is the multi-task loss balance weight parameter; By minimizing the combined loss function , perform backpropagation and weight update to generate a set of pre-trained parameters: ; Among them, is the set of pre-trained parameters; , are the network weights of the prediction branch and the scheduling branch respectively; Inputting the to-be-processed natural gas state data into the joint optimization model to generate a gas volume prediction sequence and a preliminary path scheduling vector, and performing path constraint adjustment according to the system boundary conditions to generate a task execution vector; the to-be-processed natural gas state data includes pressure distribution, flow distribution, temperature distribution, and site load requests; the system boundary conditions include the maximum allowable flow rate of nodes, pressure difference limits, and network connectivity rules; Executing the task execution vector to generate a transmission and distribution instruction, collecting the pressure offset and flow rate fluctuation during scheduling execution, and generating a dynamic risk identifier in combination with the risk knowledge rule base; the expression of the dynamic risk identifier is: ; Among them, is the dynamic risk identifier; is the risk type; is the level result; is the coefficient of variation; is the risk factor characteristic; Feeding back the gas volume prediction error, path scheduling deviation, and dynamic risk identifier to the joint optimization model, and using the backpropagation algorithm to update the gradient of the parameter set in the joint optimization model based on the correction loss function; the correction loss function is: ; Among them, is the calibration loss function; is the gas volume prediction error; is the path scheduling deviation; is a function that maps the risk level and the volatility factor to a penalty value; is the feedback loss weighting parameter.
2. The method according to claim 1, wherein wherein the is calculated based on the softmax ranking loss function and is used to optimize the scheduling path score difference.
3. The method according to claim 1, characterized in that, Wherein the task execution vector is converted into a device layer scheduling instruction through a mapping relationship, including target pressure, valve opening state, and compression ratio setting.
4. The method according to claim 1, wherein Wherein the pressure offset is the pressure difference between adjacent time steps, and the flow rate volatility is the ratio of the flow rate change per unit time.
Citation Information
Patent Citations
Power data analysis method and analysis system, power system, terminal equipment and storage medium
CN118396162A
Energy information query method, system and equipment based on large language model and medium
CN119719274A