A large model coupling working condition clustering natural gas load interval estimation method
By using a large-model coupled operating condition clustering method, and utilizing historical operating data and scenario background text of the natural gas system, the operating condition types and transitional relationships are identified. This solves the problems of operating condition identification and boundary ambiguity in natural gas load forecasting, and achieves accuracy and stability in natural gas load range forecasting.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2026-06-24
- Publication Date
- 2026-07-21
AI Technical Summary
Existing natural gas load forecasting methods struggle to effectively express the semantic attributes of operating conditions, leading to problems with operating condition identification and boundary ambiguity. Furthermore, traditional clustering methods cannot adapt to the continuous transitions and heterogeneous errors of multiple operating conditions in natural gas systems, resulting in the failure of interval estimation.
A large-model coupled operating condition clustering method is adopted. By acquiring historical operating data and scenario background text, the large model is used to identify operating condition types and transition relationships, construct semantic constraint information, and combine it with numerical feature matrix to perform fuzzy clustering, establish residual probability density distribution, and predict natural gas load range.
It achieves accuracy and stability in natural gas load range forecasting, can adapt to continuous transitions in operating conditions and changes in multiple operating conditions, and improves the coverage reliability and output stability of the forecast range.
Smart Images

Figure CN122432612A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of natural gas pipeline network operation optimization and artificial intelligence load forecasting technology, and particularly relates to a natural gas load interval estimation method based on large model coupled operating condition clustering. Background Technology
[0002] With the deepening of energy structure transformation, the proportion of natural gas in urban energy supply has been increasing year by year. Accurate natural gas load forecasting is the foundation for achieving stable pipeline network operation and emergency dispatch. Due to the influence of multiple factors such as sudden changes in meteorological conditions, complex holiday adjustment mechanisms, industrial start-up and shutdown, and long-distance pipeline storage effects, single point forecast results are difficult to meet the high reliability dispatch requirements of complex gas pipeline networks under extreme and transitional operating conditions.
[0003] Introducing forecast interval estimation methods can effectively quantify forecast uncertainties. However, existing natural gas load interval estimation techniques still face the following prominent problems:
[0004] (1) The operating conditions of natural gas systems are difficult to fully represent through simple numerical characteristics. The operating conditions of natural gas systems are not only determined by instantaneous load values and meteorological variables, but also affected by background information such as holiday arrangements, user type structure, operating rules, workday and rest day conversion methods, and upstream and downstream dispatch strategies. This information has significant semantic attributes and is difficult to characterize directly through simple numerical coding.
[0005] (2) Most existing large model applications are limited to downstream code generation or feature engineering. Although existing studies have attempted to use large models to generate feature engineering code, filter features, or assist in modeling, large models usually exist only as independent preprocessing tools and lack deep coupling with subsequent clustering and interval estimation modules, thus failing to truly solve the core problems of operating condition identification and fuzzy operating condition boundaries in natural gas load interval estimation.
[0006] (3) Traditional operating condition clustering ignores semantic operating condition relationships, which can easily lead to incorrect classification. Existing methods usually only cluster based on historical load, meteorological and other numerical characteristics, assuming that samples with similar values should be classified into the same operating condition. However, in the natural gas scenario, some samples with similar values may correspond to "pre-holiday peak rush operating conditions" and "normal weekday off-peak operating conditions" respectively, and their load evolution trends and residual distributions are significantly different; conversely, some samples with large numerical differences may have similar error formation mechanisms because they are in the same type of holiday transition state. Therefore, clustering based solely on numerical distance can easily lead to distorted operating condition classification.
[0007] (4) Natural gas operating conditions are characterized by continuous transition and ambiguous boundaries. Under scenarios such as the transition from winter to spring, the shift between day and night peak and valley, the period before and after holidays, and the triggering of extreme weather, the operating conditions of the natural gas system do not undergo hard switching, but rather transition continuously between adjacent states. Using traditional hard clustering methods will cause boundary samples to be forcibly classified into a single operating condition, resulting in abrupt changes in interval output or insufficient coverage.
[0008] (5) A uniform global residual distribution is difficult to adapt to heterogeneous errors under multiple operating conditions. The mean, variance, and tail characteristics of natural gas load forecast residuals will change with the operating conditions. If a uniform global residual distribution is used for interval estimation, it is difficult to take into account the interval reliability under both normal and extreme operating conditions.
[0009] Therefore, existing technologies urgently need a natural gas load interval estimation method that can not only introduce large models, but also allow large models to truly participate in the natural gas operating condition expression, operating condition clustering, and interval output process, thereby solving the problem of interval estimation failure under fuzzy operating conditions. Summary of the Invention
[0010] The purpose of this invention is to solve the problems existing in the prior art and to provide a method for estimating natural gas load ranges by large-scale model coupled operating condition clustering.
[0011] To achieve the above-mentioned objectives, the present invention specifically adopts the following technical solution:
[0012] In a first aspect, the present invention provides a method for estimating natural gas load intervals using large-scale model coupled operating condition clustering, comprising the following steps:
[0013] S1. Obtain historical operating data and scenario background text of the natural gas pipeline network. The historical operating data includes natural gas load data, meteorological data, and time stamps. The scenario background text includes a pipeline topology summary, a user business composition summary, a holiday arrangement description, a pipeline operation rule summary, and a scheduling constraint description. Input the working condition semantic parsing prompt words containing scenario background text, time stamps, and operating variable descriptions into the large model to instruct the large model to identify the working condition types, working condition transition relationships, and key factors affecting working condition changes in the natural gas system, thereby outputting working condition semantic constraint information.
[0014] S2. Construct a numerical feature matrix based on historical operating data, construct a clustering constraint matrix using the operating condition semantic constraint information generated by the large model, and then input the clustering constraint matrix and the numerical feature matrix into the coupled operating condition clustering model to obtain several fuzzy operating condition clusters and the membership degree of each historical operating data to each fuzzy operating condition cluster.
[0015] S3. For each fuzzy operating condition cluster, the numerical feature matrix corresponding to the historical operating data of the fuzzy operating condition cluster is used as input, and the actual value of the natural gas load data at the next moment is used as the supervision label. The natural gas load point prediction model corresponding to the fuzzy operating condition cluster is trained by supervised learning. After training, the residual between the model prediction value and the actual value at each moment is calculated. The residuals at each moment form a residual sequence. The adaptive bandwidth kernel density estimation algorithm is used to model each residual sequence to obtain the residual probability density distribution corresponding to each fuzzy operating condition cluster.
[0016] S4. For each prediction time, acquire the real-time operation data of the natural gas pipeline network and obtain its semantic matching degree and numerical membership degree relative to each fuzzy operation condition cluster. Then, fuse the semantic matching degree and numerical membership degree to obtain the real-time coupled membership weight. After the natural gas load point prediction model corresponding to each fuzzy operation condition cluster outputs the candidate point prediction value for the prediction time, extract the corresponding upper and lower quantiles of the residuals under the preset confidence level according to the residual probability density distribution established in S3. Combine the real-time coupled membership weight to dynamically and continuously weight the candidate point prediction value and the upper and lower quantiles of the residuals corresponding to each fuzzy operation condition cluster, thereby outputting the upper and lower bounds of the natural gas load prediction interval to form the natural gas load prediction interval. Then, the candidate point prediction values of each fuzzy operation condition cluster are weighted and fused according to the real-time coupled membership weight to obtain the final point prediction value for the prediction time, which is used together with the natural gas load prediction interval as the natural gas load prediction result.
[0017] Based on the above scheme, each step can be implemented in the following preferred manner.
[0018] As a preferred embodiment of the first aspect mentioned above, in step S1, the semantic constraint information of the working conditions output by the large model includes at least one or more of the following: a set of working condition type labels, working condition transition relationships, and working condition constraint rules; the set of working condition type labels includes weekday working conditions, pre-holiday peak-loading working conditions, post-holiday recovery working conditions, low-temperature high-load working conditions, and extreme rainfall disturbance working conditions; the working condition constraint rules include semantic similarity constraints that prioritize classifying semantically similar historical operating data into the same fuzzy operating condition cluster, semantic exclusivity constraints that prevent semantically exclusive historical operating data from being classified into the same fuzzy operating condition cluster, and adjacency relationship constraints between working conditions.
[0019] As a preferred embodiment of the first aspect above, in step S2, the numerical feature matrix includes at least one or more of the following: historical load time lag features, sliding statistical features, meteorological features, periodic features, and holiday offset features.
[0020] As a preferred embodiment of the first aspect above, the historical load time lag feature is based on the current natural gas load data corresponding to a certain time lag before that time, selecting natural gas load data at several lag times; the sliding statistical feature is a statistical quantity calculated based on a preset time window for natural gas load data or meteorological data, the statistical quantity including one or more of the following: load sliding mean, temperature sliding mean, load sliding variance, and cumulative rainfall sliding value; the meteorological feature includes one or more of the following: ambient temperature, rainfall, relative humidity, wind speed, and perceived temperature; the periodic feature is a time-based feature constructed based on the position of the historical operating data time in the hourly, daily, weekly, monthly, or other periods; the holiday offset feature refers to the feature characterizing the relative distance between the current historical operating data and the start or end time of a holiday.
[0021] As a preferred embodiment of the first aspect mentioned above, in step S2, each element of the clustering constraint matrix represents the semantic constraint relationship between two historical operation data. The specific method for constructing the clustering constraint matrix from the semantic constraint information of the operating conditions is as follows: According to the analysis results of the large model, when it is identified that two historical operation data correspond to the same type of operating condition semantics or that the operating conditions corresponding to the two historical operation data can be classified into the same operating stage in terms of business, then these two historical operation data are regarded as semantically similar historical operation data, and the element value of the corresponding position in the clustering constraint matrix is assigned a positive value; when it is identified that two historical operation data correspond to different types of operating condition semantics or that the operating conditions corresponding to the two historical operation data cannot be classified into the same operating stage in terms of business, then these two historical operation data are regarded as semantically contradictory historical operation data, and the element value of the corresponding position in the clustering constraint matrix is assigned a negative value; if no semantic constraint relationship between two historical operation data is identified, then the element value of the corresponding position in the clustering constraint matrix is assigned a value of 0.
[0022] As a preferred embodiment of the first aspect, in the coupled operating condition clustering model of step S2, the fuzzy C-means clustering algorithm is used to perform soft clustering on the historical operating data, and Mahalanobis distance is used as the distance metric between the historical operating data and the cluster centers. At the same time, a semantic constraint penalty term based on the clustering constraint matrix is introduced into the objective function of the fuzzy C-means clustering algorithm, so that the final fuzzy operating condition clusters can simultaneously satisfy numerical similarity and operating condition semantic consistency. The Mahalanobis distance in the objective function is calculated based on the numerical feature matrix of the historical operating data and the cluster centers of the fuzzy operating condition clusters. The semantic constraint penalty term is calculated based on the clustering constraint matrix and the membership distribution of the historical operating data on each fuzzy operating condition cluster.
[0023] Furthermore, the objective function adopted... The specific form is as follows:
[0024]
[0025] in, The total number of historical data points that participated in the clustering; The number of fuzzy operating condition clusters; For the first The historical operational data belongs to the first Membership degree of a fuzzy operating condition cluster; It is a fuzzy weighted index; for of Power of; It is a fuzzy weighted index; for and Mahalanobis distance between them; For the first Numerical feature matrix of historical running data; For the first Cluster centers of a fuzzy operating condition cluster; This is a semantic constraint penalty term; These are the semantic constraint weight coefficients.
[0026] Furthermore, the formula for calculating the semantic constraint penalty term is as follows:
[0027]
[0028] in, The first in the clustering constraint matrix Line number Column elements; and They represent the first The historical operational data for the first The and the first The membership degree of a fuzzy operating condition cluster.
[0029] As a preferred embodiment of the first aspect mentioned above, the specific process of step S4 is as follows:
[0030] S41. For each prediction time, acquire the real-time operation data of the natural gas pipeline network, replace the historical operation data in S2 with the real-time operation data, and construct the real-time numerical feature matrix according to step S2 and obtain the numerical membership degree of the real-time operation data to each fuzzy operation condition cluster; at the same time, acquire the real-time scene semantic description corresponding to the prediction time, and then input the real-time scene semantic description and the pre-acquired real-time operation variable description into the large model to obtain the semantic matching degree of the real-time operation data relative to each fuzzy operation condition cluster.
[0031] S42. Then, the numerical membership degree and semantic matching degree are fused by exponentiation and normalized to obtain the real-time coupled membership weight.
[0032] S43. Input the real-time numerical feature matrix into the pre-trained natural gas load point prediction model corresponding to each fuzzy operating condition cluster to obtain the candidate point prediction value under each fuzzy operating condition cluster at the prediction time.
[0033] S44. For each fuzzy operating condition cluster, first obtain the corresponding cumulative distribution function by integration, and then extract the upper quantile and lower quantile of the residual from the cumulative distribution function under a preset confidence level.
[0034] S45. Add the predicted values of candidate points for each fuzzy operating condition cluster to the corresponding upper quantile of the residual, and use this as the upper bound prediction result. Then, use real-time coupled membership weights to weight the predicted values, and use this as the weighted upper bound prediction result for that fuzzy operating condition cluster. Finally, sum the weighted upper bound prediction results of all fuzzy operating condition clusters to obtain the upper bound of the output natural gas load prediction interval. Add the predicted values of candidate points for each fuzzy operating condition cluster to the corresponding lower quantile of the residual, and use this as the lower bound prediction result. Then, use real-time coupled membership weights to weight the predicted values, and use this as the weighted lower bound prediction result for that fuzzy operating condition cluster. Finally, sum the weighted lower bound prediction results of all fuzzy operating condition clusters to obtain the lower bound of the output natural gas load prediction interval. The natural gas load prediction interval is formed by the upper and lower bounds of the natural gas load prediction interval.
[0035] S46. The candidate point prediction values of each fuzzy operating condition cluster are weighted and fused according to the real-time coupling membership weight to obtain the final point prediction value at the prediction time, and together with the natural gas load prediction interval, it is used as the natural gas load prediction result.
[0036] Furthermore, in step S42, the first Real-time coupling membership weights corresponding to each fuzzy operating condition cluster The calculation formula is:
[0037]
[0038] in, Indicates the total number of fuzzy operating condition clusters; For the first Numerical membership degree of a fuzzy operating condition cluster; For the first Numerical membership degree of a fuzzy operating condition cluster; For the first The semantic matching degree of a fuzzy operating condition cluster; For the first The semantic matching degree of a fuzzy operating condition cluster; and These are the adjustment parameters corresponding to numerical membership degree and semantic matching degree, respectively.
[0039] Furthermore, in step S45, the upper and lower bounds of the natural gas load forecast interval are respectively expressed as:
[0040]
[0041]
[0042] in, , , and The first Real-time coupled membership weights, candidate point prediction values, upper quantiles, and lower quantiles of residuals corresponding to each fuzzy operating condition cluster; That is, the first The upper bound prediction results of a fuzzy operating condition cluster. That is, the first The weighted upper bound prediction results of a fuzzy operating condition cluster That is, the first Lower bound prediction results for a fuzzy operating condition cluster That is, the first The weighted lower bound prediction results of a fuzzy operating condition cluster.
[0043] In a second aspect, the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, enables the natural gas load range estimation method based on large model coupled operating condition clustering as described in any of the solutions in the first aspect above.
[0044] Thirdly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a natural gas load range estimation method based on large-model coupled operating condition clustering as described in any of the solutions of the first aspect above.
[0045] Fourthly, the present invention provides a computer electronic device, which includes a memory and a processor;
[0046] The memory is used to store computer programs;
[0047] The processor is configured to, when executing the computer program, implement the natural gas load range estimation method based on large model coupled operating condition clustering as described in any of the schemes of the first aspect above.
[0048] Compared with existing technologies, the natural gas load interval estimation method based on large-model coupled operating condition clustering described in this invention has the following advantages:
[0049] Compared with traditional interval estimation methods based solely on numerical feature clustering, this invention does not use the large model as a simple feature generation or code generation tool. Instead, it utilizes the large model to extract the semantic structure information of natural gas operating conditions and transforms it into clustering constraints that are directly embedded in the operating condition partitioning process, thereby achieving deep coupling of "large model - operating condition clustering - interval output".
[0050] Compared to schemes that use large models for downstream feature derivation, the large model of this invention addresses the more critical issues of operating condition identification and operating condition boundary continuity in natural gas load range estimation.
[0051] Compared with traditional hard clustering or pure numerical fuzzy clustering methods, the present invention enables the obtained fuzzy operating condition clusters to simultaneously satisfy numerical similarity and semantic consistency, and can better reflect the true formation mechanism of natural gas load residuals.
[0052] Compared with the method of using a globally unified residual distribution, this invention establishes residual probability density distributions for each fuzzy operating condition cluster and performs continuous fusion based on real-time coupled membership weights. This enables it to better adapt to the fuzzy operating condition changes of the natural gas system under holiday switching, peak-valley transition and extreme weather conditions, and significantly improves the coverage reliability, interval compactness and output stability of the prediction interval. Attached Figure Description
[0053] Figure 1 This is a flowchart of the method of the present invention;
[0054] Figure 2 This is a schematic diagram of the coupled working condition clustering process in S2 of the present invention;
[0055] Figure 3 This is a schematic diagram of the process of outputting natural gas load prediction results in S4 of the present invention;
[0056] Figure 4 This is a schematic diagram of a computer electronic device provided by the present invention. Detailed Implementation
[0057] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below. Technical features in the various embodiments of the present invention can be combined accordingly without mutual conflict.
[0058] In the description of this invention, it should be understood that the terms "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first" and "second" may explicitly or implicitly include at least one of those features.
[0059] like Figure 1 As shown, in a preferred embodiment of the present invention, the natural gas load range estimation method based on large model coupled operating condition clustering includes the following steps S1 to S4. The specific implementation process of each step is described in detail below.
[0060] I. Semantic Parsing and Constraint Generation of Natural Gas Operating Conditions Based on Large Models
[0061] S1. Obtain historical operating data and scenario background text of the natural gas pipeline network. The historical operating data includes natural gas load data, meteorological data, and time stamps. The scenario background text includes a pipeline topology summary, a user business composition summary, a holiday arrangement description, a pipeline operation rule summary, and a scheduling constraint description. Input the working condition semantic parsing prompt words containing scenario background text, time stamps, and operating variable descriptions into the large model, instructing the large model to identify the working condition types, working condition transition relationships, and key factors affecting working condition changes in the natural gas system, thereby outputting working condition semantic constraint information.
[0062] It should be noted that, in step S1 of the present invention, the semantic constraint information of the working conditions output by the large model includes at least one or more of the following: a set of working condition type labels, a set of working condition transition relationships, and working condition constraint rules; the set of working condition type labels includes weekday working conditions, pre-holiday peak-loading working conditions, post-holiday recovery working conditions, low-temperature high-load working conditions, and extreme rainfall disturbance working conditions; the working condition constraint rules include semantic similarity constraints that prioritize classifying semantically similar historical operating data into the same fuzzy operating condition cluster, semantic exclusivity constraints that avoid classifying semantically exclusive historical operating data into the same fuzzy operating condition cluster, and adjacency relationship constraints between working conditions.
[0063] The following explains some concepts involved in this invention. In step S1, the scenario background text refers to textual information used to describe the operating environment, gas consumption scenarios, operating rules, and scheduling conditions of the natural gas pipeline network. This information is not directly represented as a single numerical variable, but it will affect the natural gas load change pattern, operating condition division, and operating condition evolution relationship. The scenario background text can originate from operating procedures, scheduling records, business instruction documents, manually compiled text, or descriptive text obtained through structured data conversion.
[0064] A pipeline topology summary is a textual description of the connections and supply structure of a natural gas pipeline network. Its content may include, but is not limited to: the number of gate stations, the connection between main pipelines and branches, the branches where key users are located, the direction of gas supply, the regional gas supply relationship, upstream and downstream connection methods, and the existence of backup channels. The purpose of the pipeline topology summary is to enable large models to understand the gas supply correlation between different regions and the potential impact of local load changes on overall operating conditions. For example, in a specific implementation scenario, it could be described as: "There are 8 branches downstream of this gate station, of which industrial users are mainly concentrated in branches 1 and 2, and residential users are mainly concentrated in branches 5 to 8. Some branches have interconnection and switching capabilities."
[0065] User type composition summary refers to textual information describing the types of natural gas users and their proportions. Its content may include the composition ratio of residential users, industrial users, commercial users, public welfare users, and combined heat and power (CHP) users, their main gas consumption patterns, and their impact on load fluctuations. This information helps large-scale models determine whether load fluctuations during a given period are more likely to be caused by factors such as residential heating, industrial production, commercial operations, and holiday shutdowns or resumptions. For example, it could be stated as: "Residential users account for a relatively high proportion in this area, with significant morning and evening peaks during the low temperatures of winter; industrial users tend to increase production before holidays and gradually recover afterward."
[0066] Holiday schedule descriptions refer to text that describes time system information such as statutory holidays, weekends, work-rest arrangements, and the distribution of dates before and after holidays. They may include the holiday name, start and end dates, adjusted work-rest dates, adjacency with normal working days, and the number of days offset from the start or end of the holiday. Since natural gas load is often affected by holiday schedules, return-to-work dates, and changes in business activities, holiday schedule descriptions are used by large models to identify operating states with temporal semantic differences, such as "working day conditions," "pre-holiday peak-loading conditions," and "post-holiday recovery conditions."
[0067] A pipeline operation rule summary is a text that provides a general description of the business rules, control logic, and empirical patterns followed in the daily operation of a natural gas pipeline network. Its content may include: peak-hour supply principles, pressure control requirements, priority gas supply strategies for key users, nighttime peak-shaving methods, principles for responding to extreme weather, valve switching restrictions, etc. This summary is used to guide large-scale models to understand operating condition boundaries and transition conditions from the perspective of industry operating rules, rather than relying solely on numerical fluctuations themselves.
[0068] Dispatch constraint descriptions are textual descriptions of the restrictions that a natural gas system must meet during dispatch operations. These may include constraints such as upper limits on gas supply, upper and lower limits on pipeline pressure, start-up and shutdown constraints on peak-shaving equipment, limits on gas transfers between upstream and downstream suppliers, and restrictions on uninterrupted service for key users. Dispatch constraint descriptions help large-scale models understand that certain load changes are not solely caused by natural demand variations but may also be influenced by gas supply capacity, dispatch strategies, or safety boundaries.
[0069] The description of operating variables refers to the textual expression of the main operating variables involved in the judgment of operating conditions and their changing states. These operating variables may include natural gas load, temperature, rainfall, humidity, wind speed, time stamps, pre-holiday / post-holiday indicators, historical load lag terms, and sliding statistics. Unlike directly inputting raw numerical values, the description of operating variables in this invention can use interpretive expressions such as "the average temperature of the day is significantly lower than the average of the previous three days," "the current time is two days before the holiday," and "the load is significantly higher than the same period of the previous day," to enhance the large model's ability to recognize the semantics of operating conditions.
[0070] The operating condition type label set refers to the label set formed by the large model after semantically classifying the operating status of the natural gas system based on the scenario background text, time labels, and descriptions of operating variables. This set is used to characterize the typical operating condition category to which historical operating data may belong. The operating condition type label set can be one or more of the preset candidate labels, or it can be an extended label output by the large model under the naming rules given in this invention. Preferably, the operating condition type label set includes at least one of the following: weekday operating condition, pre-holiday peak rush operating condition, post-holiday recovery operating condition, low temperature and high load operating condition, and extreme rainfall disturbance operating condition. It should be noted that the above operating condition labels are not intended to limit this invention to only these few types of operating conditions, but rather to illustrate the expression form of the large model's output results as an implementable example.
[0071] Specifically, the weekday operating condition in the set of operating condition types refers to the operating state corresponding to non-holidays, non-major abnormal weather, and when users' gas consumption behavior basically follows the normal production and life rhythm. This operating condition is usually characterized by industrial, commercial, and residential demand fluctuating according to normal patterns, and the load curve having relatively stable intraday periodicity and intraweekly regularity.
[0072] Pre-holiday peak-load operation refers to the operational state in which the natural gas load experiences a phased increase or structural change during the period before the start of statutory holidays or concentrated vacations due to factors such as industrial enterprises rushing to complete tasks, commercial stockpiling, and residents' early activities.
[0073] Post-holiday recovery refers to the transitional operational state that occurs after a holiday, as industry resumes production, businesses reopen, and residents gradually return to normal activities. This state typically does not return to normal workday levels instantaneously, but rather involves a gradual recovery, fluctuations, or redistribution over several days. Therefore, semantically, "post-holiday recovery" is a transitional operational state.
[0074] Low-temperature high-load operating conditions refer to a high-load operating state that occurs against the backdrop of a significant drop in temperature, an increase in heating demand, or an increase in residential gas consumption. This condition is usually strongly correlated with meteorological factors such as temperature, wind speed, and humidity, and is particularly pronounced in areas with a high proportion of residential users during winter.
[0075] Extreme rainfall disturbance conditions refer to the operating conditions under extreme rainfall weather conditions such as short-term heavy rainfall, continuous heavy rainfall, or regional rainstorms, where natural gas load exhibits sudden fluctuations, local abnormal changes, or significant deviations from the normal pattern due to the combined effects of factors such as changes in temperature and humidity, adjustments in users' residential and commercial gas consumption behaviors, upstream and downstream supply and demand disturbances, and changes in the pipeline network operating environment.
[0076] Operating condition transition relationships refer to the connections, adjacencies, or transformations between different operating condition types over time. This can indicate that a particular operating condition is more likely to transition to another under specific conditions, or that two operating conditions are temporally adjacent, have contiguous boundaries, or have a continuous transition region. In this invention, operating condition transition relationships are primarily used to guide the clustering stage in making more reasonable soft partitions of historical operational data at the boundaries, rather than requiring a unique and definite hard switch between operating conditions.
[0077] Operating condition constraint rules refer to a set of rules formed based on the semantic understanding results of operating conditions output by the large model, which can be used to constrain the clustering process. They are used to transform semantic information into formal constraints that can be used in subsequent coupled operating condition clustering.
[0078] A fuzzy operating condition cluster refers to a set of historical operating data formed under a fuzzy clustering framework. Unlike traditional hard clustering, where a historical operating data point can only belong to one cluster, in this invention, each historical operating data point can have different degrees of membership to multiple operating condition clusters, thereby characterizing the continuity and fuzziness of natural gas operating conditions in the boundary region.
[0079] Semantically contradictory historical operating data refers to historical operating data whose reflected operating conditions or load formation mechanisms are significantly different based on a comprehensive judgment of the scenario background text, temporal semantics, operating variables, and operating rules, and are not suitable for priority classification into the same operating condition cluster. Semantically similar historical operating data refers to historical operating data whose corresponding historical operating data may not be completely identical in some numerical characteristics, but whose operating condition meaning, load formation background, or evolution mechanism is similar based on a comprehensive judgment of the scenario background text, temporal semantics, direction of change of operating variables, and business rules.
[0080] In this embodiment, the specific form of the semantic parsing prompts for the working conditions input to the large model is as follows:
[0081] "Please act as an assistant for natural gas pipeline network operation analysis, combining the following information to identify the operating condition type of the current historical operating data, and summarize the transitional relationships and clustering constraints between operating conditions. The known scenario background text is as follows: 1) Pipeline topology summary: There are 8 branches downstream of this gate station, of which branches 1 and 2 are mainly for industrial users, and branches 5 to 8 are mainly for residential and commercial users. Some branches have interconnection and switching capabilities; 2) User type composition summary: Residential users account for a relatively high proportion, with significant morning and evening peaks during the winter low temperatures; Industrial users often increase production before holidays, and there is a gradual recovery process after the holidays; 3) Holiday arrangement description: The current date is 2 days before the Spring Festival, followed by a 7-day consecutive holiday; 4) Pipeline operation..." Operational rules summary: During peak hours, priority is given to residential gas supply, while industrial users can be flexibly adjusted when necessary; 5) Dispatch constraint description: There is a daily upper limit to the gas supply capacity of the gate station, and the stability of branch pressure needs to be monitored during morning and evening peak hours. The time label and operation variable description of historical operation data are as follows: 1) Date type: Weekday; 2) Holiday offset: 2 days before the start of the Spring Festival; 3) Time period: 18:00; 4) Meteorological conditions: The average temperature of the day is 4 degrees Celsius lower than the previous 3 days, with no significant rainfall; 5) Load change: The current load is 8% higher than the same period yesterday and 10% higher than the same period last week; 6) Historical trend: The load has been rising continuously in the past 24 hours, with a significant increase in industrial-related branches. Please complete the following task: A. A. Determine the possible operating condition type label corresponding to the historical operating data; B. Provide the operating condition types adjacent to or potentially transitioning to this operating condition; C. Explain the key factors affecting this operating condition judgment; D. Output clustering constraint suggestions, including: which types of historical operating data should be considered semantically similar; which types of historical operating data should be considered semantically exclusive; whether there are adjacency constraints; whether there are transition direction constraints; E. Output the results in the following JSON format: {"Operating Condition Type Label": [],"Key Influencing Factors": [],"Adjacent Operating Conditions": [],"Transition Direction": [],"Semantically Similar Rules": [],"Semantically Exclusive Rules": [],"Adjacency Constraints": [],"Boundary Hints": []}.
[0082] After inputting the above-mentioned semantic analysis prompts for operating conditions into the large model, in addition to the semantic constraints of the operating conditions mentioned above, the large model will also output operating condition boundary prompts for holidays, day-night peak-valley transitions, and extreme weather triggering conditions.
[0083] II. Numerical Feature Construction and Coupled Working Condition Clustering
[0084] S2. Construct a numerical feature matrix based on historical operating data, construct a clustering constraint matrix using the operating condition semantic constraint information generated by the large model, and then input the clustering constraint matrix and the numerical feature matrix into the coupled operating condition clustering model to obtain several fuzzy operating condition clusters and the membership degree of each historical operating data to each fuzzy operating condition cluster.
[0085] It should be noted that in step S2 of the present invention, the numerical feature matrix includes at least one or more of the following: historical load time lag features, sliding statistical features, meteorological features, periodic features, and holiday offset features.
[0086] In this embodiment, the historical load time lag feature refers to the natural gas load data selected at several lag times prior to the current natural gas load data, based on the current natural gas load data, to characterize the load time-series dependency. Preferably, the historical load time lag feature includes natural gas load data corresponding to the previous 1 hour, 24 hours, 48 hours, and 72 hours from the current time. Specifically, the natural gas load data for the previous 1 hour is used to characterize the short-term continuity of the natural gas load, the natural gas load data for the previous 24 hours and 48 hours are used to characterize the daily cycle repetition pattern, and the natural gas load data for the previous 72 hours is used to supplement and reflect the cross-day continuation effect. If the current natural gas load data... The corresponding time is recorded as The historical load time lag characteristics can then be expressed as: , ... , respectively corresponding to the previous Real-time natural gas load data, previous Real-time natural gas load data, ..., previous Natural gas load data at any given time. Among them, The preset lag time length is used. This type of feature extraction follows conventional time-series feature construction methods in this field, and can be directly obtained from the natural gas load data time series by time index. For moments with missing values, preprocessing can be performed using methods such as interpolation between adjacent moments, substitution within the same period, or removal of outlier samples.
[0087] In this embodiment, the sliding statistical feature refers to a statistical quantity calculated based on natural gas load data or meteorological data within a preset time window, used to reflect the recent fluctuation level, trend, and cumulative effect of historical operating data. The statistical quantity includes one or more of the following: load sliding mean, temperature sliding mean, load sliding variance, and rainfall sliding cumulative value. Specifically, the load sliding mean can be obtained from the average value of natural gas load data within the preset window before the current time, used to reflect the recent overall gas consumption level; the load sliding variance can be obtained from the degree of fluctuation of natural gas load data within the preset window before the current time, used to reflect recent load stability; the temperature sliding mean can be obtained from the average temperature over several hours or days before the current time, used to reflect a background of continuous cooling or warming; and the rainfall sliding cumulative value can be obtained by accumulating the rainfall within the preset window before the current time, used to reflect the lagging effect of continuous rainfall on operating conditions. Preferably, if the current time... The previous length was If the sliding time window is used as the statistical interval, then the interval can be used for statistical analysis. The invention calculates corresponding statistics based on natural gas load data or meteorological data. Examples include the 3-hour moving average of load; the 6-hour moving average of load; the 24-hour moving variance of load; and the 3-day moving average of temperature. The extraction method of the moving statistical features is a conventional time window statistical method in this field. This invention is not limited to a specific window length; the window length can be set according to the gate station load change pattern, the prediction time scale, and the validation set effect.
[0088] In this embodiment, the meteorological features include one or more of the following: ambient temperature, rainfall, relative humidity, wind speed, perceived temperature, or other meteorological variables affecting natural gas demand. These features are primarily acquired directly through meteorological observation systems, meteorological forecasting systems, or third-party meteorological data interfaces, and are aligned with natural gas load data by timestamp. Preferably, ambient temperature and rainfall are used as the basic meteorological features. Specifically, the temperature feature can directly use the measured temperature value corresponding to the current moment or the forecast temperature value; the rainfall feature can directly use the observed rainfall value corresponding to the current moment or the forecast rainfall amount. Furthermore, to enhance the ability of meteorological factors to represent changes in operating conditions, the following derived meteorological features can be constructed: the difference between the current temperature and the temperature at the same time the previous day; the difference between the current temperature and the average temperature of the past 3 days; a binary indicator of whether low-temperature weather has occurred; and a binary indicator of whether heavy rainfall weather has occurred. These binary indicators can be generated based on preset thresholds. For example, when the temperature is below a set threshold, it can be marked as low-temperature weather; when the rainfall per unit time is above a set threshold, it can be marked as heavy rainfall weather. These threshold settings can be determined based on local climate characteristics, operational experience, or historical statistical results. The acquisition and time alignment of meteorological features are conventional techniques in this field. The focus of this invention is to couple these meteorological features with time semantic features and operating condition semantic constraints.
[0089] In this embodiment, the periodic feature refers to a time-based feature constructed based on the position of the historical operating data's time within a cycle such as hour, day, week, or month, used to characterize the repetitive variation pattern of natural gas load over a fixed time period. Preferably, the periodic feature includes one or more of the following: hourly features, weekday features, monthly features, weekday / weekend identifiers, and periodic coding features. Specifically: the hourly feature can be directly obtained from the hour number corresponding to the historical operating data time, for example, a value from 0 to 23; the weekday feature can be obtained from the week number corresponding to the historical operating data date, for example, 1 to 7; the monthly feature can be obtained from the month number corresponding to the historical operating data date, for example, 1 to 12; the weekday / weekend identifier can be determined based on calendar information, with 1 for workdays and 0 for non-workdays, or by using the opposite coding method; the periodic coding feature can be obtained by mapping the hour, weekday, or month using a periodic function to enhance the model's ability to express the continuity of the beginning and end of the cycle. Preferably, the hourly and weekday features can be sine-cosine coded to obtain hourly and weekday periodic codes. If the hour number is denoted as... Then the hourly cycle code can be represented as: If the week number is recorded as Then the weekday cycle code can be represented as The purpose of using the above encoding method is to avoid incorrectly treating the beginning and end positions of a periodicity as being far apart. For example, 23:00 and 0:00 are numerically adjacent, but their differences are significant under linear encoding. Periodic encoding can more reasonably represent their temporal proximity. The method of constructing periodic features belongs to the conventional time feature extraction methods in this field.
[0090] In this embodiment, the holiday offset feature refers to the feature that characterizes the relative distance between the current historical operating data and the start or end time of the holiday. It is used to reflect the impact on natural gas load during the pre-holiday, mid-holiday, and post-holiday periods. It includes one or more of the following: whether it is a holiday, whether it is a pre-holiday period, whether it is a post-holiday period, the number of days until the start of the next holiday, the number of days until the end of the previous holiday, the number of days offset before the holiday, the number of days offset after the holiday, and the indicator of working on adjusted work days.
[0091] Preferably, the holiday attribute corresponding to each historical operating data date can be determined first based on the statutory holiday calendar, local holiday adjustment arrangements, and the company's operating calendar. Then, the time offset between the historical operating data and the recent holiday boundary can be calculated to construct a holiday offset feature. Specifically, the binary identifier of whether it is a holiday is 1 if the date corresponding to the current historical operating data is a holiday, and 0 otherwise; the binary identifier of whether it is a pre-holiday period is 1 if the date corresponding to the current historical operating data is within a preset pre-holiday window, and 0 otherwise; the binary identifier of whether it is a post-holiday period is 1 if the date corresponding to the current historical operating data is within a preset post-holiday window, and 0 otherwise; the number of days until the start of the next holiday is calculated based on the difference between the date corresponding to the current historical operating data and the start date of the next holiday; the number of days until the end of the previous holiday is calculated based on the difference between the date corresponding to the current historical operating data and the end date of the previous holiday; the pre-holiday offset days and the post-holiday offset days can be preset to any range from 1 to 7 days based on business experience. For example, the 3 days before the holiday can be defined as the pre-holiday sensitive interval, and the 3 days after the holiday can be defined as the post-holiday recovery interval. Preferably, for dates with adjusted workday arrangements, the pre-holiday and post-holiday offset days can be identified using calendar rules, and adjusted workday identifiers can be constructed to avoid simply treating these dates as ordinary workdays or ordinary rest days. It should be noted that although the holiday offset features can be directly calculated from calendar information, their role in this invention is not only as a general time feature input, but also further resonates with the semantic labels such as pre-holiday peak-loading conditions and post-holiday recovery conditions obtained from the large model in step S1, thereby jointly supporting subsequent coupled work condition clustering.
[0092] Furthermore, after obtaining the aforementioned historical load time lag characteristics, sliding statistical characteristics, meteorological characteristics, periodic characteristics, and holiday offset characteristics, the various characteristics are uniformly aligned according to the corresponding time of the historical operating data, and then concatenated column by column to form a numerical feature matrix. If there are a total of Each historical operational data point is extracted. If the numerical features are dimensional, then the numerical feature matrix is... It can be represented as: Each row represents a numerical feature vector corresponding to a historical data point, and each column represents a type of numerical feature. Preferably, after forming the numerical feature matrix, each feature can be standardized, normalized, or subjected to outlier processing to reduce the impact of dimensional differences and abnormal observations on the clustering results. This processing method is a conventional data preprocessing method in this field and will not be elaborated further here.
[0093] It should be noted that in step S2 of this invention, the aforementioned clustering constraint matrix is used to guide the clustering process by "aggregating semantically similar historical operating data" and "separating semantically contradictory historical operating data." The clustering constraint matrix is constructed from the operating condition semantic constraint information obtained in step S1 and is used to characterize the proximity, exclusion, or unconstrained relationship between different historical operating data in terms of operating condition semantics. Each element of the clustering constraint matrix represents the semantic constraint relationship between two historical operation data. The specific method for constructing the clustering constraint matrix from the semantic constraint information of the operating conditions is as follows: Based on the analysis results of the large model, when it is identified that two historical operation data correspond to the same type of operating condition semantics or that the operating conditions corresponding to the two historical operation data can be classified into the same operating stage in terms of business, then these two historical operation data are regarded as semantically similar historical operation data, and the element value of the corresponding position in the clustering constraint matrix is assigned a positive value (such as +1); when it is identified that two historical operation data correspond to different types of operating condition semantics or that the operating conditions corresponding to the two historical operation data cannot be classified into the same operating stage in terms of business, then these two historical operation data are regarded as semantically contradictory historical operation data, and the element value of the corresponding position in the clustering constraint matrix is assigned a negative value (such as -1); if no semantic constraint relationship is identified between two historical operation data, then the element value of the corresponding position in the clustering constraint matrix is assigned a value of 0. In other embodiments, different positive or negative weights can also be assigned according to the confidence level, semantic strength, or importance of business rules output by the large model.
[0094] It should be noted that in the coupled working condition clustering model of step S2 of this invention, as follows: Figure 2As shown, a fuzzy C-means clustering algorithm is used to perform soft clustering on historical operating data, and Mahalanobis distance is used as the distance metric between historical operating data and cluster centers. A semantic constraint penalty term based on the clustering constraint matrix is introduced into the objective function of the fuzzy C-means clustering algorithm, ensuring that the final fuzzy operating condition clusters simultaneously satisfy numerical similarity and semantic consistency. The Mahalanobis distance in the objective function is calculated based on the numerical feature matrix of historical operating data and the cluster centers of the fuzzy operating condition clusters; the semantic constraint penalty term is calculated based on the clustering constraint matrix and the membership distribution of historical operating data in each fuzzy operating condition cluster.
[0095] In this embodiment, unlike hard clustering which assigns each historical operating data point to only one category, the soft clustering used in this invention allows the same historical operating data point to have different membership degrees with multiple fuzzy operating condition clusters simultaneously. This better reflects the transitional states and boundary ambiguities inherent in natural gas operating conditions in real-world scenarios. Furthermore, this embodiment replaces the Euclidean distance metric in the fuzzy C-means clustering algorithm with Mahalanobis distance. Compared to Euclidean distance, Mahalanobis distance comprehensively considers the scale differences and correlations between various numerical features, thereby reducing the interference of feature correlations on the clustering results and improving the rationality of distance calculations between historical operating data points under different operating conditions. Further, this invention introduces a semantic constraint penalty term into the objective function, ensuring that the clustering optimization process considers not only the proximity of historical operating data points in the numerical feature space but also the semantic similarity or repulsion between historical operating data points.
[0096] Specifically, in this embodiment, the objective function used is... The specific form is as follows:
[0097]
[0098] in, The total number of historical data points that participated in the clustering; The number of fuzzy operating condition clusters; For the first The historical operational data belongs to the first Membership degree of a fuzzy operating condition cluster; This is a fuzzy weighting index used to adjust the degree of fuzziness in the clustering results; for of The exponentiation is used to reflect the weighting effect of historical operating data with different membership levels on the objective function; This is a fuzzy weighting index used to adjust the degree of fuzziness in the clustering results. It can be set according to the distribution characteristics of historical operating data and clustering requirements; for and Mahalanobis distance between them; For the first Numerical feature matrix of historical running data; For the first Cluster centers of a fuzzy operating condition cluster; This is a semantic constraint penalty term; The semantic constraint weight coefficient is used to adjust the relative influence of the numerical clustering term and the semantic constraint penalty term in the objective function. The value of this coefficient can be preset empirically or determined through parameter optimization based on the clustering results of historical data. When the value is small, the clustering results focus more on the similarity of historical data in the numerical feature space; when... When the value is large, the clustering results focus more on satisfying the semantic constraints of the working conditions.
[0099] Furthermore, the formula for calculating the semantic constraint penalty term is as follows:
[0100]
[0101] in, The first in the clustering constraint matrix Line number Column elements, used to represent the first The fuzzy operating condition cluster and the first The semantic constraint strength between fuzzy operating condition clusters; and They represent the first The historical operational data for the first The and the first The membership degree of each fuzzy operating condition cluster. By constraining the differences in membership degree of the same historical operating data across different fuzzy operating condition clusters, the clustering results can better conform to the preset semantic relationships of operating conditions.
[0102] Furthermore, the formula for calculating the Mahalanobis distance is as follows:
[0103]
[0104] in, This is the matrix transpose. The covariance matrix of the numerical characteristic matrix; superscript To find the inverse of the matrix.
[0105] III. Fuzzy Operating Condition Cluster Intra-point Prediction Modeling and Conditional Residual Distribution Modeling
[0106] S3. For each fuzzy operating condition cluster, the numerical feature matrix corresponding to the historical operating data of that fuzzy operating condition cluster is used as input, and the actual value of the natural gas load data at the next moment is used as the supervision label. The natural gas load point prediction model corresponding to the fuzzy operating condition cluster is trained by supervised learning. After training, the residual between the model prediction value and the actual value at each moment is calculated. The residuals at each moment constitute a residual sequence, and the adaptive bandwidth kernel density estimation algorithm is used to model each residual sequence to obtain the residual probability density distribution corresponding to each fuzzy operating condition cluster.
[0107] It should be noted that in step S3 of this invention, the natural gas load point prediction model can be a deep learning model (such as LSTM / GRU / BP network) or other regression models. In this embodiment, a Long Short-Term Memory (LSTM) network is specifically used as the natural gas load point prediction model.
[0108] Furthermore, since fuzzy clustering is used in step S2 to obtain the membership degree of each historical operating data to each fuzzy operating condition cluster, in step S3, when performing point prediction modeling for a certain fuzzy operating condition cluster, the membership degree of the historical operating data to that fuzzy operating condition cluster can be used as sample weights for training; alternatively, historical operating data with a membership degree greater than a preset threshold can be selected to form a training dataset, thereby training the natural gas load point prediction model corresponding to that fuzzy operating condition cluster. During the training process, the model outputs the predicted value of the natural gas load data, and calculates the loss function based on the error between the predicted value and the actual value. Then, the model parameters are iteratively updated based on the loss function until the model training is completed.
[0109] It should be noted that in step S3 of the present invention, for the first... The first fuzzy operating condition cluster, based on historical operating data, is used for the... The membership degree of a fuzzy operating condition cluster is used to determine the relationship with the first fuzzy operating condition cluster. The residual sequences corresponding to each fuzzy operating condition cluster are used to calculate the residual probability density distribution of the cluster based on the residual sequences using the Adaptive Bandwidth Kernel Density Estimation (AKDE) algorithm. Specifically, let's assume that the first... The residual sequence corresponding to a fuzzy operating condition cluster contains Then the nth residual value, Residual probability density distribution of a fuzzy operating condition cluster It can be represented as:
[0110]
[0111] in, The value of the residual random variable is the independent variable, which represents the location of the residual when estimating the probability density distribution of the residual; For kernel functions; For the first The residuals corresponding to each historical operating data; The local adaptive bandwidth is calculated using the following formula:
[0112]
[0113] in, This is the global initial bandwidth; Geometric mean factor; for The prior probability density estimate; This is the sensitivity parameter.
[0114] IV. Real-time Coupled Membership Calculation and Prediction Interval Output
[0115] S4. For each prediction time, acquire the real-time operation data of the natural gas pipeline network and obtain its semantic matching degree and numerical membership degree relative to each fuzzy operation condition cluster. Then, fuse the semantic matching degree and numerical membership degree to obtain the real-time coupled membership weight. After the natural gas load point prediction model corresponding to each fuzzy operation condition cluster outputs the candidate point prediction value for the prediction time, extract the corresponding upper and lower quantiles of the residuals under the preset confidence level according to the residual probability density distribution established in S3. Combine the real-time coupled membership weight to dynamically and continuously weight the candidate point prediction value and the upper and lower quantiles of the residuals corresponding to each fuzzy operation condition cluster, thereby outputting the upper and lower bounds of the natural gas load prediction interval to form the natural gas load prediction interval. Then, the candidate point prediction values of each fuzzy operation condition cluster are weighted and fused according to the real-time coupled membership weight to obtain the final point prediction value for the prediction time, which is used together with the natural gas load prediction interval as the natural gas load prediction result.
[0116] It should be noted that, as Figure 3 As shown, the specific process of step S4 of the present invention is as follows:
[0117] S41. For each prediction time... The system acquires real-time operating data of the natural gas pipeline network, replaces the historical operating data in step S2 with the real-time operating data, and constructs a real-time numerical feature matrix according to step S2 to obtain the numerical membership degree of the real-time operating data to each fuzzy operating condition cluster. At the same time, it acquires the real-time scene semantic description corresponding to the prediction time, and then inputs the real-time scene semantic description and the pre-acquired real-time operating variable description into the large model to obtain the semantic matching degree of the real-time operating data relative to each fuzzy operating condition cluster.
[0118] It should be noted that in step S41 of the present invention, the real-time scene semantic description refers to the semantic expression information of the real-time operating status of the natural gas pipeline network at the predicted time. The semantic expression information includes at least one or more of the following: holiday features, time period features, weather change features, peak-valley transition features, supply and demand change features, and sudden boundary scene features, which are used to characterize the working condition attributes of the real-time operating status at the semantic level.
[0119] S42. The numerical membership degree and semantic matching degree are then fused by exponentiation and normalized to obtain the real-time coupled membership weight, which aims to comprehensively reflect the dual membership relationship of the real-time running status in the numerical space and semantic space.
[0120] It should be noted that in step S42 of the present invention, the first Real-time coupling membership weights corresponding to each fuzzy operating condition cluster The calculation formula is:
[0121]
[0122] in, Indicates the total number of fuzzy operating condition clusters; For the first Numerical membership degree of a fuzzy operating condition cluster For the first The numerical membership degree of a fuzzy operating condition cluster is used to represent the degree of numerical belonging of the real-time numerical feature matrix to the corresponding fuzzy operating condition cluster. For the first The semantic matching degree of a fuzzy operating condition cluster. For the first The semantic matching degree of a fuzzy operating condition cluster is used to represent the degree of semantic matching between the real-time scene semantic description and the corresponding fuzzy operating condition cluster. and These are the adjustment parameters corresponding to numerical membership degree and semantic matching degree, respectively, used to control the influence weight of the two in the coupling process.
[0123] Furthermore, each real-time coupling membership weight is non-negative, and the sum of all real-time coupling membership weights is 1, which satisfies:
[0124]
[0125]
[0126] in, For the first Real-time coupled membership weights corresponding to a fuzzy operating condition cluster.
[0127] Through the above coupling method, the real-time prediction results depend not only on the position of the real-time running data in the numerical feature space, but also on the running semantic interpretation corresponding to the real-time scene semantic description. This enables more accurate handling of natural gas load prediction problems in scenarios with ambiguous boundaries, such as holiday switching, peak-valley transition, and extreme weather.
[0128] S43. Input the real-time numerical feature matrix into the pre-trained natural gas load point prediction model corresponding to each fuzzy operating condition cluster to obtain the candidate point prediction value under each fuzzy operating condition cluster at the prediction time.
[0129] S44. For each fuzzy operating condition cluster, first obtain the corresponding cumulative distribution function by integration, and then extract the upper quantile and lower quantile of the residual from the cumulative distribution function under a preset confidence level.
[0130] It should be noted that in step S44 of the present invention, the first... Taking a fuzzy operating condition cluster as an example, its corresponding cumulative distribution function Calculate using the following formula:
[0131]
[0132] in, For integration variables. Then, at the preset confidence level. Below, the cumulative distribution function described above is used to extract the first... Upper quantile of the residuals corresponding to a fuzzy operating condition cluster and residual lower quantile This is used to characterize the probability fluctuation range of the prediction error under the corresponding fuzzy operating condition cluster, specifically:
[0133]
[0134]
[0135] in, For the significance level, correspondingly, The corresponding confidence level; Indicates the first The inverse function of the cumulative distribution function corresponding to each fuzzy operating condition cluster.
[0136] S45. Add the candidate point prediction values of each fuzzy operating condition cluster to the corresponding upper quantile of the residual, and use this as the upper bound prediction result. Then, use real-time coupled membership weights to weight the sum, and use this as the weighted upper bound prediction result for that fuzzy operating condition cluster. Finally, sum the weighted upper bound prediction results of all fuzzy operating condition clusters to obtain the upper bound of the output natural gas load prediction range. The predicted candidate points of each fuzzy operating condition cluster are added to their corresponding lower quantiles of the residuals to obtain the lower bound prediction result. This lower bound prediction result is then weighted using real-time coupled membership weights. Finally, the weighted lower bound prediction results of all fuzzy operating condition clusters are summed to obtain the lower bound of the output natural gas load prediction range. The natural gas load forecasting range is formed by the upper and lower boundaries of the natural gas load forecasting range. .
[0137] It should be noted that, in step S45 of this invention, the upper and lower boundaries of the natural gas load prediction interval are respectively represented as:
[0138]
[0139]
[0140] in, , , and The first Real-time coupled membership weights, candidate point prediction values, upper quantiles, and lower quantiles of residuals corresponding to each fuzzy operating condition cluster; That is, the first The upper bound prediction results of a fuzzy operating condition cluster. That is, the first The weighted upper bound prediction results of a fuzzy operating condition cluster That is, the first Lower bound prediction results for a fuzzy operating condition cluster That is, the first The weighted lower bound prediction results of a fuzzy operating condition cluster.
[0141] S46. The candidate point prediction values of each fuzzy operating condition cluster are weighted and fused according to the real-time coupling membership weight to obtain the final point prediction value at the prediction time, and together with the natural gas load prediction interval, it is used as the natural gas load prediction result.
[0142] It should be noted that in step S46 of this invention, the final point prediction value The calculation formula is:
[0143]
[0144] The aforementioned final point prediction value is used to characterize the central value of natural gas load prediction, and the natural gas load prediction interval is used to characterize the possible fluctuation range of natural gas load under the preset confidence level. Therefore, this invention uses the final point prediction value and its corresponding natural gas load prediction interval together as the natural gas load prediction result at the prediction time, so as to characterize the central prediction level of natural gas load and its uncertainty range, and provide a basis for natural gas pipeline network scheduling and control, operation risk early warning, supply and demand balance analysis and auxiliary decision-making.
[0145] To better demonstrate the specific implementation and technical effects of the present invention, the method for estimating natural gas load intervals by large model coupled operating condition clustering shown in steps S1 to S4 of the above preferred implementation is applied to a specific example.
[0146] Example
[0147] In this embodiment, the semantic constraint information of the operating condition is generated according to the aforementioned step S1, and then the real-time coupled membership weight calculation and prediction interval output are performed according to step S4. This realizes the natural gas load interval estimation method using "semantic-numerical coupled operating condition clustering - intra-cluster conditional residual quantile modeling - dynamic continuous weighted fusion". The specific implementation process of this method is as described above and will not be repeated here. The following mainly shows its specific implementation details and technical effects.
[0148] This embodiment selects a natural gas gate station in a certain region as the data source for case verification, and uses the standard condition total of the eight branches of the gate station as the natural gas load data. At the same time, meteorological data (including temperature and rainfall) and time labels (such as hour, day / night segment, weekday / weekend, holiday / non-holiday, pre-holiday / post-holiday offset, etc.) are collected synchronously with the observation period.
[0149] The data was divided into training, validation, and test sets in chronological order, with a 70% training, 15% validation, and 15% test set ratio. The training set was used to train the coupled load case clustering model and to model predictions and residual quantiles within each load case cluster; the validation set was used for model selection and hyperparameter calibration under different cluster numbers / fusion and confidence level settings; and the test set was used to evaluate the final interval estimation performance.
[0150] Then, the corresponding scene background text, time labels, and runtime variable descriptions from the training set are input into the large model, causing it to output clustering-oriented semantic constraint information for operating conditions. Simultaneously, this embodiment constructs a numerical feature matrix containing historical load time lag features, meteorological features, holiday offset features, and periodic features. To achieve matching with subsequent clustering calculations, this embodiment further standardizes / normalizes the features and adds feature weight normalization before clustering. The aforementioned semantic constraint information for operating conditions is then mapped into a clustering constraint matrix for coupling operating condition clustering, and constraint weights and applicable ranges are set, thereby simultaneously constraining "numerical similarity" and "semantic consistency" in the objective function.
[0151] After obtaining the above clustering constraint matrix and numerical feature matrix, this invention further employs a fuzzy C-means clustering algorithm based on Mahalanobis distance for soft clustering. (Number of clusters for fuzzy operating conditions) The selection is based on a combination of two metrics: validation set interval coverage performance and interval compactness.
[0152] In this embodiment, the cluster number candidate set is set as For each candidate number of clusters, the corresponding clustering results are used to train a coupled operating condition clustering model, and the representativeness of the test set interval indicators is evaluated to ultimately determine the number of clusters. .
[0153] For each fuzzy operating condition cluster, historical operating data within the cluster is first extracted to train the natural gas load point prediction model and obtain the predicted value. After the model is trained, the final prediction result can be generated according to step S4. It is worth noting that in this embodiment, to ensure the smoothness and robustness of the interval estimation, the bandwidth of AKDE modeling is calibrated using a validation set, maintaining a consistent bandwidth strategy during the testing phase, and extreme anomalies in the residual distribution are removed according to a threshold rule (the anomaly threshold can be implemented using a scaling threshold based on quantile statistics) to reduce the interference of outlier residuals on quantile estimation.
[0154] This embodiment uses the Interval Coverage Probability Index (PICP) and the Interval Average Width Index (PINAW) for evaluation. The variance of the interval width change between adjacent time points is used to measure continuous transition capability. The results are shown in Table 1. Here, PINC represents the nominal confidence level; PICP is the proportion of the actual load in the test set sample that falls into the corresponding prediction interval; and PINAW is the normalized average of the upper and lower bounds of the prediction interval width.
[0155] Table 1 Performance Indicators for Prediction Intervals
[0156] As shown in Table 1, when the nominal confidence level is 60%~90%, the PICP and PINC of the prediction interval maintain a high degree of consistency, indicating that the method of the present invention has strong coverage reliability; at the same time, PINAW increases with the increase of confidence level, reflecting a reasonable trade-off between interval compactness and coverage.
[0157] Table 2 Continuity Indicators of Operating Condition Boundaries
[0158] In Table 2, the variance of the interval width change between adjacent time points is used to measure the fluctuation of the interval width at the transition boundary of the operating condition. As shown in Table 2, the variance of the predicted interval width change at the transition boundary of the operating condition in this embodiment is small, indicating that dynamic continuous weighted fusion can achieve smooth interval transition and avoid abrupt interval changes.
[0159] To illustrate the contributions of semantic-numerical coupled clustering and dynamic fusion, this embodiment presents a comparative ablation experiment, the results of which are shown in Table 3. As can be seen from Table 3, compared to numerical clustering alone or a fixed fusion method, this embodiment reduces the average interval width and improves interval compactness while maintaining the coverage probability.
[0160] Table 3 Comparison of ablation test results
[0161] In summary, this embodiment successfully validated the interval prediction process using real-world operational data from a natural gas gate station in a specific region. The results demonstrate that the proposed method achieves high interval coverage reliability under complex and fuzzy operating conditions while maintaining interval compactness. Furthermore, dynamic continuous weighted fusion enhances the smooth transition capability at operating condition boundaries.
[0162] It is understood that the natural gas load range estimation method based on large-model coupled operating condition clustering described in S1-S4 above can essentially be implemented by a computer program. Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer program product corresponding to the natural gas load range estimation method based on large-model coupled operating condition clustering provided in the above embodiments. This product includes a computer program / instruction that, when executed by a processor, can implement the natural gas load range estimation method based on large-model coupled operating condition clustering as described in the above embodiments.
[0163] Similarly, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer electronic device corresponding to the natural gas load interval estimation method of large model coupled operating condition clustering provided in the above embodiments, such as... Figure 4 As shown, it includes a memory and a processor;
[0164] The memory is used to store computer programs;
[0165] The processor is configured to, when executing the computer program, implement a natural gas load range estimation method based on large-model coupled operating condition clustering as described in the above embodiments.
[0166] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0167] Therefore, based on the same inventive concept, another preferred embodiment of the present invention also provides a computer-readable storage medium corresponding to the natural gas load range estimation method of large model coupled operating condition clustering provided in the above embodiment. The storage medium stores a computer program, which, when executed by a processor, can realize the natural gas load range estimation method of large model coupled operating condition clustering in the above embodiment.
[0168] Specifically, in the computer-readable storage medium of the above three embodiments, the stored computer program is executed by a processor, which can perform the aforementioned steps S1 to S4.
[0169] It is understood that the aforementioned storage media may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Furthermore, the storage media may also be various media capable of storing program code, such as USB flash drives, external hard drives, magnetic disks, or optical discs.
[0170] It is understood that the processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0171] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the invention. Therefore, all technical solutions obtained through equivalent substitution or transformation fall within the protection scope of the present invention.
Claims
1. A method for estimating natural gas load intervals using large-scale model coupled operating condition clustering, characterized in that, Includes the following steps: S1. Obtain historical operating data and scenario background text of the natural gas pipeline network. The historical operating data includes natural gas load data, meteorological data, and time stamps. The scenario background text includes a pipeline topology summary, a user business composition summary, a holiday arrangement description, a pipeline operation rule summary, and a scheduling constraint description. Input the working condition semantic parsing prompt words containing scenario background text, time stamps, and operating variable descriptions into the large model to instruct the large model to identify the working condition types, working condition transition relationships, and key factors affecting working condition changes in the natural gas system, thereby outputting working condition semantic constraint information. S2. Construct a numerical feature matrix based on historical operating data, construct a clustering constraint matrix using the operating condition semantic constraint information generated by the large model, and then input the clustering constraint matrix and the numerical feature matrix into the coupled operating condition clustering model to obtain several fuzzy operating condition clusters and the membership degree of each historical operating data to each fuzzy operating condition cluster. S3. For each fuzzy operating condition cluster, the numerical feature matrix corresponding to the historical operating data of the fuzzy operating condition cluster is used as input, and the actual value of the natural gas load data at the next moment is used as the supervision label. The natural gas load point prediction model corresponding to the fuzzy operating condition cluster is trained by supervised learning. After training, the residual between the model prediction value and the actual value at each moment is calculated. The residuals at each moment form a residual sequence. The adaptive bandwidth kernel density estimation algorithm is used to model each residual sequence to obtain the residual probability density distribution corresponding to each fuzzy operating condition cluster. S4. For each prediction time, obtain the real-time operation data of the natural gas pipeline network and obtain its semantic matching degree and numerical membership degree relative to each fuzzy operation condition cluster; then, fuse the semantic matching degree and numerical membership degree to obtain the real-time coupled membership weight. After the candidate point prediction values for the prediction time are output by the natural gas load point prediction model corresponding to each fuzzy operating condition cluster, the upper and lower quantiles of the corresponding residuals are extracted according to the residual probability density distribution established in S3 under the preset confidence level. The candidate point prediction values and the upper and lower quantiles of the residuals corresponding to each fuzzy operating condition cluster are dynamically and continuously weighted by real-time coupled membership weights, so as to output the upper and lower bounds of the natural gas load prediction interval and form the natural gas load prediction interval. Then, the candidate point prediction values of each fuzzy operating condition cluster are weighted and fused according to the real-time coupling membership weight to obtain the final point prediction value at the prediction time, which is then used together with the natural gas load prediction interval as the natural gas load prediction result.
2. The natural gas load range estimation method based on large-model coupled operating condition clustering as described in claim 1, characterized in that, In step S1, the semantic constraint information of the working conditions output by the large model includes at least one or more of the following: a set of working condition type labels, working condition transition relationships, and working condition constraint rules. The set of working condition type labels includes weekday working conditions, pre-holiday peak-loading working conditions, post-holiday recovery working conditions, low-temperature high-load working conditions, and extreme rainfall disturbance working conditions. The working condition constraint rules include semantic similarity constraints that prioritize classifying semantically similar historical operating data into the same fuzzy operating condition cluster, semantic exclusivity constraints that prevent semantically exclusive historical operating data from being classified into the same fuzzy operating condition cluster, and adjacency relationship constraints between working conditions.
3. The natural gas load range estimation method based on large-model coupled operating condition clustering as described in claim 1, characterized in that, In step S2, the numerical feature matrix includes at least one or more of the following: historical load time lag features, sliding statistical features, meteorological features, periodic features, and holiday offset features.
4. The natural gas load range estimation method based on large-model coupled operating condition clustering as described in claim 3, characterized in that, The historical load time lag feature is based on the current natural gas load data at a certain time lag before that time. The sliding statistical feature is a statistical quantity calculated from natural gas load data or meteorological data based on a preset time window. The statistical quantity includes one or more of the following: load sliding mean, temperature sliding mean, load sliding variance, and cumulative rainfall sliding value. The meteorological feature includes one or more of the following: ambient temperature, rainfall, relative humidity, wind speed, and perceived temperature. The periodic feature is a time-based feature constructed based on the position of the historical operating data at a certain time in a period such as hour, day, week, or month. The holiday offset feature refers to the feature that characterizes the relative distance between the current historical operating data and the start or end time of a holiday.
5. The natural gas load range estimation method based on large-model coupled operating condition clustering as described in claim 1, characterized in that, In step S2, each element of the clustering constraint matrix represents the semantic constraint relationship between two historical operation data. The specific way to construct the clustering constraint matrix from the semantic constraint information of the working conditions is as follows: According to the analysis results of the large model, when it is identified that two historical operation data correspond to the same type of working condition semantics or the working conditions corresponding to two historical operation data can be classified into the same operating stage in terms of business, then these two historical operation data are regarded as historical operation data with similar semantics, and the element values of the corresponding positions in the clustering constraint matrix are assigned positive values. When it is identified that two historical operation data correspond to different types of working conditions or that the working conditions corresponding to two historical operation data cannot be classified into the same operating stage in terms of business, these two historical operation data are regarded as semantically contradictory historical operation data, and the element values of the corresponding positions in the clustering constraint matrix are assigned negative values. If the semantic constraint relationship between two historical data points is not identified, the element value at the corresponding position in the clustering constraint matrix will be assigned the value 0.
6. The natural gas load interval estimation method based on large-model coupled operating condition clustering as described in claim 1, characterized in that, In the coupled operating condition clustering model of step S2, the fuzzy C-means clustering algorithm is used to perform soft clustering on the historical operating data, and Mahalanobis distance is used as the distance metric between the historical operating data and the cluster centers. At the same time, a semantic constraint penalty term based on the clustering constraint matrix is introduced into the objective function of the fuzzy C-means clustering algorithm, so that the final fuzzy operating condition clusters can simultaneously satisfy numerical similarity and operating condition semantic consistency. The Mahalanobis distance in the objective function is calculated based on the numerical feature matrix of the historical operating data and the cluster centers of the fuzzy operating condition clusters. The semantic constraint penalty term is calculated based on the clustering constraint matrix and the membership distribution of the historical operating data in each fuzzy operating condition cluster.
7. The natural gas load range estimation method based on large-model coupled operating condition clustering as described in claim 1, characterized in that, The specific process of step S4 is as follows: S41. For each prediction time, acquire the real-time operation data of the natural gas pipeline network, replace the historical operation data in S2 with the real-time operation data, and construct the real-time numerical feature matrix according to step S2 and obtain the numerical membership degree of the real-time operation data to each fuzzy operation condition cluster; at the same time, acquire the real-time scene semantic description corresponding to the prediction time, and then input the real-time scene semantic description and the pre-acquired real-time operation variable description into the large model to obtain the semantic matching degree of the real-time operation data relative to each fuzzy operation condition cluster. S42. Then, the numerical membership degree and semantic matching degree are fused by exponentiation and normalized to obtain the real-time coupled membership weight. S43. Input the real-time numerical feature matrix into the pre-trained natural gas load point prediction model corresponding to each fuzzy operating condition cluster to obtain the candidate point prediction value under each fuzzy operating condition cluster at the prediction time. S44. For each fuzzy operating condition cluster, first obtain the corresponding cumulative distribution function by integration, and then extract the upper quantile and lower quantile of the residual from the cumulative distribution function under a preset confidence level. S45. Add the predicted values of candidate points for each fuzzy operating condition cluster to the corresponding upper quantile of the residual, and use this as the upper bound prediction result. Then, use real-time coupled membership weights to weight the predicted values, and use this as the weighted upper bound prediction result for that fuzzy operating condition cluster. Finally, sum the weighted upper bound prediction results of all fuzzy operating condition clusters to obtain the upper bound of the output natural gas load prediction range. Add the predicted values of candidate points for each fuzzy operating condition cluster to the corresponding lower quantile of the residual, and use this as the lower bound prediction result. Then, use real-time coupled membership weights to weight the predicted values, and use this as the weighted lower bound prediction result for that fuzzy operating condition cluster. Finally, sum the weighted lower bound prediction results of all fuzzy operating condition clusters to obtain the lower bound of the output natural gas load prediction range. The natural gas load forecast interval is formed by the upper and lower boundaries of the natural gas load forecast interval; S46. The candidate point prediction values of each fuzzy operating condition cluster are weighted and fused according to the real-time coupling membership weight to obtain the final point prediction value at the prediction time, and together with the natural gas load prediction interval, it is used as the natural gas load prediction result.
8. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it can implement the natural gas load range estimation method of large model coupled operating condition clustering as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the natural gas load range estimation method based on large model coupled operating condition clustering as described in any one of claims 1 to 7.
10. A computer electronic device, characterized in that, Including memory and processor; The memory is used to store computer programs; The processor is configured to, when executing the computer program, implement the natural gas load range estimation method based on large model coupled operating condition clustering as described in any one of claims 1 to 7.