Load side data sample marking method fusing data mode and adjustment criterion
By constructing a multi-dimensional load feature system and secondary AP clustering, combined with physical criteria and supervised learning, structured labels are generated, which solves the problems of subjectivity and low accuracy in load sample labeling, realizes the accuracy and engineering applicability of load potential identification, and provides interpretable labeling results.
Patent Information
- Application Number
- CN202511458994.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2026-01-20
AI Technical Summary
In existing technologies, load sample labeling methods rely on human experience or a single indicator, which leads to strong subjectivity and low accuracy in identifying load adjustability potential. Furthermore, the lack of physical adjustment capability verification results in poor engineering practicality of the identification results, and the relationship between features and response potential is opaque, making it difficult to interpret and reuse.
A multi-dimensional load characteristic system is constructed. By combining secondary AP clustering and three physical criteria, a structured load sample labeling system is generated, including electricity consumption scale, fluctuation characteristics, regularity characteristics and sensitivity characteristics to external factors. Combined with adjustable continuity, rate and capacity ratio, supervised learning is used to identify key features and output structured labels.
It improves the accuracy and engineering applicability of load adjustability potential identification, provides high-quality data support, provides standardized support for sample selection and extrapolation in different power grid control scenarios, and the labeling results are interpretable and reusable.
Smart Images

Figure CN121365261A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of smart grid and power system automation, and particularly relates to a load side data sample labeling method fusing data mode and regulation criterion. BACKGROUND
[0002] With high proportion of renewable energy access to the power grid, the flexible regulation capability of load side resources (such as industrial and commercial users, temperature control loads, electric vehicles, etc.) becomes an important support to maintain the balance between supply and demand of the power grid. However, whether the load sample has "adjustable potential" is an implicit feature that cannot be directly observed, and the traditional method relies on artificial experience labeling or single index screening, which has strong subjectivity, poor generalization ability and insufficient engineering applicability.
[0003] In the prior art, clustering analysis is commonly used for power user grouping, which is mostly based on load curve shape (such as bimodal type, smooth type) or electricity consumption scale for division, aiming to realize user portrait and differentiated management. However, such methods generally have a key defect: equating "large electricity consumption scale" with "high adjustable potential", ignoring the physical adjustability and behavioral flexibility of users' actual participation in demand response. Practice shows that not all industrial and commercial users are suitable as adjustable resources. For example, first-class important electricity users, large shopping malls, hospitals, etc., although their total load is large and the peak and valley are obvious, due to the high requirement of production continuity or the rigidity of business hours, it is difficult for them to actively adjust the electricity consumption behavior even under the price incentive, and their willingness to participate in peak clipping and valley filling is low. On the contrary, some manufacturing users, electric vehicle charging stations, park energy storage, etc., although the individual regulation capacity is limited, their load has strong time flexibility and interruptibility, and has more actual regulation value.
[0004] In addition, the prior art lacks quantitative verification of the physical regulation capacity of users when generating "potential labels", resulting in a disconnection between the identification results and the scheduling practicability; at the same time, the correlation between the features and the response potential is not transparent, making it difficult to explain "why is this user determined as high potential", which limits the application of labels in digital twin and scenario deduction. SUMMARY
[0005] The technical problems to be solved by the present application are to provide a load-side data sample labeling method fusing data patterns and adjustment criteria, and to generate sample labeling results that can be directly used for sample screening and combination under different operation scenarios (normal, alert, fault, emergency, recovery), to provide high-quality data support for power grid control typical scenario deduction, to effectively improve the accuracy and engineering applicability of high-potential adjustable load identification, and to solve the technical problems of strong subjectivity and low accuracy of load adjustable potential identification caused by relying on artificial experience or a single indicator, poor engineering practicability of identification results caused by lack of physical adjustment capability verification, and weak interpretability and difficulty in reuse of labeling results caused by opaque correlation between features and potential.
[0006] The present application adopts the following technical solutions: A load-side data sample labeling method fusing data patterns and adjustment criteria, comprising the following steps: Constructing a load feature set for describing user electricity consumption behavior and response characteristics to external environment, the load features including electricity consumption scale features, electricity consumption fluctuation features, electricity consumption regularity features, and external factor sensitivity features; First, the users are grouped by secondary AP clustering to identify candidate groups with adjustable behavior characteristics, and then the individual users in the candidate groups are verified for adjustment capability in combination with three physical criteria of adjustable duration, adjustable rate, and adjustable capacity ratio to construct a binary classification response potential label; Modeling the user adjustment potential identification problem as a binary classification supervised learning task, training the binary classification supervised learning task based on the constructed load feature set and the constructed binary classification response potential label, identifying key features through feature importance analysis, and outputting a structured load sample labeling system.
[0007] Preferably, the specific process of constructing a multi-dimensional load feature system comprises: The electricity consumption scale features include average load and extreme value load, wherein the average load is obtained by calculating the average value of load at each time during the investigation period or the ratio of electricity consumption to time, and the extreme value load is obtained by determining the maximum or minimum value of load during the investigation period, or using the maximum 3-day average load to describe the peak value characteristics of the power load; The electricity consumption fluctuation features include peak-valley difference rate and load rate, the peak-valley difference rate is obtained by calculating the ratio of daily load peak-valley difference to daily maximum load, and is used to reflect the degree of daily load peak-valley change, and the load rate is obtained by calculating the ratio of average load to maximum load during the investigation period, and is used to reflect the degree of imbalance of load distribution; The electricity usage regularity features include a flexibility coefficient and a regularity coefficient, the flexibility coefficient is obtained by calculating average fluctuation level of daytime electricity usage load mode within a time scale, and the regularity coefficient is obtained by calculating consistency and stability of electricity usage behavior within a time scale by using a Spearman correlation coefficient; The external factor sensitivity features include a temperature sensitivity coefficient, a humidity sensitivity coefficient, an electricity price sensitivity coefficient, and a holiday sensitivity coefficient, the temperature sensitivity coefficient, the humidity sensitivity coefficient, and the electricity price sensitivity coefficient are obtained by using a Spearman correlation coefficient to measure consistency degree of change trend of a daily load curve of a user with a temperature curve, a humidity curve, and an electricity price curve, and the holiday sensitivity coefficient is obtained by calculating change of total electricity usage load on holidays and weekdays.
[0008] Preferably, the specific process of the secondary AP clustering includes: The first stage: average values of multi-day load data of each user are calculated to obtain typical daily load curves, score normalization processing is performed on the typical daily load curves to eliminate absolute difference of total electricity usage among users, AP clustering algorithm is combined with cosine similarity measurement to cluster the normalized typical daily load curves, and a plurality of typical behavior mode clusters are output; The second stage: based on the clustering result of the first stage, users with abnormal or special electricity usage modes are removed, AP clustering algorithm is combined with Euclidean distance measurement to cluster the filtered user load data again, user electricity usage scale is distinguished on the basis of similar behavior modes, and a user candidate set with adjustable potential is obtained.
[0009] Preferably, the specific process of verifying the adjustment ability of the individual user in combination with the three physical criteria includes: S2021, determining a load change direction, identifying a transferable rate calculation interval, an adjustable interval, and a post-transfer recovery rate calculation interval, calculating an adjustable duration proportion, and determining that the user meets the adjustable duration requirement when the adjustable duration proportion is not less than 30% of the total sampling duration; S2022, calculating a maximum ramp rate of historical load of the user, taking percentage of the maximum rising rate to equivalent rated capacity as a representation of the adjustable rate, and determining that the user meets the adjustable rate requirement when the adjustable rate is not less than 5% / min, wherein the equivalent rated capacity is the 95% quantile load of the user; S2023, identifying a high-position operation interval of a typical peak period of the load, finding time periods in which the load is lower than a preset threshold in the interval and calculating average load of these time periods, determining adjustable capacity of the user in the peak period, calculating a ratio of the adjustable capacity to the equivalent rated capacity to obtain an adjustable capacity ratio, and determining that the user meets the adjustable capacity ratio requirement when the adjustable capacity ratio is not less than 10%.
[0010] Preferably, in step S2021, the specific process of identifying the transferable rate calculation interval, the adjustable interval and the post-transfer recovery rate calculation interval is as follows: A variable is set and a threshold is set, the variable records the positive and negative of the load change direction, the load change direction variable is read, if the variable is non-negative, it is assigned as 1, otherwise it is assigned as -1; The interval with the variable value being continuously 1 and the number of monitoring points being greater than or equal to the threshold is defined as the transferable rate calculation interval, the interval with the variable value being continuously -1 and the number of monitoring points being greater than or equal to the threshold is defined as the post-transfer recovery rate calculation interval, and the interval between the transferable rate calculation interval and the post-transfer recovery rate calculation interval is defined as the adjustable interval.
[0011] Preferably, in step S2023, the specific process of determining the adjustable capacity of the user in the peak period is as follows: The low load persistence method is used to find the period in which the load is continuously lower than the preset threshold in the high position operation interval of the typical peak period of the load, and the average load of these periods is calculated; The difference between the average load in the high position operation interval and the average load calculated above is determined as the adjustable capacity of the user in the peak period.
[0012] Preferably, the specific process of constructing the binary classification response potential label is as follows: The load pattern of each cluster obtained by analyzing the secondary AP clustering is analyzed to determine a number of high-potential candidate categories with adjustable behavior characteristics, and the users in the high-potential candidate categories constitute a candidate set; If the user belongs to the high-potential candidate category and simultaneously satisfies the adjustable persistence, adjustable rate and adjustable capacity ratio three physical criteria, the user is marked as a high-potential adjustable load, and the label is marked as y=1; if the user does not satisfy any of the above conditions, the user is marked as a low-potential adjustable load, and the label is marked as y=0.
[0013] Preferably, the specific process of key feature identification and sample label optimization includes: S301, using the multi-dimensional load feature system constructed in step S1 as the input feature vector, using the binary classification response potential label constructed in step S2 as the model output, and using the random forest classifier to construct a binary classification supervised learning model; S302, the Gini importance is used to evaluate the contribution of each input feature to the model classification decision, the feature importance ranking result is output, and the key features that have a significant impact on the user adjustment potential are identified; S303, constructing a structured load sample labeling system including a basic label layer, a potential level layer, and a semantic rule layer based on the feature importance analysis result and the secondary AP clustering result, wherein the basic label layer records the behavior pattern category and the physical capability compliance of the user, the potential level layer divides the response potential level of the user according to the feature importance and the scene weight, and the semantic rule layer generates an interpretable composite label by combining the high importance features and the clustering result.
[0014] Preferably, in step S303, the division process of the response potential level is as follows: According to the feature importance analysis result, the weight of each key feature is determined, the scene weight is set in combination with different power grid operation scenes, the response potential comprehensive score of the user is calculated based on the value of the user on each key feature, the key feature weight and the scene weight, and the response potential of the user is divided into three levels of high, medium and low according to the distribution of the response potential comprehensive score, so as to form the labeling content of the potential level layer. Preferably, in step S303, the generation process of the interpretable composite label of the semantic rule layer is as follows: The high importance features with high importance are screened, and the effective value range of each high importance feature is determined; the association rules between the value of the high importance feature and the behavior pattern category are established in combination with the user behavior pattern category obtained by the secondary AP clustering; The interpretable composite label is generated based on the association rules, and the interpretable composite label can reflect the comprehensive influence of the high importance features and the behavior pattern category of the user on the response potential.
[0015] In a second aspect, an embodiment of the present application provides a load side data sample labeling system integrating data patterns and adjustment criteria, characterized in that it comprises: A feature module is configured to construct a load feature set for describing the electricity consumption behavior of the user and the response characteristics to the external environment, wherein the load features include electricity consumption scale features, electricity consumption fluctuation features, electricity consumption law features, and external factor sensitivity features. A sample module is configured to first group the users by the secondary AP clustering, identify a candidate group with adjustable behavior characteristics, and then verify the adjustment capability of the individual users in the candidate group in combination with three physical criteria of adjustable persistence, adjustable rate, and adjustable capacity ratio, to construct a binary classification response potential label. A labeling module is configured to model the user adjustment potential identification problem as a binary classification supervised learning task, train the binary classification supervised learning task based on the load feature set constructed in step S1 and the binary classification response potential label constructed in step S2, identify key features through feature importance analysis, and output a structured load sample labeling system.
[0016] Preferably, the feature module comprises: The electricity consumption scale feature extraction unit is configured to calculate the average load at each time in the observation period or the ratio of electricity consumption to time to obtain an average load, determine the maximum or minimum value of the load in the observation period, or adopt a maximum 3-day average load to describe the peak value characteristics of the power load to obtain a maximum or minimum load; The electricity consumption fluctuation feature extraction unit is configured to calculate the ratio of the daily load peak-valley difference to the daily maximum load to obtain a peak-valley difference rate, and calculate the ratio of the average load to the maximum load in the observation period to obtain a load rate. The electricity consumption regularity feature extraction unit is configured to calculate the average fluctuation level of the daily electricity consumption load mode in a certain time scale to obtain a flexibility coefficient, and calculate the consistency and stability of the electricity consumption behavior in a certain time scale to obtain a regularity coefficient. The external sensitivity feature extraction unit is configured to use the Spearman correlation coefficient to measure the consistency degree of the daily load curve of the user and the change trend of the temperature, humidity, and electricity price curve to obtain a temperature sensitivity coefficient, a humidity sensitivity coefficient, and an electricity price sensitivity coefficient, and calculate the change of the total electricity consumption load on holidays and weekdays to obtain a holiday sensitivity coefficient.
[0017] Preferably, the sample module comprises: The secondary clustering unit is configured to average the multi-day load data of each user to obtain a typical daily load curve and perform score normalization processing, perform electricity consumption mode clustering by using an AP clustering algorithm combined with cosine similarity measurement, perform comprehensive characteristic clustering by using an AP clustering algorithm combined with Euclidean distance measurement after excluding electricity consumption mode abnormalities or special users, and obtain a user candidate set with adjustable potential. The physical criterion verification unit is configured to determine the load change direction and each interval, calculate the adjustable duration proportion to verify the adjustable duration, calculate the percentage of the maximum rising rate to the equivalent rated capacity to verify the adjustable rate, determine the adjustable capacity and calculate the adjustable capacity ratio to verify the adjustable capacity ratio. The label generation unit is configured to analyze the load pattern of each clustering cluster to determine a high-potential candidate category, and generate a binary classification response potential label according to whether the user belongs to the high-potential candidate category and whether the three physical criteria are met.
[0018] Preferably, the marking module comprises: The model construction unit is configured to use a multi-dimensional load feature system as an input feature vector, use a binary classification response potential label as a model output, and construct a binary classification supervised learning model by using a random forest classifier. The feature importance analysis unit is configured to evaluate the contribution of each input feature to the model classification decision by using Gini importance, output a feature importance ranking result, and identify key features. The marking system construction unit is configured to construct a structured load sample marking system including a basic label layer, a potential level layer and a semantic rule layer based on the feature importance analysis result and the secondary AP clustering result.
[0019] Preferably, the marking system construction unit comprises: The basic label generation subunit is configured to record the behavior pattern categories obtained through the secondary clustering and the physical ability compliance obtained through the physical criterion verification. The potential level division subunit is configured to determine the key feature weight according to the feature importance, calculate a user response potential comprehensive score in combination with a scene weight, and divide the response potential level according to the score. The semantic label generation subunit is configured to filter high importance features and determine their effective value range, establish an association rule between the high importance feature value and the behavior pattern category, and generate an interpretable composite label.
[0020] In a third aspect, a computer device comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the load side data sample marking method of fusing data patterns and adjustment criteria when executing the computer program.
[0021] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium comprising a computer program, and the computer program implements the steps of the load side data sample marking method of fusing data patterns and adjustment criteria when executed by a processor.
[0022] In a fifth aspect, a chip comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the load side data sample marking method of fusing data patterns and adjustment criteria when executing the computer program.
[0023] In a sixth aspect, an embodiment of the present application provides an electronic device comprising a computer program, and the computer program implements the steps of the load side data sample marking method of fusing data patterns and adjustment criteria when executed by the electronic device.
[0024] Compared with the prior art, the present application has at least the following beneficial effects: A load-side data sample labeling method fusing data mode and regulation criterion, first constructs a load feature set from four dimensions such as power consumption scale, then constructs a binary classification label through secondary AP clustering and three physical criteria, and finally identifies key features by supervised learning and outputs a structured labeling system; Breakthrough the limitations of traditional single index or artificial experience, form a complete link of "multi-dimensional description - dual verification of behavior and physics - interpretable optimization", not only ensure the comprehensiveness of load feature description, but also exclude misjudgment through clustering and physical criterion combination, and rely on supervised learning to realize the transparency of feature and potential correlation, Provide standardized support for sample screening in different scenarios of power grid, significantly improve the identification accuracy and engineering applicability of high potential load.
[0025] Further, a comprehensive and deep load feature portrait is constructed. This system not only includes traditional power consumption scale and fluctuation characteristics, but also innovatively introduces flexibility coefficient and regularity coefficient to measure the internal law of user behavior, as well as sensitivity coefficient to external incentives and special periods. This multi-dimensional and multi-perspective feature construction can reveal the electricity elasticity of users from different aspects, providing rich and reliable information input for subsequent clustering analysis and potential identification, which is a solid foundation to ensure the accuracy and effectiveness of the whole method.
[0026] Further, an intelligent grouping strategy is adopted, which is progressive from coarse to fine. In the first stage, cosine similarity is used to focus on the pure load curve shape, ignoring the absolute value of power consumption, so as to accurately identify early peak type, double peak type and other behavior patterns, solving the problem that traditional clustering methods cover up the similarity of shape due to scale difference. In the second stage, on the basis of the first stage, Euclidean distance is used to further distinguish the scale, realizing the fine distinction of users with the same mode but different scale. This two-stage strategy does not need to preset the number of clusters, has high automation degree, and can effectively exclude the interference of abnormal users. The final candidate set has the characteristics of "adjustable behavior" and "observable scale", providing a high-quality target group for subsequent physical verification.
[0027] Further, the abstract "adjustable potential" is quantified into three hard indicators with clear engineering significance. The adjustable persistence ensures that the user has certain adjustment endurance, rather than instantaneous response; The adjustable rate reflects the user's agility in responding to grid instructions; The adjustable capacity ratio defines the user's adjustment scale value from the total amount. These three criteria together constitute a multi-dimensional capability evaluation system. Only users who meet all the criteria are identified as high potential, which completely changes the previous extensive mode of relying on single index or experience for judgment, so that the identification result can directly support the actual dispatching decision of the power grid, and the engineering practicability is extremely strong.
[0028] Further, an objective and calculable process for automatically identifying the adjustment window based on the direction of load change is provided. By setting variables to track the positive and negative of load change, and defining the transferable rate calculation interval, the adjustable interval and the recovery interval based on continuous monitoring point number, the subjective interval judgment is converted into a data-based and repeatable algorithm step. This not only improves the automation level and reliability of the method, avoids the subjective randomness of manual division, but also provides technical support for accurately calculating the key indicator of adjustable duration ratio.
[0029] Further, the innovative concept of low load persistence method is proposed to estimate the actual reducible space of users in peak period. This method is not simply the maximum load minus the minimum load, but focuses on the "typical peak high position running interval" within which the user can continuously maintain low load operation. By identifying the period within the interval when the load is below the threshold and calculating the average load, the actual feasible capacity of users participating in peak shaving and regulation in the key period can be more truly reflected, making the capacity evaluation result more conservative, reliable and close to the actual dispatching scenario.
[0030] Further, a dual-verification collaborative decision-making mechanism is established. The mechanism stipulates that users must meet both the conditions of belonging to the high potential candidate category and meeting the three physical criteria to be labeled as high potential. This logic is the core of the high accuracy of the invention, which effectively eliminates two types of misjudgment: one is the user whose electricity consumption behavior appears to be adjustable but is physically rigid, and the other is the user who has physical ability but has no regular, predictable and schedulable electricity consumption behavior. This ensures the high credibility of the final label.
[0031] Further, a leap from black box identification to white box explanation and optimization is achieved. By using the features and labels generated by the previous steps as training data and using a random forest model for learning, not only a high-precision classifier can be obtained, but more importantly, the contribution of each feature to the high potential decision can be quantified through methods such as Gini importance. This reveals the key driving factors that affect load regulation potential, making the decision-making process of the model interpretable and traceable. Based on this, the structured labeling system greatly enriches the information dimension and application flexibility of the label.
[0032] Further, a dynamic and configurable fine-grained classification scheme is provided. Instead of simply classifying users as yes or no, the method introduces feature weights and scenario weights, calculates a comprehensive score, and then classifies them into high, medium and low levels. This makes the labeling results adaptable to the differentiated needs of different grid operation scenarios, such as in emergency situations, where the adjustment rate may be more important, and its weight is correspondingly increased. This dynamic classification mechanism makes the sample library more useful and targeted, making it possible for precise scenario deduction and resource allocation.
[0033] Further, the intelligent analysis result of the application is converted into business knowledge that can be directly understood and used by human beings. By associating the effective value range of the high importance feature with the specific behavior mode category, semantic labels such as the following are generated: if belonging to bimodal type and temperature sensitivity > 0.6→ high temperature and high potential adjustable. These labels greatly improve the explainability and usability of the labeled result, so that the operation personnel can quickly grasp the core characteristics and potential reasons of the user without understanding complex algorithms, support fast sample retrieval and combination based on natural semantics, and effectively support the construction and application of intelligent scenarios.
[0034] It can be understood that the beneficial effects of the above-mentioned second aspect to the sixth aspect can be referred to the related description in the first aspect, which will not be repeated here.
[0035] In summary, the application accurately identifies high potential adjustable load by multi-dimensional feature construction, secondary AP clustering combined with physical criteria, supervised learning optimization, solves the problems of subjectivity, no physical verification and opaque correlation of traditional methods, and the label engineering is practical, the labeling system is structured and reusable, and is suitable for power grid multi-scenario dispatching.
[0036] The technical solutions of the application will be further described in detail below with reference to the drawings and embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0037] Figure 1 is a whole flowchart of the method of the application; Figure 2 is a secondary AP clustering flowchart of the application; Figure 3 is an adjustable ability criterion interval definition diagram of the application; Figure 4 is a feature importance ranking column chart of the application; Figure 5 is a schematic diagram of a computer device provided by an embodiment of the application; Figure 6 is a block diagram of a chip according to an embodiment of the application.
[0038] Among them, 60. Computer device; 61. Processor; 62. Memory; 63. Computer program; 600. Electronic device; 610. Processing unit; 620. Storage unit; 6201. Random access storage unit; 6202. Cache storage unit; 6203. Read-only storage unit; 6204. Program / utility; 6205. Program module; 630. Bus; 640. Display unit; 650. Input / output interface; 660. Network adapter; 700. External device. DETAILED DESCRIPTION
[0039] The application provides a load-side data sample labeling method fusing data patterns and regulation criteria.
[0040] Please refer to Figure 1 The application provides a load-side data sample labeling method fusing data patterns and regulation criteria, and the method comprises the following steps. S1, multi-dimensional load feature system construction In order to comprehensively depict the user electricity consumption behavior and the response characteristics to the external environment, a load feature set covering four dimensions of electricity consumption scale, fluctuation characteristics, behavior regularity and external sensitivity is constructed.
[0041] S101, electricity consumption scale feature Average load From the statistical point of view, the average load refers to the average value of the load at each time in the observation period; from the physical point of view, the average load refers to the ratio of the electricity consumption to the time in the observation period, and the average load defined in this way is also called average power load.
[0042] Average load of the i-th day is expressed as: (1) wherein, is the power value of the i-th day and the l-th sampling time point of the user, is the number of time points.
[0043] For the average load of a user for consecutive N days is expressed by the following formula: (2) Extreme load The maximum or minimum load recorded in the investigation period (such as day, month, year) is referred to as the maximum or minimum load. The time span of the load measured by the electric energy meter is instantaneous, 15 minutes, 30 minutes, and 60 minutes. The maximum (minimum) load of the day refers to the maximum (minimum) value measured by the hourly electric energy meter. The calculation of the maximum (minimum) load of the day, month, and year is shown in the following formula (only the maximum load is taken as an example, and the minimum load is calculated in the same way): The maximum load of the day = max{the load record value at the i th hour in a day} i = 1, 2,..., 24 (3) The maximum load of the month = max{the maximum load of the i th day in a month} i = 1, 2,..., 31 (4) The maximum load of the year = max{the maximum load of the i th month in a year} i = 1, 2,..., 12 (5) According to related research, the "maximum load" should represent a highest load level state, rather than an upper limit value with accidental factors (such as caused by individual abnormal weather, sudden events), and the use of a single load instantaneous or integral point maximum value will have certain limitations. In order to exclude the maximum load value caused by accidental factors, it is suggested to use the maximum 3-day average load to describe the peak value characteristics of the electric power load, and the index can be obtained by taking the average of the maximum 3-day load values in a month.
[0044] For a user, the average maximum / minimum load of N days is , represented by the following formula: (6) (7) S102, power fluctuation characteristics The fluctuation index used to evaluate the user's power demand and demand response potential in the present application includes two indexes of peak-valley difference rate and load rate, wherein the peak-valley difference rate is a positive index, and the load rate is a negative index. The greater the positive index value and the smaller the negative index value, the higher the user's demand response potential, and vice versa.
[0045] Peak-valley difference rate The peak-valley difference rate reflects the degree of change of daily load peak and valley. The peak-valley difference is the difference between the maximum load and the minimum load, that is, the absolute value of the system power load change range. The definition of the i th day load peak-valley difference is as follows: (8) wherein, and These represent the maximum and minimum loads on day i, respectively. Based on the daily peak-to-valley difference, the maximum peak-to-valley difference and the average peak-to-valley difference can be obtained. The maximum peak-to-valley difference refers to the largest daily peak-to-valley difference during the observation period, while the average peak-to-valley difference refers to the average daily peak-to-valley difference during the observation period. Specific indicators include monthly maximum peak-to-valley difference, annual maximum peak-to-valley difference, monthly average daily peak-to-valley difference, and annual average daily peak-to-valley difference.
[0046] Peak-valley difference only reflects the magnitude of load change, but not the degree of change. Peak-valley difference rate, on the other hand, reflects the relative magnitude of peak-valley difference and maximum load.
[0047] Among them, the load peak-valley difference rate on day i The definition of is: (9) Indicators such as the monthly maximum peak-to-valley difference rate, the monthly average daily peak-to-valley difference rate, the annual maximum peak-to-valley difference rate, and the annual average daily peak-to-valley difference rate can be calculated based on the daily peak-to-valley difference rate.
[0048] For a user's average load peak-valley difference rate over N days Expressed as follows: (10) Load factor Reflecting the comparison between the average and peak load, the load rate is the ratio of the average load to the maximum load during the observation period, and can be used to describe the degree of unevenness in load distribution.
[0049] Daily load factor on day i The definition of is: (11) For a user's average load rate over N days Expressed as follows: (12) S103. Characteristics of Electricity Consumption Patterns Flexibility coefficient Reflecting the average fluctuation level of daytime electricity load patterns over a certain time scale, the user's flexibility coefficient over a total of N days. The calculation formula is as follows: (13) in, For user number Heavenly Daily load curve power values at each sampling time; For users The average power of the daily load curve at sampling time t within a day; This represents the total number of sampling points for the daily load curve.
[0050] Regularity coefficient Reflecting the consistency and stability of electricity consumption behavior over a certain time scale, the Spearman correlation coefficient is used in statistics to evaluate the consistency of the directions of change of two vectors. Its calculation formula is shown below: (14) in, For vectors The Spearman correlation coefficient between the two vectors ranges from [-1, 1]. The larger the absolute value, the stronger the consistency in the direction of change of the two vectors, and vice versa. For vector dimensions; Representing vectors The difference between the serial numbers of the elements in descending order.
[0051] This invention uses Spearman correlation coefficient to calculate the user's total Daily regularity coefficient The specific calculation formula is as follows: (15) in, These represent the daily load curve vectors for the user on day i and day j, respectively.
[0052] S104, Sensitivity Characteristics to External Factors Temperature / humidity / electricity price sensitivity This invention uses the Spearman correlation coefficient to measure the consistency between the user's daily load curve and the trends of temperature, humidity, and electricity price curves, and takes its absolute value to characterize the user's electricity sensitivity.
[0053] The formulas for calculating the user's temperature sensitivity coefficient M1, humidity sensitivity coefficient M2, and electricity price sensitivity coefficient M3 over a total of N days are as follows: (16) in, Indicates the first Daily user load data vector; They represent the first Vector data of temperature, humidity, and electricity price for the day.
[0054] Holiday Sensitivity This invention uses the changes in total electricity load between holidays and weekdays to reflect users' sensitivity to holidays. The user's holiday sensitivity over a total of N days is calculated. The calculation formula is as follows: (17) in, This represents the average daily electricity load for users during holidays. This represents the user's average daily electricity load during weekdays.
[0055] S2. Construction of sample labels for fusion data patterns and adjustment criteria. This invention proposes a binary classification response potential label that combines data-driven group behavior pattern recognition with individual physical adjustability criterion verification. ) Construction mechanism.
[0056] First, secondary AP clustering is used to refine the grouping of all industrial and commercial users, resulting in several typical load pattern categories (such as "morning peak type," "evening peak type," "dual-peak production type," and "stable commercial type"). The purpose of this step is not to directly determine the regulation potential, but to identify candidate groups with "transferable" and "interruptible" characteristics in their electricity consumption behavior. For example, "dual-peak production type" users may have equipment downtime during midday, possessing load transfer potential; "charging station type" users have load concentrated at night, possessing valley filling capabilities.
[0057] Then, based on the clustering results, the capabilities of individual users in the candidate groups are verified by combining three physical criteria: adjustable persistence, rate, and capacity ratio. Only those users who belong to both the "behavioralally adjustable" and "physically adjustable" groups are marked as high-potential adjustable loads.
[0058] This mechanism effectively avoids misjudging users with "high but rigid electricity consumption" (such as key enterprises and large shopping malls) as high-potential resources, thus improving the accuracy of sample labeling and the practicality of scheduling.
[0059] S201, Data Pattern Recognition To overcome the shortcomings of traditional clustering methods, such as requiring a pre-defined number of clusters and ignoring both pattern and scale characteristics, a two-stage AP clustering strategy is proposed: S2011, Phase 1: Electricity Consumption Pattern Clustering (Morphological Clustering) Objective: To identify user groups with similar electricity consumption patterns, ignoring differences in total electricity consumption and focusing on the behavioral characteristic of "adjustability".
[0060] Implementation process: Construction of typical daily load curves: The average load data of each user over multiple days is calculated to obtain their typical daily load curves, thus achieving dimensionality reduction in the time dimension; Score normalization: (18) Eliminate the absolute difference in total electricity consumption among users while preserving the shape characteristics of the load curve; AP clustering + cosine similarity measure: Affinity Propagation (AP) clustering algorithm is adopted to automatically identify the inherent structure of data without presetting the number of clusters. The distance metric adopts Cosine Similarity to measure the directional consistency of two curves, and is not sensitive to the amplitude. Output: several typical behavior pattern clusters, such as "morning peak type", "evening peak type", "double-peak production type", "stable business type", etc.
[0061] S2012, second stage: comprehensive characteristic clustering (refinement classification) Objective: To further distinguish the user electricity scale while maintaining the consistency of behavior patterns, and to provide capacity basis for subsequent resource regulation and evaluation.
[0062] Implementation process: Input data: based on the results of the first stage, eliminate a small number of electricity pattern abnormal or special (such as random fluctuations, holiday dominated, peak electricity) users to improve the stability of subsequent clustering. Use the screened cluster curve, and no further processing is done on the cluster curve to retain the true electricity level; AP clustering + Euclidean distance metric: AP clustering algorithm is adopted again, and the distance metric is changed to Euclidean distance (Euclidean Distance) to reflect the comprehensive difference of load curve in amplitude and time; Output: on the basis of "similar behavior patterns", further refinement classification is carried out to obtain user clusters with adjustable potential, which are used as candidate sets.
[0063] S202, regulation criterion design Referring to the evaluation range, evaluation index and evaluation method of the access ability, dispatching operation ability and automatic power control ability of the adjustable load participating in the grid regulation in the "Technical Specification for Adjustable Load Grid Connection and Control", the present application selects three indexes of adjustable persistence, adjustable rate and adjustable capacity for quantifying the physical regulation ability of users participating in demand response. The present application adopts a data-driven method to approximately model based on historical power curves. To ensure that the label has practical scheduling significance, the users in the candidate set are further subjected to the following three physical constraint conditions of adjustable ability: Adjustable persistence: the adjustable time proportion is not less than 30% of the total sampling time, which measures the stability and endurance of the regulation behavior.
[0064] Adjustable rate: the maximum change rate of load per unit time is not less than 5% / min, which reflects the speed ability of users responding to grid dispatching instructions.
[0065] Adjustable capacity ratio: the user can adjust the load to be not less than 10% of the rated capacity under typical operating conditions, reflecting the user's ability to participate in peak load shifting.
[0066] The meanings and calculation methods of each index are as follows: S2021, adjustable persistence First, the transferable rate calculation interval, adjustable interval, and post-transfer recovery rate calculation interval need to be determined.
[0067] The determination steps are as follows: Determine the load change direction. Set the variable and the threshold value , the variable records the positive and negative of . Read . Determine the positive and negative, if non-negative, assign to 1, otherwise, assign to -1.
[0068] Identify the adjustment interval, draw the value of as shown in Figure 3 , define the interval where the number of monitoring points that persist for 1 is greater than or equal to the threshold value as interval 1, i.e., the transferable rate calculation interval; define the interval where the number of monitoring points that persist for -1 is greater than or equal to the threshold value as interval 3, i.e., the post-transfer recovery rate calculation interval; define the interval between the two as interval 2, i.e., the transferable interval.
[0069] Calculate the adjustable period ratio. The adjustable duration ratio is the proportion of interval 2 in the total duration, which is a positive index, and the expression is: (19) wherein is the adjustable duration ratio, , are the sampling point numbers of interval 2 and the whole interval, respectively.
[0070] S2022, adjustable rate Calculate the maximum ramp rate of the user's historical load, and represent it as "maximum rise rate percentage of equivalent rated capacity", which is usually required to be not less than 1% rated capacity / minute. Since the maximum load of most industrial and commercial users is close to their design capacity, and the maximum load of a charging station is approximately equal to the full-load power of all charging piles, this project takes the 95% quantile load as the "equivalent rated capacity".
[0071] Maximum ramp rate The maximum increase of load in the transferable rate calculation interval is defined as: (20) The sampling point number of interval 1 is defined, and further, the adjustable rate ratio is defined as the percentage of the maximum increase rate and the equivalent rated capacity is: (21) S2023, adjustable capacity ratio To further identify the adjustable potential of users from historical load behavior, the invention focuses on the high operating interval in the typical daily load curve-interval 2 (i.e. the continuous running period before the load enters the stable high position after the continuous rise, usually corresponding to the daytime peak of electricity consumption). In this interval: S20231, estimate the load space that can be reduced by the user from the historical load, use the low load persistence method to find the period when the load is lower than a certain threshold (such as 70% ), to judge whether the load is significantly deviated from its typical high operating level; S20232, identify the period in interval 2 when the load is continuously lower than the threshold, calculate the average load of these periods; S20233, define the adjustable capacity of the user in the peak period as: (22) This index reflects the user's ability to maintain a lower load operation in the typical peak period, i.e. its potential peak shaving adjustment capacity; Further, the adjustable capacity ratio is defined as: (23) The larger the ratio, the more significant the load reduction space of the user in the peak period, indicating higher demand response participation potential.
[0072] S203, label generation: mode and criterion collaborative decision A two-stage label construction mechanism of "candidate set screening + criterion verification" is adopted: S2031, candidate set generation: analyze the load pattern of each cluster (such as whether there is an obvious peak / valley), determine a number of "high potential candidate categories", and the users constitute the candidate set; S2032, label determination: only when the user meets the following conditions at the same time: belongs to the high potential candidate category; At the same time, meet the three adjustable capacity criteria; be marked as y=1 (high potential adjustable), otherwise y=0.
[0073] Mathematically, it is expressed as: (24) S3, Key feature identification and sample labeling optimization S301, Model task definition This study models the user adjustment potential identification problem as a binary classification supervised learning task: Input feature vector: The multi-dimensional load feature system constructed in the first phase is adopted, including four categories of features: electricity consumption scale, electricity fluctuation, electricity regularity, and external sensitivity, a total of M dimensions, denoted as .
[0074] Output: The "high potential user" label constructed in the second phase is used as the model output, denoted as , where indicates that the user has significant load adjustment capability, indicates low adjustment potential.
[0075] Model: Random Forest Classifier The model aims to learn the mapping relationship from the input features to the potential label , and in the process, quantifies the contribution of each feature to the classification decision.
[0076] S302, Feature importance analysis Gini importance (Gini Importance) is used to evaluate the contribution of each feature to the label, and then the feature importance ranking is output.
[0077] S303, Load sample labeling system Instead of a single label, this phase outputs a structured load sample labeling system that supports on-demand combination and dynamic update, including: S3031, Basic label layer: the behavior pattern category (such as 'double peak production type') that the user belongs to, and the physical ability compliance situation; S3032, Force level layer: 'high / medium / low' response potential level calculated according to feature importance and scenario weight; S3033 Semantic rule layer: interpretable composite labels generated by combining high importance features and clustering results, such as "if belongs to double peak type and temperature sensitivity > 0.6 → high temperature high potential adjustable".
[0078] Supports quick retrieval, combination, and editing of samples by label, serving intelligent scene construction.
[0079] In still another embodiment of the present application, a load-side data sample labeling system integrating data patterns and adjustment criteria is provided, which can be used to implement the load-side data sample labeling method integrating data patterns and adjustment criteria. Specifically, the load-side data sample labeling system integrating data patterns and adjustment criteria comprises a feature module, a sample module, and a labeling module.
[0080] The feature module is configured to construct a load feature set for characterizing user electricity consumption behavior and response characteristics to external environment, and the load features include electricity consumption scale features, electricity consumption fluctuation features, electricity consumption regularity features, and external factor sensitivity features. The sample module is configured to first group users through secondary AP clustering, identify candidate groups with adjustable behavior characteristics, and then verify adjustment capabilities of individual users in the candidate groups in combination with three physical criteria of adjustable persistence, adjustable rate, and adjustable capacity ratio, to construct a binary classification response potential label. The labeling module is configured to model the user adjustment potential identification problem as a binary classification supervised learning task, train the binary classification supervised learning task based on the constructed load feature set and the constructed binary classification response potential label, identify key features through feature importance analysis, and output a structured load sample labeling system.
[0081] The feature module comprises: The electricity consumption scale feature extraction unit is configured to calculate average loads or ratios of electricity consumption to time at each time in the observation period, determine maximum or minimum values of the loads in the observation period, or obtain maximum or minimum values of the loads by using maximum 3-day average loads to describe peak value characteristics of the electricity loads. The electricity consumption fluctuation feature extraction unit is configured to calculate a peak-valley difference rate by calculating a ratio of daily load peak-valley difference to daily maximum load, and calculate a load rate by calculating a ratio of average load to maximum load in the observation period. The electricity consumption regularity feature extraction unit is configured to calculate a flexibility coefficient by calculating average fluctuation levels of daily electricity consumption load patterns in a certain time scale, and calculate a regularity coefficient by using Spearman correlation coefficients to calculate consistency and stability of electricity consumption behavior in the certain time scale. The external sensitivity feature extraction unit is configured to obtain temperature sensitivity coefficients, humidity sensitivity coefficients, and price sensitivity coefficients by using Spearman correlation coefficients to measure consistency degrees of daily load curves of users and change trends of temperature, humidity, and price curves, and obtain a holiday sensitivity coefficient by calculating changes of total electricity consumption loads on holidays and weekdays.
[0082] The sample module comprises: A secondary clustering unit is configured to obtain a typical daily load curve by averaging multi-day load data of each user, perform score normalization processing, and perform electricity usage mode clustering by using an AP clustering algorithm combined with cosine similarity measurement. After excluding electricity usage mode anomalies or special users, comprehensive characteristic clustering is performed by using an AP clustering algorithm combined with Euclidean distance measurement to obtain a user candidate set with adjustable potential. A physical criterion verification unit is configured to determine load change direction and intervals, calculate adjustable duration proportion to verify adjustable persistence, calculate maximum rise rate percentage of equivalent rated capacity to verify adjustable rate, determine adjustable capacity and calculate adjustable capacity ratio to verify adjustable capacity ratio. A label generation unit is configured to analyze load patterns of each clustering cluster to determine a high-potential candidate category, and generate a binary classification response potential label according to whether a user belongs to the high-potential candidate category and whether three physical criteria are met.
[0083] The marking module comprises: A model construction unit is configured to use a multi-dimensional load feature system as an input feature vector, use a binary classification response potential label as model output, and use a random forest classifier to construct a binary classification supervised learning model. A feature importance analysis unit is configured to evaluate the contribution of each input feature to model classification decision by using Gini importance, output feature importance ranking results, and identify key features. A label system construction unit is configured to construct a structured load sample labeling system comprising a basic label layer, a potential level layer, and a semantic rule layer based on feature importance analysis results and secondary AP clustering results.
[0084] The label system construction unit comprises: A basic label generation subunit is configured to record behavior mode categories obtained by secondary clustering and physical capability compliance obtained by physical criterion verification. A potential level division subunit is configured to determine key feature weights according to feature importance, calculate user response potential comprehensive scores combined with scene weights, and divide response potential levels according to the scores. A semantic label generation subunit is configured to filter high-importance features and determine their effective value ranges, establish association rules between high-importance feature values and behavior mode categories, and generate interpretable composite labels.
[0085] The application provides a terminal device, which comprises a processor and a memory, the memory is used for storing a computer program, the computer program comprises program instructions, and the processor is used for executing the program instructions stored in the computer storage medium. The processor can be a central processing unit (CPU), and can also be other general-purpose processors, graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components and the like, which are the computing core and control core of the terminal, and are suitable for implementing one or more instructions, and are specifically suitable for loading and executing one or more instructions to implement a corresponding method flow or a corresponding function; the processor in the embodiment of the application can be used for the operation of the load side data sample labeling method fusing data mode and adjustment criterion, comprising: A load feature set for describing user power consumption behavior and response characteristics to external environment is constructed, the load features include power consumption scale features, power consumption fluctuation features, power consumption law features and external factor sensitivity features; a user is grouped by secondary AP clustering, a candidate group with adjustable behavior characteristics is identified, and then individual users in the candidate group are verified for adjustment capability by combining three physical criteria of adjustable persistence, adjustable rate and adjustable capacity ratio, to construct a binary classification response potential label; a user adjustment potential identification problem is modeled as a binary classification supervised learning task, the binary classification supervised learning task is trained based on the constructed load feature set and the constructed binary classification response potential label, key features are identified through feature importance analysis, and a structured load sample labeling system is output.
[0086] Please refer to Figure 5 , the terminal device is a computer device, the computer device 60 of the embodiment comprises a processor 61, a memory 62 and a computer program 63 stored in the memory 62 and capable of running on the processor 61, and the computer program 63 realizes the method for estimating the concentration of radioactive iodine species in the containment after an accident in the embodiment when executed by the processor 61, to avoid repetition, which will not be described here. Alternatively, the computer program 63 realizes the functions of each model / unit in the load side data sample labeling system fusing data mode and adjustment criterion when executed by the processor 61, to avoid repetition, which will not be described here.
[0087] The computer device 60 can be a desktop computer, a notebook computer, a palm computer, a cloud server, and the like. The computer device 60 can include, but is not limited to, a processor 61, a memory 62. Those skilled in the art can understand that the processor 61 and the memory 62 can be connected through a bus, and the bus can be a peripheral component interconnect (PCI) bus, a serial advanced technology attachment (SATA) bus, a universal serial bus (USB), or the like. Figure 5 The computer device 60 is only an example and does not constitute a limitation on the computer device 60, and can include more or fewer components than shown, or combine certain components, or different components, for example, the computer device can also include an input / output device, a network access device, a bus, and the like.
[0088] The processor 61 can be a central processing unit (CPU), and can also be other general-purpose processors, graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, and the like. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0089] The memory 62 can be an internal storage unit of the computer device 60, such as a hard disk or a memory of the computer device 60. The memory 62 can also be an external storage device of the computer device 60, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, and the like.
[0090] Further, the memory 62 can include both an internal storage unit and an external storage device of the computer device 60. The memory 62 is used to store computer programs and other programs and data required by the computer device. The memory 62 can also be used to temporarily store data that has been output or will be output.
[0091] Please refer to Figure 6The terminal device is an electronic device 600, which is in the form of a general computing device. Components of the electronic device can include, but are not limited to, at least one processing unit 610, at least one storage unit 620, a bus 630 that connects the different platform components, including the storage unit 620 and the processing unit 610, a display unit 640, and the like.
[0092] The storage unit stores program code that can be executed by the processing unit 610 to cause the processing unit 610 to perform the steps described in the method part of the specification above according to various exemplary embodiments of the present application. For example, the processing unit 610 can perform the steps shown in FIG. 6. Figure 1
[0093] The storage unit 620 can include a readable medium in the form of volatile storage such as a random access memory (RAM) 6201 and / or cache memory 6202, and can further include a read-only memory (ROM) 6203.
[0094] The storage unit 620 can also include program / utility 6204 having a set of at least one program modules 6205, including but not limited to, an operating system, one or more application programs, other program modules, and program data, each or some combination thereof, which can include implementation of a network environment.
[0095] The bus 630 can be representative of one or more of several types of bus structures, including a storage bus or bus controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of a variety of bus structures.
[0096] The electronic device 600 can also communicate with one or more external devices 700 such as a keyboard or pointing device, a Bluetooth device, etc.; other devices that enable a user to interact with the electronic device 600; and / or one or more devices that enable the electronic device 600 to communicate with one or more other computing devices. Such communication can be via the input / output interface 650. The electronic device 600 can further communicate with one or more networks, such as a local area network, a wide area network, and / or the public switched telephone network, such as the Internet, via a network adapter 660. The network adapter 660 can communicate with the other components of the electronic device 600 via the bus 630. It should be appreciated that although the network adapter 660 is shown as a single component, the network adapter 660 can comprise a plurality of components that work in cooperation to provide the functionality described herein. It should also be appreciated that although not shown, other hardware and / or software components that are commonly used in computing devices, such as microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc. can be used with the electronic device 600.
[0097] Embodiment 4 The present application further provides a storage medium, specifically a computer readable storage medium, which is a memory device in the terminal device, used for storing programs and data. It can be understood that the computer readable storage medium herein can include the built-in storage medium in the terminal device, and of course can include the expansion storage medium supported by the terminal device, and can be any tangible medium containing or storing programs, which can be used by or in combination with the instruction execution system, device or apparatus. The computer readable storage medium provides a storage space, which stores the operating system of the terminal. Moreover, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, which can be one or more computer programs (including program codes). It should be noted that more specific examples of the computer readable storage medium include an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0098] The computer readable storage medium further includes a data signal carried in the baseband or as a part of a carrier wave, in which readable program codes are borne. Such a propagated data signal can take various forms, including but not limited to electro-magnetic signal, optical signal or any suitable combination of the above. The readable storage medium can also be any readable medium other than the readable storage medium, which can send, propagate or transmit programs for use by or in combination with the instruction execution system, device or apparatus. The program codes contained in the readable storage medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical cable, radio frequency, etc. or any suitable combination of the above.
[0099] The program code for performing the operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, C++, etc., and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server. In the latter scenario, the remote computing device can be connected to the user's computing device through any type of network, including a local area network or a wide area network, or the connection can be made to an external computing device, such as through the Internet using an Internet Service Provider.
[0100] The one or more instructions stored in the computer readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the load side data sample labeling method related to the fusion of the data mode and the adjustment criterion in the above embodiments; the one or more instructions stored in the computer readable storage medium are loaded and executed by the processor to implement the following steps: A load feature set for describing user electricity consumption behavior and response characteristics to external environment is constructed, and the load features include electricity consumption scale features, electricity consumption fluctuation features, electricity consumption law features and external factor sensitivity features; a user is first grouped through secondary AP clustering, a candidate group with adjustable behavior characteristics is identified, then an individual user in the candidate group is verified for adjustment capability by combining three physical criteria of adjustable persistence, adjustable rate and adjustable capacity ratio, and a binary classification response potential label is constructed; a user adjustment potential identification problem is modeled as a binary classification supervised learning task, the binary classification supervised learning task is trained based on the constructed load feature set and the constructed binary classification response potential label, key features are identified through feature importance analysis, and a structured load sample labeling system is output.
[0101] The database involved in each embodiment provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a blockchain, and the like, without being limited thereto. The processor involved in each embodiment provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, and the like, without being limited thereto.
[0102] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application but not all embodiments of the present application. The components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work are within the scope of protection of the present application.
[0103] Please refer to Figure 4, based on the historical load data of a provincial power grid, the load samples are identified and labeled by the method. From about 1000 industrial and commercial users, about 500 candidate users with typical adjustable behavior patterns are first screened out by two-stage AP clustering; then combined with three physical criteria of adjustable persistence, adjustable rate and adjustable capacity ratio, about 200 high-potential adjustable load users are further identified; finally, based on the random forest model output feature importance ranking, regularity coefficient, price sensitivity, temperature sensitivity and flexibility coefficient are identified as key driving factors, and structured labeling results including "behavior mode class", "price sensitivity class", "response potential level" and "semantic composite label" are generated for all users in the sample library, realizing systematic feature labeling of load samples, supporting fast retrieval and combination according to labels, and serving the digital construction of power grid regulation scene.
[0104] In summary, the load-side data sample labeling method fusing data mode and adjustment criterion has the following characteristics: The load sample identification is more accurate: the two-stage AP clustering is used to realize the fine grouping of user behavior patterns, and the three physical criteria of adjustable persistence, rate and capacity ratio are combined to verify the ability, effectively excluding the misjudgment users with "large power consumption but rigidity", avoiding subjective labeling depending on artificial experience, and significantly improving the identification accuracy and engineering practicability of high-potential adjustable load; The label engineering practicability is enhanced: by introducing the quantitative criteria of adjustable persistence, adjustable rate and adjustable capacity ratio, it is ensured that the users labeled as "high potential" have the physical ability to actually participate in demand response; The feature-potential relationship is interpretable: based on the binary classification potential label, the random forest model is trained, the feature importance ranking is output, the temperature sensitivity and regularity coefficient are identified as key driving factors, and the transformation from "black box identification" to "interpretable labeling" is realized; The method has strong reusability: the constructed feature system and discrimination process form a standardized framework, and after new load data is connected, feature extraction, pattern recognition and label generation can be automatically completed, supporting dynamic updating of the sample library and cross-scene migration application.
Claims
1. A load side data sample labeling method that fuses data patterns and regulation criteria, characterized by, The method comprises the following steps: S1, constructing a load feature set for characterizing user electricity consumption behavior and external environment response characteristics, the load features including electricity consumption scale features, electricity consumption fluctuation features, electricity consumption regularity features, and external factor sensitivity features; S2, first clustering users through secondary AP clustering to identify candidate groups with adjustable behavior characteristics, then verifying the adjustment capability of individual users in the candidate groups in combination with three physical criteria of adjustable persistence, adjustable rate, and adjustable capacity ratio to construct a binary classification response potential label; S3, modeling the user adjustment potential identification problem as a binary classification supervised learning task, training the binary classification supervised learning task based on the load feature set constructed in step S1 and the binary classification response potential label constructed in step S2, identifying key features through feature importance analysis, and outputting a structured load sample labeling system.
2. The method of claim 1, wherein the method further comprises: The specific process of constructing a multi-dimensional load feature system comprises: The electricity consumption scale features include average load and extreme value load, wherein the average load is obtained by calculating the average value of the load at each time during the observation period or the ratio of the electricity consumption to time, and the extreme value load is obtained by determining the maximum or minimum value of the load during the observation period or using the maximum 3-day average load to describe the peak value characteristics of the power load; The electricity consumption fluctuation features include peak-valley difference rate and load rate, the peak-valley difference rate is obtained by calculating the ratio of the daily load peak-valley difference to the daily maximum load, and is used to reflect the degree of daily load peak-valley change, and the load rate is obtained by calculating the ratio of the average load to the maximum load during the observation period, and is used to reflect the imbalance degree of load distribution; The electricity consumption regularity features include flexibility coefficient and regularity coefficient, the flexibility coefficient is obtained by calculating the average fluctuation level of daily electricity consumption load mode within a certain time scale, and the regularity coefficient is obtained by calculating the consistency and stability of electricity consumption behavior within a certain time scale using the Spearman correlation coefficient; The external factor sensitivity features include temperature sensitivity coefficient, humidity sensitivity coefficient, electricity price sensitivity coefficient, and holiday sensitivity coefficient, the temperature sensitivity coefficient, humidity sensitivity coefficient, and electricity price sensitivity coefficient are obtained by using the Spearman correlation coefficient to measure the consistency degree of the daily load curve of the user with the change trend of the temperature, humidity, and electricity price curve, and the holiday sensitivity coefficient is obtained by calculating the change of total electricity consumption load on holidays and weekdays.
3. The method of claim 1, wherein the method further comprises: The specific process of secondary AP clustering comprises: First stage: obtaining a typical daily load curve by averaging the multi-day load data of each user, performing score normalization processing on the typical daily load curve to eliminate the absolute difference in total electricity consumption between users, and performing clustering on the normalized typical daily load curve by using the AP clustering algorithm combined with cosine similarity measurement to output several typical behavior mode clusters; Second stage: based on the clustering results of the first stage, eliminating users with abnormal or special electricity consumption modes, performing re-clustering on the filtered user load data by using the AP clustering algorithm combined with Euclidean distance measurement, and realizing the differentiation of user electricity consumption scale based on similar behavior modes to obtain a candidate set of users with adjustable potential.
4. The method of claim 1, wherein the method further comprises: The specific process of verifying the adjustment capability of individual users in combination with the three physical criteria comprises: S2021, determine the load change direction, identify the transferable rate calculation interval, adjustable interval, transfer after the recovery rate calculation interval, calculate the adjustable duration ratio, when the adjustable duration ratio is not less than 30% of the total sampling duration, it is determined that the user meets the adjustable duration requirement; S2022, calculate the maximum climbing rate of the user's historical load, the maximum climbing rate is represented by the percentage of the equivalent rated capacity, when the adjustable rate is not less than 5% / min, it is determined that the user meets the adjustable rate requirement, wherein the equivalent rated capacity is the 95% quantile load of the user; S2023, identify the high operating interval of the load typical peak period, find out the period when the load is lower than the preset threshold in the interval and calculate the average load of these periods, determine the adjustable capacity of the user in the peak period, calculate the ratio of the adjustable capacity to the equivalent rated capacity to obtain the adjustable capacity ratio, when the adjustable capacity ratio is not less than 10%, it is determined that the user meets the adjustable capacity ratio requirement.
5. The method of claim 4, wherein the step of fusing the data patterns and the regulation criterion comprises: In step S2021, the specific process of identifying the transferable rate calculation interval, adjustable interval and transfer after the recovery rate calculation interval is as follows: Set variables and thresholds, the variable records the positive and negative of the load change direction, read the load change direction variable, if the variable is non-negative, assign it to 1, otherwise assign it to -1; The interval with variable value continuously equal to 1 and monitoring point number greater than or equal to threshold is defined as the transferable rate calculation interval, the interval with variable value continuously equal to -1 and monitoring point number greater than or equal to threshold is defined as the transfer after the recovery rate calculation interval, and the interval between the transferable rate calculation interval and the transfer after the recovery rate calculation interval is defined as the adjustable interval.
6. The method of claim 4, wherein the step of fusing the data patterns and the regulation criterion with the load side data sample signatures further comprises: In step S2023, the specific process of determining the adjustable capacity of the user in the peak period is as follows: Using the low load persistence method, find out the period when the load is continuously lower than the preset threshold in the high operating interval of the load typical peak period, and calculate the average load of these periods; The difference between the average load in the high operating interval and the average load calculated above is determined as the adjustable capacity of the user in the peak period.
7. The method of claim 1, wherein the method further comprises: The specific process of constructing the binary classification response potential label is as follows: Analyze the load shape of each cluster obtained by secondary AP clustering, determine several high-potential candidate categories with adjustable behavior characteristics, and the users in the high-potential candidate categories constitute the candidate set; If the user belongs to the high-potential candidate category and meets the adjustable duration, adjustable rate and adjustable capacity ratio physical criteria at the same time, the user is marked as a high-potential adjustable load, and the label is marked as y=1; if the user does not meet any of the above conditions, the user is marked as a low-potential adjustable load, and the label is marked as y=0.
8. The method of claim 1, wherein the method further comprises: The specific process of key feature identification and sample labeling optimization includes: S301, use the multi-dimensional load feature system constructed in step S1 as the input feature vector, use the binary classification response potential label constructed in step S2 as the model output, and use the random forest classifier to construct a binary classification supervised learning model; S302, evaluate the contribution of each input feature to the model classification decision using Gini importance, output the feature importance ranking result, and identify the key features that have a significant impact on the user adjustment potential; S303, based on the feature importance analysis result and the secondary AP clustering result, a structured load sample labeling system including a basic label layer, a potential level layer, and a semantic rule layer is constructed, wherein the basic label layer records the behavior mode category and the physical capability compliance of the user, the potential level layer divides the response potential level of the user according to the feature importance and the scene weight, and the semantic rule layer generates an interpretable composite label by combining the high importance features and the clustering result.
9. The method of claim 8, wherein the step of fusing the data patterns and the regulation criterion comprises: In step S303, the division process of the response potential level is as follows: According to the feature importance analysis result, the weights of the key features are determined, and the scene weights are set in combination with different power grid operation scenes; based on the values of the user on the key features, the response potential comprehensive score of the user is calculated in combination with the key feature weights and the scene weights; According to the distribution of the response potential comprehensive score, the response potential of the user is divided into high, medium and low levels, and the labeling content of the potential level layer is formed.
10. The method of claim 8, wherein the step of fusing the data patterns and the regulation criterion with the load side data sample signatures further comprises: In step S303, the generation process of the interpretable composite label of the semantic rule layer is as follows: High importance features are screened according to the importance ranking, and the effective value range of each high importance feature is determined; the association rules between the high importance feature values and the behavior mode categories are established in combination with the user behavior mode categories obtained by the secondary AP clustering; Based on the association rules, an interpretable composite label is generated, which can reflect the comprehensive influence of the high importance features and the behavior mode categories of the user on the response potential.
11. A load side data sample tagging system that fuses data patterns with regulatory criteria, characterized by, It includes: A feature module constructs a load feature set for describing the user's electricity consumption behavior and the response characteristics to the external environment, including electricity consumption scale features, electricity consumption fluctuation features, electricity consumption regularity features, and external factor sensitivity features; A sample module first groups users through secondary AP clustering, identifies candidate groups with adjustable behavior characteristics, and then verifies the adjustment ability of individual users in the candidate groups based on three physical criteria: adjustable persistence, adjustable rate, and adjustable capacity ratio, to construct a binary classification response potential label; A labeling module models the user adjustment potential identification problem as a binary classification supervised learning task, trains the binary classification supervised learning task based on the load feature set constructed in step S1 and the binary classification response potential label constructed in step S2, identifies key features through feature importance analysis, and outputs a structured load sample labeling system.
12. The load side data sample tagging system that fuses data patterns with regulation criteria of claim 11, wherein, The feature module includes: An electricity consumption scale feature extraction unit is configured to calculate the average load or the ratio of electricity consumption to time at each time during the observation period to obtain the average load, determine the maximum or minimum value of the load during the observation period, or obtain the maximum or minimum value of the load by using the maximum 3-day average load to describe the peak value characteristics of the power load; An electricity consumption fluctuation feature extraction unit is configured to calculate the ratio of the daily load peak valley difference to the daily maximum load to obtain the peak valley difference rate, and calculate the ratio of the average load to the maximum load during the observation period to obtain the load rate; An electricity consumption regularity feature extraction unit is configured to calculate the average fluctuation level of the daily electricity consumption load mode in a certain time scale to obtain a flexibility coefficient, and use the Spearman correlation coefficient to calculate the consistency and stability of the electricity consumption behavior in a certain time scale to obtain a regularity coefficient; An external sensitivity feature extraction unit is configured to measure the consistency degree of the daily load curve of a user with the temperature, humidity, and electricity price curve trend by using a Spearman correlation coefficient to obtain a temperature sensitivity coefficient, a humidity sensitivity coefficient, and an electricity price sensitivity coefficient, and to obtain a holiday sensitivity coefficient by calculating the change of total electricity load on holidays and weekdays.
13. The load side data sample tagging system that fuses data patterns with regulation criteria of claim 11, wherein, The sample module includes: A secondary clustering unit is configured to average the multi-day load data of each user to obtain a typical daily load curve and perform score normalization processing, to perform electricity consumption mode clustering by using an AP clustering algorithm combined with cosine similarity measurement, to perform comprehensive characteristic clustering by using an AP clustering algorithm combined with Euclidean distance measurement after excluding electricity consumption mode anomalies or special users, and to obtain a user candidate set with adjustable potential. A physical criterion verification unit is configured to determine the load change direction and each interval, to calculate the adjustable duration ratio to verify the adjustable continuity, and to calculate the percentage of the maximum rising rate to the equivalent rated capacity to verify the adjustable rate, and to determine the adjustable capacity and calculate the adjustable capacity ratio to verify the adjustable capacity ratio. A label generation unit is configured to analyze the load pattern of each clustering cluster to determine a high-potential candidate category, and to generate a binary classification response potential label according to whether the user belongs to the high-potential candidate category and whether the three physical criteria are met.
14. The load side data sample tagging system that fuses data patterns with regulation criteria of claim 11, wherein, The marking module includes: A model construction unit is configured to use a multi-dimensional load feature system as an input feature vector and a binary classification response potential label as a model output, and to construct a binary classification supervised learning model by using a random forest classifier. A feature importance analysis unit is configured to evaluate the contribution of each input feature to the model classification decision by using Gini importance, to output a feature importance ranking result, and to identify key features. A marking system construction unit is configured to construct a structured load sample marking system including a basic label layer, a potential level layer, and a semantic rule layer based on the feature importance analysis result and the secondary AP clustering result.
15. The load side data sample tagging system that fuses data patterns with regulation criteria of claim 14, wherein, The marking system construction unit includes: A basic label generation subunit is configured to record the behavior mode category obtained by secondary clustering and the physical ability compliance obtained by physical criterion verification. A potential level division subunit is configured to determine the key feature weight according to the feature importance, to calculate the user response potential comprehensive score in combination with the scene weight, and to divide the response potential level according to the score. A semantic label generation subunit is configured to filter high-importance features and determine their effective value range, to establish the association rules between the high-importance feature values and the behavior mode categories, and to generate interpretable composite labels.
16. A computer-readable storage medium storing one or more programs, the one or more programs comprising instructions for: The one or more programs include instructions that, when executed by a computing device, cause the computing device to perform the method of any one of claims 1-10.
17. A computing device, comprising: The one or more processors, the memory, and the one or more programs are included, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing the steps in the method of any one of claims 1-10.