An intelligent classification method based on rainfall sensitivity and decision tree

By calculating the sensitivity index of user electricity consumption relative to changes in rainfall and fusing environmental attribute data, an environmentally robust decision tree model is constructed, which solves the problem of insufficient user classification accuracy and stability in existing technologies and achieves high-precision user identification under different climatic conditions.

CN120974287BActive Publication Date: 2026-01-23STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511508123.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-01-23
Estimated Expiration
2045-10-22

AI Technical Summary

Technical Problem

Existing technologies are insufficient in quantitatively describing the dynamic relationship between rainfall and electricity consumption behavior in user classification. Traditional methods are difficult to adapt to the diverse changes in user behavior under different seasons and extreme climate conditions, resulting in insufficient classification accuracy and stability.

Method used

By calculating the sensitivity index of user electricity consumption relative to changes in rainfall, a sensitivity feature vector is constructed and fused with environmental attribute data to build a decision tree model based on the Gini coefficient splitting criterion. Cross-seasonal sample balancing and rainfall anomaly compensation mechanisms are introduced to form an environmentally robust training model.

Benefits of technology

It achieves high accuracy and stability in user classification under different climatic conditions, can dynamically identify user categories, reduce the need for manual intervention, and adapt to scenarios such as power dispatching and demand-side management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974287B_ABST
    Figure CN120974287B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of intelligent classification method based on rainfall sensitivity and decision tree.The method carries out standardization processing to multi-source original data, constructs data set;Calculate the sensitivity index of the change of user electricity consumption relative to rainfall, and generate sensitivity feature vector;Characteristic vector is fused with seasonal characteristics and crop type characteristics, and the sensitivity comprehensive feature set for describing the dynamic response of user electricity consumption behavior to climatic factor is obtained;The comprehensive feature set is input into the decision tree model based on gini coefficient split criterion, and the cross-seasonal sample balance mechanism and rainfall anomaly compensation mechanism are introduced in the training process, to obtain the classification model with environmental robustness;New multi-source data is automatically identified by the model.The present application can effectively improve the accuracy and adaptability of user classification, avoid the shortcomings of traditional methods relying on artificial experience or fixed rules, and has wide application value in power dispatching, differentiated electricity price and demand side management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of machine learning technology, and in particular to an intelligent classification method based on rainfall sensitivity and decision trees. Background Technology

[0002] In existing technologies, power companies and research institutions commonly utilize users' historical electricity consumption data, meteorological data, and some environmental information to classify and analyze users. Common methods include cluster-based load feature identification, rule-based threshold determination, and the application of traditional decision tree models to segment users based on behavioral patterns. These methods can, to some extent, reveal the impact of weather factors on electricity load, such as distinguishing user groups through temperature, humidity, or seasonal indicators, and have found some application in demand-side management, energy-saving dispatch, and electricity price optimization.

[0003] However, the aforementioned existing methods generally suffer from problems such as coarse data processing granularity, insufficient response to meteorological factors, and limited robustness of classification models. Specifically, most existing studies focus on the impact of common meteorological factors such as temperature, with few quantitative descriptions of the dynamic relationship between rainfall and electricity consumption behavior; traditional classification methods rely on fixed thresholds or empirical rules, making it difficult to adapt to the diverse changes in user behavior under different seasons and extreme climate conditions; furthermore, existing models often exhibit decreased classification accuracy when cross-seasonal sample distribution is uneven or when abnormal rainfall data exists, making it difficult to meet the accuracy and stability requirements of power systems for user classification in complex environments.

[0004] Therefore, there is an urgent need to propose a new intelligent classification method. Summary of the Invention

[0005] This application provides an intelligent classification method based on rainfall sensitivity and decision tree to improve the accuracy and adaptability of user classification.

[0006] This application provides an intelligent classification method based on rainfall sensitivity and decision trees, including:

[0007] Standardize the multi-source raw data, which includes user electricity consumption data, rainfall data, and environmental attribute data, to generate a standardized dataset with unified dimensions and scale constraints.

[0008] Based on the standardized dataset, a sensitivity index for user electricity consumption relative to changes in rainfall is calculated. The sensitivity index is obtained by comparing the differences between rainfall levels and corresponding user electricity consumption at different time periods, and a sensitivity feature vector is constructed with the sensitivity index as the core.

[0009] The sensitivity feature vector is fused with the seasonal and crop type features in the environmental attribute data to obtain a comprehensive sensitivity feature set, which is used to characterize the multidimensional dynamic response of user electricity consumption behavior to climate factors.

[0010] The sensitivity comprehensive feature set is input into a decision tree classification model constructed based on the Gini coefficient splitting criterion. During the training process, a cross-seasonal sample balancing mechanism and a rainfall anomaly compensation mechanism are introduced to form a training model with environmental robustness. The training model is used to learn the nonlinear mapping relationship between the sensitivity comprehensive feature set and the user category.

[0011] The training model is used to classify and identify new multi-source raw data. Without the need for manually setting fixed thresholds, dynamic and high-precision user category identification is achieved based on the user's sensitivity to changes in rainfall and related environmental factors.

[0012] The beneficial effects of the technical solution provided in this application include:

[0013] (1) This application calculates the sensitivity index of user electricity consumption relative to changes in rainfall and constructs a sensitivity feature vector based on it. This can quantitatively reflect the impact of meteorological factors on user electricity consumption behavior. Compared with traditional methods that rely on single meteorological factors such as temperature, the dimension is more comprehensive and helps to improve the accuracy of user classification. (2) This application integrates the sensitivity feature vector with seasonal features and crop type features to generate a comprehensive sensitivity feature set, thereby realizing the multi-dimensional dynamic response modeling of user electricity consumption behavior to climate factors. This can effectively distinguish different groups such as agricultural users and residential users, and solve the problem of insufficient identification ability of existing methods for specific user scenarios. (3) This application introduces a cross-seasonal sample balancing mechanism and a rainfall anomaly compensation mechanism in the training process of the decision tree classification model, so that the model has stronger environmental robustness and can maintain classification accuracy under conditions of uneven sample distribution and extreme rainfall, avoiding the distortion caused by training data bias in traditional models. (4) The classification model obtained by training in this application can achieve dynamic and high-precision discrimination of new user data without the need for manual setting of fixed thresholds. This significantly reduces the need for manual intervention, improves the automation and adaptability of the classification method, and is suitable for promotion and application in scenarios such as power dispatch, differentiated electricity pricing and demand-side management. Attached Figure Description

[0014] Figure 1 This is a flowchart of an intelligent classification method based on rainfall sensitivity and decision tree provided in the first embodiment of this application. Detailed Implementation

[0015] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.

[0016] The first embodiment of this application provides an intelligent classification method based on rainfall sensitivity and decision trees. Please refer to... Figure 1 This figure is a schematic diagram of the first embodiment of this application. The following is in conjunction with... Figure 1 The first embodiment of this application provides a detailed description of an intelligent classification method based on rainfall sensitivity and decision tree.

[0017] To quickly understand the specific application scenarios of this invention, let's first illustrate with an example. For instance, within the power supply area managed by a certain regional power company, there are two types of users: one type is agricultural users who mainly grow rice, and the other type is ordinary residential households. The power company wants to be able to automatically distinguish between these two types of users in order to implement time-of-use pricing strategies, optimize grid dispatch, and achieve precise demand-side management.

[0018] To this end, the power company collected multi-source raw data from various users, including daily electricity consumption data for the past two years, daily rainfall data provided by the local meteorological department, and basic environmental attribute data declared by users (such as whether they are engaged in agricultural production, the main crop types, and the seasonal climate characteristics corresponding to their geographical location).

[0019] The system first standardizes the raw data, converting data such as electricity consumption, rainfall, and seasonal indicators to a unified scale, thereby forming a standardized dataset to ensure comparability between different features in the subsequent modeling process.

[0020] Based on this, the system calculates a rainfall sensitivity index by comparing the differences in electricity consumption among users on "rainy days" and "dry days." For example, some agricultural users experience a significant decrease in electricity consumption during periods of continuous rainfall, while their electricity consumption increases significantly during periods of drought when irrigation equipment is activated; conversely, residential users show little difference in daily electricity consumption regardless of rainfall. The system quantifies this sensitivity index and further constructs a sensitivity feature vector, assigning each user a set of numerical features reflecting the "degree of impact of rainfall."

[0021] Next, the system fuses this sensitivity feature vector with the user's seasonal characteristics (such as summer and winter load differences) and crop type characteristics (such as water requirement patterns of crops like rice and corn) to obtain a comprehensive sensitivity feature set. This feature set can characterize the dynamic coupling relationship between the user's electricity consumption behavior and the climate environment in multiple dimensions.

[0022] During the model training phase, the system inputs the aforementioned comprehensive feature set into a decision tree classification model constructed based on the Gini coefficient splitting criterion. To prevent abnormal data bias caused by excessive concentration of training samples in a particular season or by extreme rainstorms, the system specifically introduces a cross-seasonal sample balancing mechanism and a rainfall anomaly compensation mechanism. This enhances the model's environmental robustness when learning the non-linear relationship between rainfall sensitivity and user category.

[0023] Ultimately, the model trained by the system can automatically identify new user data. For example, when a new user's electricity consumption and rainfall data are input, the model can determine whether the user is a typical agricultural irrigation user or an ordinary residential user. In actual operation, power companies can use this classification result to implement differentiated electricity pricing and dispatching strategies for agricultural users during the dry and flood seasons, while maintaining a stable power supply for residential users, thereby achieving dynamic and highly accurate user category identification.

[0024] Step S101: Standardize the multi-source raw data containing user electricity consumption data, rainfall data, and environmental attribute data to generate a standardized dataset with unified dimensions and scale constraints.

[0025] In step S101, the multi-source raw data, including user electricity consumption data, rainfall data, and environmental attribute data, needs to be standardized to generate a standardized dataset with unified dimensions and scale constraints. This step is the foundation of the entire method, and its purpose is to eliminate the differences in dimensions and value ranges between data from different sources, so that subsequent feature calculations and model training can be carried out under a unified numerical system, avoiding model bias or calculation distortion caused by numerical inconsistencies.

[0026] In the specific implementation process, it is first necessary to clarify the composition of the multi-source raw data. User electricity consumption data can include daily electricity consumption, hourly load curves, peak and valley electricity statistics, etc., recording the user's electricity consumption in different time periods in time series form. Rainfall data usually comes from meteorological monitoring systems and can record information such as total rainfall, duration of rainfall, and rainfall intensity in hourly, daily, or weekly units. Environmental attribute data can include the user's geographical location, industry type, whether it is an agricultural user, type of crop grown, and typical seasonal climate characteristics of the region. These data themselves often have different units of measurement and numerical ranges in different acquisition systems. For example, electricity consumption is expressed in kilowatt-hours, rainfall in millimeters, seasonal characteristics may be represented by qualitative labels or numerical codes, and crop type may be a discrete categorical variable.

[0027] In the standardization process, the first step is to map various types of data into numerical form. For non-numerical environmental attribute data, one-hot encoding or ordinal encoding can be used to convert them into numerical variables. Then, a unified standardization method is applied to all numerical variables. For example, linear normalization can be used to map all variables to the [0,1] interval, or zero-mean standardization can be used to convert the data into a distribution with a mean of 0 and a variance of 1. For time-series data, time alignment and interpolation can be performed on data from the same user at different time periods to ensure that rainfall and electricity consumption correspond on the same time scale. For missing values, mean imputation, nearest neighbor interpolation, or imputation strategies based on similar users can be used to ensure data integrity and continuity.

[0028] During the processing, it is also necessary to pay attention to the importance and stability of different features. For example, the fluctuation values ​​of user electricity consumption can be standardized by extracting their mean, variance, and peak-to-valley difference, while rainfall needs to be cumulatively calculated in conjunction with time windows to form a basis for comparison. In this way, data from different sources and with different dimensions are transformed into a numerical set under a unified standard, constructing a standardized dataset with unified dimensions and scale constraints. This dataset can not only contain single numerical features, but also standardized time series, categorical variable mapping values, and statistically processed combined features, thus providing an accurate, consistent, and comparable data foundation for the subsequent calculation of sensitivity indicators.

[0029] This step eliminates dimensional differences and scale biases among various types of raw data, ensuring that operations can be performed using a unified standard in subsequent sensitivity calculations and feature fusion processes.

[0030] Furthermore, the standardization process for the multi-source raw data, including user electricity consumption data, rainfall data, and environmental attribute data, to generate a standardized dataset with unified dimensions and scale constraints includes:

[0031] A unified time-series index is established for the multi-source raw data, and user electricity consumption data, rainfall data and environmental attribute data are synchronously aligned within the same time window to obtain a multi-source aligned dataset with time-series consistency.

[0032] The multi-source aligned dataset is subjected to climate pattern-based interval normalization. The interval normalization is dynamically adjusted according to the joint probability of historical rainfall distribution and user electricity consumption distribution, so that the feature values ​​under different seasons and different climate conditions are mapped to a unified scale, resulting in a normalized dataset with dynamic adaptability.

[0033] The normalized dataset with dynamic adaptability is subjected to feature decoupling based on mutual information constraints. The mutual information strength between user electricity consumption data, rainfall data and environmental attribute data is calculated. Feature dimensions with redundancy exceeding the threshold are removed, and independent contributing variables are retained through orthogonal mapping to generate a denoised key feature set.

[0034] The scale of the denoised key feature set is recalibrated based on environmental sensitivity factors. A scale coupling factor is constructed by combining rainfall intensity fluctuation rate and electricity consumption sensitivity. Key features from different sources are uniformly measured and corrected to generate a standardized dataset with uniform dimensions and scale constraints.

[0035] When standardizing multi-source raw data containing user electricity consumption data, rainfall data, and environmental attribute data, it is necessary to ensure that all data are consistent not only in time but also in numerical scale and physical meaning, thereby eliminating biases and incomparability between different data sources. This step first requires establishing a unified time-series index. A unified time-series index refers to aligning data from different data sources according to a fixed time window. For example, if the system selects a daily time window, it must ensure that the user's daily electricity consumption data, the rainfall data for the same day, and the corresponding environmental attribute data for that day strictly correspond to the same timestamp. If some data is missing, linear interpolation, nearest neighbor completion, or estimation methods based on statistical distribution can be used to fill in the gaps, ensuring that complete multi-source data records are obtained within each time window. After this processing step, a multi-source aligned dataset with temporal consistency is obtained, where each record contains electricity consumption, rainfall, and environmental attribute values ​​within the same time window, avoiding calculation errors caused by time mismatches between different data sources.

[0036] After obtaining the multi-source aligned dataset, it needs to be normalized based on climate patterns. Unlike conventional normalization which uses fixed intervals, the method of this invention dynamically adjusts the normalization interval based on the joint probability of historical rainfall distribution and user electricity consumption distribution. Specifically, the system first analyzes the joint distribution of rainfall and electricity consumption in the historical time series to obtain the probability distribution characteristics under different climate conditions and seasons. Then, it determines the upper and lower limits of the normalization interval based on this joint distribution. For example, a wider normalization interval may be used in the rainy season than in the dry season to ensure that changes in electricity consumption under extreme rainfall conditions can still be effectively mapped. The result of this processing is that the feature values ​​of different seasons and different climate conditions are mapped to a comparable uniform scale. After this processing, a dynamically adaptive normalized dataset is generated, which can reflect the true differences in data under different climate patterns rather than the distortion caused by normalization.

[0037] After obtaining a dynamically adaptive normalized dataset, further feature decoupling based on mutual information constraints is required. Mutual information is an indicator that measures the correlation and dependence between two variables, reflecting nonlinear relationships better than simple correlation coefficients. In this step, the system calculates the mutual information strength between user electricity consumption data, rainfall data, and environmental attribute data. If the redundancy between certain feature dimensions exceeds a set threshold, they will be removed to avoid duplicate information interfering with the model. While removing redundant features, the remaining variables are reconstructed through orthogonal mapping, making the new variables independent of each other and preserving the original data information to the greatest extent. The final result is a denoised key feature set, which is more concise in its information representation and reduces the interference caused by high correlation.

[0038] After obtaining the denoised key feature set, scale recalibration based on environmental sensitivity factors is required. The core of this process lies in constructing a scale coupling factor, which is jointly formed by rainfall intensity volatility and electricity consumption sensitivity, used to unify the measurement standards of different features. Rainfall intensity volatility can be obtained by calculating the variance or standard deviation of rainfall within a certain time window, used to characterize the severity of rainfall changes. Electricity consumption sensitivity can be calculated by analyzing the relative response of user electricity consumption to rainfall changes, such as the difference in electricity consumption under light and heavy rain conditions. By combining rainfall intensity volatility and electricity consumption sensitivity, a comprehensive scale coupling factor can be obtained, used to uniformly correct key features from different sources. The corrected features are not only numerically under the same dimensions and scale constraints, but also reflect the differences in environmental sensitivity of different features, thus ensuring that subsequent model training can be carried out under fair and consistent conditions. Finally, after this series of processes, the obtained standardized dataset has unified dimensions and scale constraints and can be directly used for the calculation of sensitivity indicators, ensuring the reliable implementation of subsequent steps in this invention.

[0039] For example, in a certain region during a summer month, the recorded daily rainfall data are: 10 mm, 15 mm, 0 mm, 40 mm, 20 mm, 25 mm, 0 mm. By calculating the standard deviation of the rainfall within this time window, the rainfall intensity volatility can be obtained. If the average rainfall is 15.7 mm and the standard deviation is 13.4 mm, then this standard deviation value represents the rainfall intensity volatility, indicating that the rainfall varies considerably within this window.

[0040] Within the same time window, a user's average daily electricity consumption was recorded, corresponding to the aforementioned number of rainy days: 100 kWh, 95 kWh, 110 kWh, 80 kWh, 90 kWh, 85 kWh, and 115 kWh. Electricity consumption sensitivity can be calculated by comparing the average electricity consumption on rainy days with that on dry days. Assuming the average electricity consumption on dry days (two days, each with 0 mm of rainfall) is 112.5 kWh, and the average electricity consumption on rainy days (five days) is 90 kWh, then the electricity consumption sensitivity is defined as the relative rate of change: (90 − 112.5) / 112.5 = −0.20, meaning that electricity consumption decreases by an average of 20% under rainy conditions.

[0041] Next, the rainfall intensity volatility and electricity consumption sensitivity are combined to form a scale coupling factor. The combination can be a product or a weighted average; this invention uses a product to enhance the amplification effect on sensitivity differences under extreme weather conditions. For example, if the rainfall intensity volatility is 13.4 and the sensitivity is −0.20, then the scale coupling factor = 13.4 × (−0.20) = −2.68. This value indicates that a negative amplification correction needs to be applied to the rainfall-related features in this user's feature set to highlight the significant decrease in electricity consumption under conditions of strong rainfall volatility.

[0042] Step S102: Based on the standardized dataset, calculate the sensitivity index of user electricity consumption relative to changes in rainfall. The sensitivity index is obtained by comparing the differences between rainfall levels and corresponding user electricity consumption at different time periods, and a sensitivity feature vector is constructed with the sensitivity index as the core.

[0043] In step S102, based on the standardized dataset obtained in the previous step, a sensitivity index for user electricity consumption relative to changes in rainfall needs to be calculated, and a sensitivity feature vector is constructed using this index as the core. This sensitivity index is a quantitative measure used to describe the degree to which user electricity consumption behavior responds under different rainfall conditions. It reflects the magnitude and direction of the difference in user electricity consumption with changes in rainfall, thereby revealing the user's dependence on or resistance to interference from climate conditions.

[0044] In the specific implementation process, the first step is to align the rainfall data and electricity consumption data in the standardized dataset according to a unified time scale. For example, on a daily basis, the daily rainfall is paired with the corresponding user's daily electricity consumption. For each user, the data can be divided into several control groups, each containing a "rainy period" and a "no-rain or low-rain period." The difference in electricity consumption between the two periods is calculated to reflect the impact of rainfall on the user's electricity consumption. If the electricity consumption during the rainy period is significantly lower than that during the no-rain period, it indicates that the user is highly sensitive to rainfall; conversely, if the change in electricity consumption is not significant, it indicates that the user is less sensitive to rainfall.

[0045] Sensitivity indices can be calculated using various methods, such as defining them as a rate of difference. This means the sensitivity value equals the difference between the average electricity consumption under rainfall conditions and the average electricity consumption under dry conditions, divided by the average electricity consumption under dry conditions, thus obtaining a relative rate of change. For example, if an agricultural user's average daily electricity consumption is 120 kWh during dry periods and 80 kWh during periods of continuous rainfall, their sensitivity index is calculated as (80 − 120) / 120 = −0.33, a negative number, indicating that rainfall reduces electricity consumption by approximately 33%. Conversely, if a residential user's average daily electricity consumption is 50 kWh during dry periods and 52 kWh during rainfall periods, their sensitivity index is (52 − 50) / 50 = 0.04, indicating that rainfall has a minimal impact on their electricity consumption, with only a 4% increase. Using this calculation method, each user can obtain a quantified sensitivity value.

[0046] After obtaining the sensitivity index, it is necessary to further construct a sensitivity feature vector. The feature vector not only contains a single sensitivity value but can also include multiple sensitivity components under different rainfall levels. For example, sensitivity values ​​under light, moderate, and heavy rain conditions can be calculated separately and combined into a single vector, thus forming a more hierarchical and fine-grained feature representation. For the agricultural user in the example above, if the sensitivity is -0.15 under light rain, -0.30 under moderate rain, and -0.45 under heavy rain, then its sensitivity feature vector can be represented as (-0.15, -0.30, -0.45). Such a feature vector can comprehensively describe the differences in the user's electricity consumption behavior under different rainfall intensities, enabling the model to more accurately distinguish user types.

[0047] In this way, sensitivity indices and sensitivity feature vectors can transform raw electricity consumption and rainfall data into structured numerical features that can be directly used for model training. This process not only preserves the key patterns of rainfall's impact on user electricity consumption behavior, but also eliminates dimensional differences and noise interference in the data through standardization and hierarchical processing.

[0048] Furthermore, based on the standardized dataset, a sensitivity index for user electricity consumption relative to changes in rainfall is calculated. This sensitivity index is obtained by comparing the differences in rainfall levels and corresponding user electricity consumption at different time periods. A sensitivity feature vector is constructed using this sensitivity index as the core, including:

[0049] The rainfall data in the standardized dataset is segmented into multiple levels of intensity, dividing the rainfall conditions into four levels: light rain, moderate rain, heavy rain, and extreme rainfall. The average rainfall level is calculated for each level interval to generate multi-level rainfall intensity reference values.

[0050] The user electricity consumption data in the standardized dataset is grouped and statistically analyzed accordingly. The average electricity consumption under different rainfall intensity reference values ​​is calculated and compared with the average electricity consumption under no-rain conditions to obtain a multi-level electricity consumption difference rate set.

[0051] The multi-level electricity consumption difference rate set is subjected to time-series windowing processing. The mean and variance of the difference rate are calculated within a set continuous time window to capture the stability of users' response to rainfall under short-term fluctuations and long-term trends, and to generate a windowed sensitivity curve.

[0052] The windowed sensitivity curve is normalized and offset corrected by introducing temperature and humidity data from environmental attributes as auxiliary variables and eliminating response bias caused by non-rainfall factors to obtain the corrected sensitivity sequence.

[0053] The corrected sensitivity sequence is vectorized and combined, and the corrected values ​​under light rain, moderate rain, heavy rain and extreme rainfall conditions are arranged in sequence to form a complete sensitivity feature vector. A stability index obtained by variance measurement is added to this vector to characterize the user's overall response to rainfall.

[0054] In the implementation of calculating the sensitivity index of user electricity consumption relative to changes in rainfall based on the standardized dataset, and constructing a sensitivity feature vector with the sensitivity index as the core, a four-level intensity division is first established based on the daily rainfall in the standardized dataset. For samples with daily rainfall greater than zero, threshold intervals for light rain, moderate rain, and heavy rain are determined according to the 25th, 50th, and 75th percentiles of the historical distribution. Samples exceeding the 75th percentile are defined as extreme rainfall, and samples equal to zero are used as a separate no-rain condition for comparison. Using historical samples from the same region over the past two years as the statistical range, the threshold calculation is directly completed within the standardized dataset to ensure that the thresholds for the same user in the same region are consistent and reproducible. After obtaining the segments, the daily rainfall is aggregated for each intensity interval, and the average rainfall for that interval is calculated. The average of the four intervals is the multi-level rainfall intensity reference value. If a user has fewer than three days of samples in a certain interval, samples from the same interval in adjacent time periods (fifteen days forward and fifteen days backward) are used to supplement the data. If this is still insufficient, the midpoint of the upper and lower boundaries of the interval is used as a substitute to ensure that each interval has a usable multi-level rainfall intensity reference value.

[0055] After establishing multi-level rainfall intensity reference values, electricity consumption comparisons under different rainfall intensities were calculated using a rainless condition as a unified baseline. The daily electricity consumption of the same user under rainless conditions was arithmetically averaged to obtain the baseline average electricity consumption under rainless conditions. The daily electricity consumption was then arithmetically averaged for each of the four intervals: light rain, moderate rain, heavy rain, and extreme rainfall, yielding the average electricity consumption for each interval. The relative difference between the average electricity consumption for each interval and the baseline average electricity consumption was calculated using the formula: "Interval average electricity consumption minus baseline average electricity consumption, then divided by baseline average electricity consumption." The four relative difference values ​​were then aggregated in interval order to form a multi-level electricity consumption difference rate set. To suppress bias caused by imbalanced sample sizes, a sample size weight was introduced for the relative difference values ​​of each interval, with intervals having more samples receiving higher weights. When an interval had only the fewest samples, its relative difference was smoothed first-order with the relative differences of adjacent intervals according to the sample ratio to avoid amplifying the results due to occasional extreme days.

[0056] In terms of time dimension, to characterize the stability and phased changes of sensitivity, the multi-level electricity consumption difference rate set is processed temporally using a sliding window. The window length is 30 days per calendar month, with a step size of one day, and a center-aligned window strategy is adopted to align the timestamps. Within each window, the mean and variance of the relative difference values ​​of the four intervals are calculated separately, and four windowed sensitivity curves are formed with time as the horizontal axis and the window mean as the vertical axis. Intervals with less than three days of samples within a window are not directly calculated, but are instead merged with adjacent windows until a threshold is met. During merging, the center position of the timestamp is kept unchanged to ensure that the windowed sensitivity curves are continuously usable. For abrupt changes that may occur near seasonal transitions, a three-point median filter is introduced to perform a one-time light smoothing of the window mean, which preserves the trend while suppressing isolated spikes.

[0057] To eliminate system bias caused by non-rainfall factors, a normalized offset correction is performed on the windowed sensitivity curve. Still using each window as a unit, both temperature and humidity auxiliary variables within that window are simultaneously read from the standardized dataset. The temperature and humidity deviations relative to the historical mean of that window are calculated separately. A univariate or bivariate linear correction coefficient is established between the windowed sensitivity curve value for that window and the two deviations. These coefficients are estimated offline during training using least squares and are fixed during application. The estimated coefficients are used to perform offset regression on the window values ​​to obtain a corrected value sequence after eliminating the combined influence of temperature and humidity. This correction is performed for each of the four intensity intervals, and the summation is the corrected sensitivity sequence. If temperature or humidity data is missing in a window, it is replaced by the mean of the two adjacent windows, and a missing data indicator is marked in that window. The missing data indicator does not participate in numerical calculations but only serves as an optional input in subsequent modeling to maintain data source transparency.

[0058] After obtaining the corrected sensitivity sequence, the corrected values ​​for each user at the same alignment time for the four intervals of light rain, moderate rain, heavy rain, and extreme rainfall are arranged in a fixed order to form a principal vector of length four. To simultaneously express response stability, a weighted average of the variances of the four interval corrected values ​​within the most recent three windows is calculated at the same alignment time. This scalar is used as a stability index and concatenated with the principal vector to form a five-dimensional sensitivity feature vector. The weights of the weighted average are determined according to the proportion of valid sample days within the window; windows with more sufficient samples have higher weights. When there are no valid samples in a certain interval of a window, the weight of that interval in that window is reset to zero and normalized according to the weights of the remaining intervals, ensuring that the stability index is always contributed by truly valid intervals. The components of the sensitivity feature vector for all users are normalized to zero mean and unit variance, retaining the normalization parameters for consistent processing of new samples during application.

[0059] In a detailed embodiment, consider the complete data processing flow for a user within a 30-day calendar month. The recorded rainfall in the user's region is as follows: 10 days without rain, 7 days with light rain, 8 days with moderate rain, 3 days with heavy rain, and 2 days with extreme torrential rain. When processing the standardized dataset, daily rainfall is first categorized into four levels—light rain, moderate rain, heavy rain, and extreme rainfall—based on the statistical percentiles of historical rainfall distribution. For example, by statistically analyzing rainfall samples from the past two years, daily rainfall of 0–10 mm is defined as light rain, 10–25 mm as moderate rain, 25–50 mm as heavy rain, and greater than 50 mm as extreme torrential rain, while 0 mm is considered a separate condition for no rain. Subsequently, the average rainfall level for each level is calculated; for example, the average rainfall for light rain is 8 mm, for moderate rain it is 17 mm, for heavy rain it is 32 mm, and for extreme torrential rain it is 65 mm, thus generating multi-level rainfall intensity reference values.

[0060] After obtaining the reference value for rainfall intensity, the electricity consumption data of users within the same time window was extracted and grouped statistically according to level. Statistical results show that the user's average daily electricity consumption is 120 kWh under rainless conditions, 108 kWh under light rain conditions, 102 kWh under moderate rain conditions, 95 kWh under heavy rain conditions, and 88 kWh under extreme rainstorm conditions. Using 120 kWh under rainless conditions as the baseline value, the electricity consumption difference rate for each level was calculated: (108−120) / 120=−0.10 for light rain, (102−120) / 120=−0.15 for moderate rain, (95−120) / 120=−0.21 for heavy rain, and (88−120) / 120=−0.27 for extreme rainstorms, thus obtaining a multi-level electricity consumption difference rate set containing four difference rates.

[0061] To further capture user response stability, the multi-level electricity consumption difference rate set is processed using time-series windowing. In this example, a 15-day sliding window is selected, with each day sliding forward one day to form a new window. Within each window, the mean and variance of the difference rates for light rain, moderate rain, heavy rain, and extreme rainstorms are calculated. For example, in the window from day 10 to 24, the mean difference rate for light rain is -0.11 with a variance of 0.02, the mean difference rate for moderate rain is -0.14 with a variance of 0.03, the mean difference rate for heavy rain is -0.20 with a variance of 0.04, and the mean difference rate for extreme rainstorms is -0.26 with a variance of 0.05, thus generating four windowed sensitivity curves that update as the window moves over time.

[0062] After obtaining the windowed sensitivity curve, it needs to be normalized and offset corrected by introducing temperature and humidity as auxiliary variables. Assuming the average temperature within a window is 2°C higher than the historical average and the average humidity is 5% higher, the windowed sensitivity curve is corrected using the temperature coefficient of 0.02 and the humidity coefficient of 0.01, estimated in advance during the training phase using the least squares method. For example, if the difference rate for light rain in this window is −0.11, the correction value is −0.11−(0.02×2)−(0.01×5)=−0.11−0.04−0.05=−0.20; similarly, the difference rate for moderate rain is adjusted from −0.14 to −0.23, for heavy rain from −0.20 to −0.29, and for extreme torrential rain from −0.26 to −0.35, thus obtaining the corrected sensitivity sequence. This processing can eliminate non-rainfall interference from temperature and humidity on user electricity consumption, ensuring that the indicators focus on the direct relationship between rainfall and electricity consumption.

[0063] After obtaining the corrected sensitivity sequence, it needs to be vectorized and combined. Using each window as an alignment unit, the values ​​are arranged in a fixed order of light rain, moderate rain, heavy rain, and extreme rainstorm, resulting in a vector containing four components. For example, the corrected values ​​for this user in a certain window are -0.20, -0.23, -0.29, and -0.35, respectively. Simultaneously, to characterize the stability of the user's response to rainfall, a stability index needs to be calculated. In this example, the variance-weighted average of the corrected values ​​for each rainfall level within the most recent three windows is used as the stability measure. If the result is 0.04, this value is appended to the end of the above vector. Finally, the user's sensitivity feature vector for this window is [-0.20, -0.23, -0.29, -0.35, 0.04].

[0064] This embodiment first generates reference values ​​by segmenting rainfall intensities at multiple levels, then obtains a set of electricity consumption difference rates through grouped statistics, further generates a time-dimensional sensitivity curve using a sliding window, then performs normalization offset correction using temperature and humidity data, and finally vectorizes the corrected sensitivity sequence and adds a stability index. The resulting sensitivity feature vector reflects not only the differences in user electricity consumption behavior under different rainfall intensities but also its consistency and stability over time, ensuring that subsequent steps of the invention can continue to be executed under clear and uniform input conditions.

[0065] Step S103: The sensitivity feature vector is fused with the seasonal features and crop type features in the environmental attribute data to obtain a comprehensive sensitivity feature set. The comprehensive sensitivity feature set is used to characterize the multidimensional dynamic response of user electricity consumption behavior to climate factors.

[0066] In step S103, the sensitivity feature vector obtained in the previous step needs to be fused with the seasonal and crop type features from the environmental attribute data to form a comprehensive sensitivity feature set. This fusion is not simply piecing together different data, but rather organically combining these features within a unified time window and user identifier, ensuring comparability and consistency. The sensitivity feature vector itself reflects the user's electricity consumption response under different rainfall conditions, while the seasonal features represent electricity consumption patterns within natural time cycles, and the crop type features reflect the degree to which agricultural users depend on climate conditions. Therefore, the fusion of these three types of information can more comprehensively depict user behavior patterns.

[0067] In practical implementation, time alignment is the first step. All data must be based on the same time scale, such as establishing a time window by day or week. Within this window, sensitivity feature vectors, seasonal features, and crop type features are all included as part of the same record. For example, if a user's sensitivity to light rain, moderate rain, and heavy rain within a month is -0.15, -0.30, and -0.45 respectively, then the user's feature record will include these three sensitivity values. Simultaneously, if the user is in summer, the indicator value for "summer" in the seasonal feature will be 1, while the indicator value for other seasons will be 0; if the user's declared crop type is rice, then the indicator value for "rice" in the crop feature will be 1, while other crops will be 0. In this way, the user's multidimensional information within that time window can be fully reflected in a single feature record.

[0068] During the fusion process, interaction effects also need to be considered. Individual sensitivity values ​​only describe the overall impact of rainfall on users, but in many cases, this impact is moderated by season and crop type. For example, rice may be more dependent on rainfall during its growing season, while the impact is less during the harvest season. To capture this difference, interaction features can be constructed by multiplying the sensitivity component by seasonal or crop type features to generate new columns. For example, the feature value of "Summer × Moderate Rain Sensitivity" is the product of the summer indicator and the moderate rain sensitivity. Similarly, the feature value of "Rice × Heavy Rain Sensitivity" can reflect the specific response of rice users under heavy rain conditions. Through these interaction features, the model can more accurately learn the moderating effects of season and crop on sensitivity.

[0069] When a user engages with multiple crops, crop features can be constructed based on the proportion of each crop. For example, if a user grows 70% rice and 30% corn, then within that time window, the rice indicator would be 0.7, the corn indicator 0.3, and the remaining crops 0. In this way, sensitivity values ​​are proportionally allocated when generating interaction features. For instance, if the sensitivity to moderate rain is -0.30, then "rice × moderate rain sensitivity" would be -0.21, and "corn × moderate rain sensitivity" would be -0.09. This weighted approach accurately reflects the user's composite response across different crop types.

[0070] After completing the feature fusion process, the resulting feature set needs to be standardized to maintain consistent numerical scales across different features and prevent biased results during model training due to excessive differences in feature ranges. Simultaneously, for extreme values ​​or occasional outliers, smoothing techniques such as moving averages or differencing can be used to ensure the feature set retains its trend without being overly affected by a single outlier. The final sensitivity feature set is represented as a high-dimensional matrix, where each row corresponds to a complete feature record of a user within a specific time window, and each column corresponds to a fused and standardized feature or interaction item.

[0071] Furthermore, the sensitivity feature vector is fused with the seasonal and crop type features in the environmental attribute data to obtain a comprehensive sensitivity feature set. This comprehensive sensitivity feature set is used to characterize the multidimensional dynamic response of user electricity consumption behavior to climate factors, including:

[0072] The sensitivity feature vector and the seasonal feature are time-series aligned, and the response values ​​of different rainfall levels in the sensitivity feature vector are mapped to the corresponding seasonal windows to generate a sensitivity distribution table indexed by season.

[0073] Crop type weights are introduced into the sensitivity distribution table. These crop type weights are proportionally allocated based on the user's main planting structure at different times. They are used to weight and correct the rainfall response values ​​under different crop types, generating a crop sensitivity correction matrix.

[0074] The crop sensitivity correction matrix is ​​interactively extended to construct a three-dimensional interactive feature tensor including rainfall level, season, and crop type, so that each feature unit can reflect the response characteristics of a specific season and a specific crop under specific rainfall conditions, thereby obtaining a complete interactive feature tensor.

[0075] The interaction feature tensor is normalized and compressed. The normalization coefficient is calculated based on the variability of each unit in the time dimension. Units with high information redundancy are dimensionality reduced, and units with high variability are retained and amplified to obtain the compressed core feature set.

[0076] The core feature set is concatenated with the stability index in the sensitivity feature vector, and a cross-seasonal smoothing coefficient is added to form the final comprehensive sensitivity feature set, which is used to comprehensively characterize the multi-dimensional dynamic response relationship of user electricity consumption behavior to rainfall, season and crop type.

[0077] When fusing sensitivity feature vectors with seasonal and crop type features, the first step is to ensure consistency across the time dimension. Temporal alignment refers to placing a user's sensitivity features at different times into corresponding time windows based on the season. For example, this can be done by dividing the data into spring, summer, autumn, and winter based on regional climate habits, or by using the power industry's commonly used wet season, normal season, and dry season. Each user's sensitivity feature vector at any given time contains values ​​for four levels of rainfall: light rain, moderate rain, heavy rain, and extreme rainfall. Through alignment, these sensitivity values ​​are categorized into their respective seasonal windows. This results in a sensitivity distribution table indexed by season, with each row corresponding to a season and each column corresponding to a rainfall level. The values ​​in the table represent the user's typical sensitivity performance within that season. For seasons lacking complete samples, data from adjacent seasons can be used to supplement the data, or the median value of the user's historical data can be used as a substitute, ensuring that every combination of season and rainfall level has available values.

[0078] After obtaining the seasonal distribution, crop type weights are introduced to correct these data. Crop type weights refer to the proportion of different crops in a user's overall electricity consumption within a season. For example, in summer, if rice accounts for 60% of total electricity consumption, corn for 30%, and other crops for 10%, then in the calculation of summer sensitivity, the weights related to rice would be set at 0.6, corn at 0.3, and other crops at 0.1. Through this allocation, seasonal sensitivity can be combined with crop structure, and the sensitivity values ​​under each rainfall level can be weighted and corrected to obtain a new matrix. The rows of the matrix represent rainfall levels, the columns represent crop types, and each cell value represents the sensitivity performance of a particular crop under a specific rainfall level. If it is necessary to further reflect the sensitivity of crops to rainfall at different growth stages, additional proportions of crop growth stages can be introduced into the weights, such as the heading stage being more important than the dormancy stage, thus making the corrected results more consistent with reality.

[0079] After forming the crop correction matrix, the factors of season, crop, and rainfall level are combined to construct a three-dimensional interaction feature tensor. This tensor can be understood as a cube, with the three directions representing rainfall level, season, and crop type, respectively. Each small square within the cube represents the electricity sensitivity of a particular crop under a specific rainfall condition in a particular season. The advantage of this three-dimensional structure is that it can simultaneously reveal the interaction relationships between different dimensions, rather than simply adding or splicing them together. For example, it can reveal differential patterns such as "summer rice is highly sensitive to heavy rain conditions, while winter wheat is less sensitive to light rain conditions."

[0080] Since 3D interaction tensors often contain a large amount of data, direct use may introduce redundancy, thus requiring compression and filtering. Compression is based on temporal variability. Specifically, this involves observing the variation of the same feature unit across different time windows. If the value of a feature unit remains almost unchanged throughout the observation period, its contribution to classification is small and it can be merged or weakened. Conversely, if a unit shows significant temporal variation, it contains stronger discriminative information and should be retained with higher weight in the results. Simultaneously, it's necessary to detect highly similar change patterns between different units. If two units are almost identical in time, the one with the more significant change can be retained, while the other can be discarded to avoid information redundancy. Through this processing, the originally complex 3D tensor is compressed into a core feature set, retaining the most informative and differentiated features.

[0081] In the final step, this core feature set needs to be combined with the stability index from the sensitivity feature vector. The stability index, calculated in the previous step, measures the stability of a user's response across different time windows. During fusion, it is directly appended to the end of the core feature set. Simultaneously, a cross-seasonal smoothing coefficient needs to be introduced. This coefficient measures whether a user's sensitivity changes abruptly at seasonal transitions. It is calculated by comparing the sensitivity difference between two adjacent seasons. If the average difference is small, it indicates that the user's behavior is stable at seasonal transitions, and the smoothing coefficient is close to 1; if the difference is large, it indicates that the user's behavior differs significantly across seasons, and the smoothing coefficient is small. This coefficient helps the subsequent model identify user behavioral characteristics during cross-seasonal transitions.

[0082] After completing these steps, the final comprehensive sensitivity feature set can be formed. This feature set not only includes the user's basic sensitivity under different rainfall intensities, but also incorporates weighted corrections based on crop structure, showcasing three-dimensional interaction relationships. Furthermore, it highlights the most informative parts after compression and filtering, and introduces stability indicators and cross-seasonal smoothing factors to ensure that the results are both comprehensive and discriminative.

[0083] For example, consider an agricultural user located in an area primarily growing rice and corn. Within the summer window of the observation year, based on the sensitivity feature vector calculated in the previous step, the user's average sensitivity values ​​for light rain, moderate rain, heavy rain, and extreme rainfall levels are -0.12, -0.18, -0.25, and -0.30, respectively. Here, the sensitivity value represents the relative rate of change; a negative value indicates that the greater the rainfall, the more significant the decrease in electricity consumption. Within the same summer window, the user's crop planting structure is 60% rice and 40% corn. Meanwhile, the sensitivity feature vectors for the autumn window are -0.05, -0.09, -0.15, and -0.20, with the main crop structure being 70% corn and 30% wheat.

[0084] First, time-series alignment is performed, placing the summer and autumn data into their respective seasonal windows to obtain a sensitivity distribution table indexed by season. In this table, the values ​​corresponding to the four rainfall levels for summer are -0.12, -0.18, -0.25, and -0.30, as mentioned above, while the values ​​for autumn are -0.05, -0.09, -0.15, and -0.20.

[0085] Based on this, crop type weights are introduced. In summer, rice has a weight of 0.6 and corn has a weight of 0.4. Therefore, the summer sensitivity distribution will be corrected to the results calculated separately for rice and corn. For example, the summer sensitivity under light rain conditions is -0.12, and the correction value for rice is -0.12 × 0.6 = -0.072, and the correction value for corn is -0.12 × 0.4 = -0.048. Similarly, the summer sensitivity under extreme rainfall conditions is -0.30, and after correction, it is -0.18 for rice and -0.12 for corn. The resulting summer crop sensitivity correction matrix is ​​as follows: rice is -0.072, -0.108, -0.150, and -0.180 under light rain, moderate rain, heavy rain, and extreme rainfall, respectively, and corn is -0.048, -0.072, -0.100, and -0.120 under the corresponding levels. The same calculation is performed for autumn, with a weight of 0.7 for maize and 0.3 for wheat. The correction values ​​under light rain conditions are −0.05×0.7=−0.035 (maize) and −0.05×0.3=−0.015 (wheat), while the correction values ​​under extreme rainfall conditions are −0.20×0.7=−0.14 (maize) and −0.20×0.3=−0.06 (wheat). This process is repeated to obtain the crop sensitivity correction matrix for autumn.

[0086] The above results were organized into a three-dimensional interaction feature tensor. The three dimensions of this tensor are rainfall level, season, and crop type, respectively. In the cell for light summer rain, the sensitivity is -0.072 for rice and -0.048 for maize; in the cell for extreme summer rainfall, it is -0.180 for rice and -0.120 for maize; in the cell for light autumn rain, it is -0.035 for maize and -0.015 for wheat; in the cell for extreme autumn rainfall, it is -0.14 for maize and -0.06 for wheat. The resulting cube structure fully represents the interaction relationship of "rainfall level × season × crop type".

[0087] After construction, this tensor needs to be compressed and filtered. Observing the changes of the same tensor unit for the user over multiple consecutive years reveals that the sensitivity value of summer rice under heavy rain conditions fluctuates significantly, for example, -0.16, -0.20, and -0.18 in the past three years, indicating significant temporal variability, and therefore it is retained as a key feature. In contrast, the sensitivity of autumn wheat under light rain conditions shows almost no change, remaining between -0.015 and -0.017 in the past three years, indicating little discriminative significance, and it can be weakened or merged. Simultaneously, it is necessary to check whether different units are highly similar. For example, the change curves of autumn corn under moderate rain conditions are very similar to those under heavy rain conditions. If the correlation exceeds a threshold, only one can be retained. After this series of operations, the size of the core feature set is reduced, but those units with greater temporal variability are retained.

[0088] Finally, this core feature set is concatenated with the existing stability index in the sensitivity feature vector. For example, within the summer window, a user's sensitivity stability index is 0.05, indicating that their sensitivity fluctuates somewhat within the time window. Simultaneously, a cross-seasonal smoothing coefficient is calculated to measure the sensitivity difference between summer and autumn. If the average sensitivity difference between summer and autumn under light rain conditions is 0.07, moderate rain is 0.09, heavy rain is 0.10, and extreme rainfall is 0.12, the overall difference level is not drastic, and the cross-seasonal smoothing coefficient can be taken as 0.8, indicating that the user's sensitivity changes relatively smoothly across seasons. Appending the stability index and the cross-seasonal smoothing coefficient to the end of the core feature set forms the final comprehensive sensitivity feature set.

[0089] Step S104: Input the sensitivity comprehensive feature set into the decision tree classification model constructed based on the Gini coefficient splitting criterion. In the training process, introduce a cross-seasonal sample balancing mechanism and a rainfall anomaly compensation mechanism to form a training model with environmental robustness. The training model is used to learn the nonlinear mapping relationship between the sensitivity comprehensive feature set and the user category.

[0090] In step S104, the sensitivity comprehensive feature set obtained in the previous step needs to be input into a decision tree classification model built based on the Gini coefficient splitting criterion, so that the model can learn the complex relationship between these features and user categories. The core of the decision tree is to divide the input data into several subsets according to different value ranges of features by continuously splitting nodes. Each split follows the principle of minimizing the Gini coefficient, so that the purity of classification is gradually improved. The Gini coefficient is an indicator that measures the impurity of data, with a value between zero and one. When almost all the samples in a node belong to the same category, the Gini coefficient of that node is close to zero; when the categories are severely mixed, the Gini coefficient is close to a higher value. During training, the model will examine each feature in the sensitivity comprehensive feature set in turn, including sensitivity indicators, seasonal features, crop type features, and their interaction terms. By comparing the Gini coefficients of different features under different splitting thresholds, the model selects the feature that can minimize impurity to the greatest extent as the splitting condition, thereby gradually establishing a decision path.

[0091] To avoid over-reliance on a specific season or type of sample during training, which could lead to a decline in model generalization ability, this step introduces a cross-seasonal sample balancing mechanism. The basic idea of ​​this mechanism is to balance the data samples from different seasons, ensuring that the model does not become biased towards summer classification patterns simply because there are far more samples in summer than in winter. This can be achieved by using oversampling techniques in seasons with fewer samples (i.e., expanding the sample size by copying or synthesizing data) or undersampling techniques in seasons with more samples (i.e., randomly reducing some samples) to ensure a balanced proportion of different seasons in the training set. For example, if a region has 10,000 user samples in summer and only 3,000 in winter, to prevent the model from becoming overly reliant on summer patterns, the summer samples can be reduced to 4,000, or the winter samples expanded to 6,000, ultimately allowing the model to learn relatively reliable classification patterns across both summer and winter.

[0092] Meanwhile, a compensation mechanism needs to be designed to address abnormal rainfall conditions. Abnormal rainfall refers to extreme weather events occurring within certain time windows, such as continuous heavy rain or prolonged drought, causing users' electricity consumption behavior to deviate significantly from normal patterns. Directly using this abnormal data for training may lead to the model misjudging normal conditions during prediction. Therefore, abnormal data should be identified and corrected during training. Identification methods can be based on statistical thresholds; for example, a day's rainfall exceeding three standard deviations of the historical average can be marked as an abnormal rainfall day. For these samples, compensation strategies can be employed, such as generating an adjustment value through interpolation or regression based on adjacent normal day data, restoring electricity consumption behavior from abnormal fluctuations to a range closer to normal levels before participating in training. Alternatively, a specific identifying feature can be added to the abnormal data, enabling the model to distinguish abnormal samples from normal samples when splitting nodes, rather than simply mixing them together. This compensation mechanism can effectively prevent extreme weather conditions from damaging the overall model training results.

[0093] By integrating cross-seasonal sample balancing and rainfall anomaly compensation mechanisms, the decision tree model gradually establishes a more robust classification structure. The features selected for each split node not only reflect users' electricity consumption sensitivity under daily climate conditions but also take into account seasonal and crop type differences. Simultaneously, the model avoids misleading data skew through compensation and balancing. This training method ensures that the model maintains high classification accuracy under various environmental conditions without losing reliability due to imbalanced training samples or the presence of extreme values.

[0094] After training, the resulting decision tree model faithfully reflects the nonlinear mapping between the comprehensive sensitivity feature set and user categories. In other words, the model learns to distinguish user categories based on different users' responses to rainfall, electricity consumption patterns in different seasons, and the demand characteristics of different crops. Because environmental uncertainties and anomalies were fully considered during training, the model exhibits strong robustness and generalization ability when facing real-world data.

[0095] In step S104, the sensitivity comprehensive feature set is input into the decision tree classification model constructed based on the Gini coefficient splitting criterion. The model's parameters are adjusted and its structure optimized through the training process. A cross-seasonal sample balancing mechanism and a rainfall anomaly compensation mechanism are introduced during training to enhance the environmental robustness of the resulting decision tree classification model. The trained model is the "trained model" referred to in this application. This trained model is essentially the completed decision tree classification model, capable of accurately learning the nonlinear mapping relationship between the sensitivity comprehensive feature set and user categories.

[0096] Furthermore, the sensitivity comprehensive feature set is input into a decision tree classification model constructed based on the Gini coefficient splitting criterion. During training, a cross-seasonal sample balancing mechanism and a rainfall anomaly compensation mechanism are introduced to form an environmentally robust training model. This training model is used to learn the nonlinear mapping relationship between the sensitivity comprehensive feature set and user categories, including:

[0097] Cross-seasonal stratified sampling is performed on the sensitivity comprehensive feature set, and data in each season are extracted in the same proportion to generate a cross-seasonal balanced training set with a balanced number of samples, in order to avoid the bias caused to the training model by the excessive proportion of samples in a certain season.

[0098] Rainfall anomaly detection is performed on the cross-seasonal balanced training set. A reference interval is established using historical rainfall distribution. Extreme rainfall samples that exceed the interval threshold are individually labeled and their impact on split node calculation is reduced through a weight decay strategy to generate an anomaly-compensated cross-seasonal balanced training set.

[0099] The cross-seasonal balanced training set after anomaly compensation is input into the decision tree construction module. The degree of classification purity reduction is calculated on each candidate feature dimension according to the Gini coefficient splitting criterion. The feature that can reduce impurity to the greatest extent is selected as the current splitting condition to generate the initial split tree structure.

[0100] The cross-seasonal robustness of the split tree structure is evaluated by comparing the classification purity differences of sub-samples from different seasons at each split node. When the difference exceeds a set threshold, the influence of seasonal differences is reduced by adjusting the sample weights or introducing additional regularization factors, thereby obtaining a cross-seasonal robust split tree structure.

[0101] The cross-seasonal robust split tree structure is subjected to multiple rounds of pruning operations. The error rate is minimized based on the validation set. Leaf nodes and branches that do not contribute enough information are deleted while retaining the environmental sensitivity features, resulting in a simplified cross-seasonal robust split tree structure.

[0102] The simplified cross-seasonal robust split tree structure is jointly corrected with the stability index of the sensitivity integrated feature set. When the output category prediction is performed at each leaf node, the stability index is introduced as a correction factor. The confidence of feature combinations with large fluctuations is reduced to obtain the final training model, which is used to maintain stable classification performance in the face of cross-seasonal imbalance and rainfall anomalies.

[0103] In this section, the sensitivity-integrated feature set needs to be trained on a decision tree classification model based on the Gini coefficient splitting criterion. To ensure that the training results not only learn the nonlinear mapping relationship between user categories, but also remain robust under uneven seasonal distribution and extreme rainfall conditions, a cross-seasonal sample balancing mechanism and a rainfall anomaly compensation mechanism must be introduced.

[0104] First, cross-seasonal stratified sampling is required for the sensitivity feature set. Cross-seasonal stratified sampling involves dividing all samples into several tiers according to their respective seasons, such as spring, summer, autumn, and winter. Within each season, samples are drawn from the dataset in a uniform proportion, ensuring that the proportion of each season in the training set remains consistent. This approach prevents an excessively large sample size for any one season, which could cause the decision tree to prioritize learning the features of that season during the splitting process, neglecting the performance of other seasons. For example, if summer samples account for 60% of the original data while winter samples only account for 10%, the summer samples will be reduced and the winter samples expanded during sampling, making their proportions in the training set closer, for example, 25% each. This cross-seasonal balanced training set provides a balanced input foundation for subsequent models.

[0105] After obtaining the cross-seasonal balanced training set, rainfall anomaly detection is required. Rainfall anomaly detection is defined as comparing historical rainfall distributions to determine whether certain rainfall levels constitute extreme cases. Specifically, this involves statistically analyzing rainfall data over several years to identify the typical rainfall distribution range. For example, if 90% of daily rainfall in a region falls between 0 and 50 millimeters, then daily rainfall exceeding 100 millimeters can be considered extreme. When a sample in the training set corresponds to a rainfall level outside this reference range, it needs to be separately marked as an anomalous sample. These anomalous samples are not directly deleted; instead, their training weights are reduced to control their impact on the decision tree splitting results. Weight decay is achieved by reducing the sample's proportion in the Gini coefficient calculation, for example, retaining only half or one-third of the weight of normal samples. This process generates an anomaly-compensated cross-seasonal balanced training set that includes extreme rainfall scenarios without allowing these few extreme samples to distort the overall model's training direction.

[0106] After obtaining the cross-seasonal balanced training set with anomaly compensation, the system proceeds to the decision tree construction module. Here, each candidate feature is examined sequentially, and the improvement in data purity is calculated under different splitting thresholds. The Gini coefficient splitting criterion determines which feature is most suitable as the splitting condition for the current node by comparing the changes in the degree of class mixing of samples before and after the split. The lower the Gini coefficient, the purer the child nodes after the split. When comparing multiple features, the one that causes the largest decrease in the Gini coefficient is selected as the splitting criterion. This process is repeated recursively layer by layer, gradually generating the initial split tree structure. At this stage, although the split tree can fit the data, it has not yet undergone cross-seasonal robustness adjustment.

[0107] Subsequently, the initial split tree structure needs to be evaluated for its robustness across seasons. At each split node, it's necessary to assess not only whether the split improves overall purity but also to compare the classification performance of subsamples from different seasons. If the classification results of the same node differ significantly between summer and winter—for example, a substantial increase in purity in summer but almost no improvement in winter—it indicates a bias in the applicability of the split condition to different seasons. To address this issue, sample weights can be dynamically adjusted during training, increasing the proportion of samples from weaker seasons, or additional regularization factors can be introduced to penalize split conditions that overly rely on a particular season. After such modifications, a robust split tree structure across seasons can be obtained, ensuring that the decision path possesses relatively consistent discriminative ability across different seasons.

[0108] After forming the cross-seasonal robust split tree structure, multiple rounds of pruning are required. The core of pruning is to test on the validation set and remove nodes and branches that contribute little to the classification results or may even cause overfitting. When performing pruning, it's not simply about reducing the tree size, but rather about carefully removing redundant parts while ensuring the model retains its environmentally sensitive features. For example, if a branch only slightly improves accuracy but introduces a large number of noise-sensitive features, it will be pruned. After multiple iterations, a streamlined cross-seasonal robust split tree structure is finally obtained, which is neither redundant nor fails to retain truly critical feature paths.

[0109] Finally, the stability index of the simplified cross-seasonal robust split tree structure and the sensitivity-integrated feature set needs to be jointly corrected. The stability index, defined in previous steps, measures the degree of fluctuation in user sensitivity across different time windows. This stability index can be used as a correction factor when predicting the output category at the leaf nodes. If the stability index for a feature combination shows high volatility, the confidence of the prediction result is reduced; if the stability index for a feature combination shows high stability, its confidence is maintained or increased. In this way, the training model not only relies on the splitting path of the decision tree but also utilizes stability information to correct the final output, thereby enhancing the model's reliability in real-world complex environments.

[0110] After this complete process, the final trained model can still maintain stable and high-precision classification performance when faced with problems such as cross-seasonal sample imbalance and abnormal rainfall.

[0111] Here is a specific example:

[0112] In a sample database containing nearly two years of data, the data was statistically analyzed by season, with 2,400 records for spring, 6,000 for summer, 1,800 for autumn, and 800 for winter, categorized as "agricultural irrigation type," "residential type," and "industrial and commercial type." To obtain a cross-seasonal balanced training set, the goal was to select 2,000 records from each season, totaling 8,000 for training: 2,000 records were randomly downsampled from the 6,000 for summer; 2,000 from the 2,400 for spring; 2,000 from the 1,800 for autumn through bootstrapping; and 2,000 from the 800 for winter through bootstrapping and perturbation with samples from adjacent weeks. This resulted in the "cross-seasonal balanced training set."

[0113] Rainfall anomaly detection was performed on these 8,000 samples. Historical statistics showed that 95% of daily rainfall in this region did not exceed 65 mm, and rainfall exceeding 80 mm was considered extreme rainfall; meanwhile, 10 or more consecutive days with daily rainfall not exceeding 1 mm were considered extreme drought. The training set was found to contain 320 samples in extreme rainfall and 190 samples in extreme drought. These two types of samples were labeled separately and subjected to "rainfall anomaly compensation": the weight of extreme rainfall samples in the training was reduced to half that of normal samples; the weight of extreme drought samples was reduced to 70% of that of normal samples. The weight of the remaining samples remained at 1. After this processing, a "cross-seasonal balanced training set after anomaly compensation" was obtained, and all subsequent splits and evaluations were counted according to these weights.

[0114] After entering the decision tree construction module, the "improvement of classification purity" is calculated for each candidate feature under different splitting thresholds. The candidate features at the root node include: heavy rain sensitivity, extreme rainfall sensitivity, seasonal coding (summer), crop type (rice) weight, interaction feature (crop type (rice) × heavy rain sensitivity), and cross-seasonal smoothing coefficient. Comparison revealed that the "interaction feature (crop type (rice) × heavy rain sensitivity)" provides the greatest purity improvement at a threshold around -0.18: after splitting at this threshold, the left subset is dominated by "agricultural irrigation type" (1,560 agricultural irrigation type entries and 240 other types by weight), while the right subset is dominated by "residential / industrial and commercial type" (3,120 residential / industrial and commercial type entries and 1,120 agricultural irrigation type entries by weight), significantly reducing the heterogeneity compared to not splitting. Therefore, the root node uses this feature and this threshold to generate the first layer of the initial split tree structure.

[0115] A cross-seasonal robustness assessment was conducted on the initial split tree structure. Taking the two child nodes after the root node split as an example, the "classification purity" of each season within the child nodes was statistically analyzed. The seasonal purity of the left child node (dominated by agricultural irrigation) was: summer 0.89, spring 0.83, autumn 0.81, and winter 0.72; the difference between the maximum and minimum was 0.17, which is higher than the preset threshold of 0.10, indicating that the split is not robust enough to winter. To reduce the impact of seasonal differences, the weights of samples from different seasons within this node were dynamically adjusted: the weight of winter samples was multiplied by 1.3, the weight of summer samples was multiplied by 0.9, and the weights of spring and autumn samples remained unchanged; at the same time, a narrow-range regularization was applied to the threshold range of "seasonal coding_winter × heavy rain sensitivity" to make further splits in the winter direction more likely to occur. After applying the above adjustments, the seasonal purity of the same node was reassessed: summer 0.85, spring 0.83, autumn 0.82, and winter 0.82, and the difference between seasons decreased to 0.03, meeting the threshold requirement. The robustness assessment and weight fine-tuning of the remaining nodes in the tree are performed in the above manner to obtain the "cross-seasonal robust split tree structure".

[0116] After growing to a maximum depth of 6 layers, a larger tree was obtained. To avoid overfitting, a separate validation set of 2,000 samples (balanced sampling of 500 samples each season) was prepared, and the "cross-seasonal robust split tree structure" was pruned multiple times. The initial validation error without pruning was 12.6%. Pruning a leaf pair at layer 5 subdivided by "extreme rainfall sensitivity" (which only covered 180 samples in the training set and was predominantly extreme samples) reduced the validation error from 12.6% to 11.9%. Pruning another small branch subdivided by "cross-seasonal smoothing coefficient" (covering 140 samples in training with limited validation improvement) further reduced the error to 11.6%. When a third, more widely covered branch was pruned, the validation error rose to 12.0%, so this pruning was reversed. Finally, a "simplified cross-seasonal robust split tree structure" with a depth of 4 layers was obtained, which performed best on the validation set.

[0117] The simplified cross-seasonal robust split tree structure was jointly calibrated, and a stability index was introduced to adjust the prediction confidence of leaf nodes. The stability index, derived from the aforementioned sensitivity comprehensive feature set, measures the fluctuation of the same user's response to rainfall within adjacent time windows. A simple and operable rule was established: when the stability index does not exceed 0.05, the tree is considered stable, and the original confidence of the leaf node can be slightly increased by 5%; when the stability index is between 0.05 and 0.10, no adjustment is made; when the stability index exceeds 0.10, the tree is considered to have large fluctuations, and the confidence of the leaf node is decreased by 20%. Two specific samples illustrate this: Sample A falls on a leaf node, with an original prediction of "agricultural irrigation type" with a confidence level of 0.86 and a stability index of 0.12, indicating high volatility. After a 20% reduction in confidence level, the score is 0.69, still higher than the system's decision threshold of 0.60, therefore it is classified as "agricultural irrigation type." Sample B falls on another leaf node, with an original prediction of "residential type" with a confidence level of 0.76 and a stability index of 0.03, indicating stability. After a 5% increase in confidence level, the score is 0.80, and it is classified as "residential type." This joint correction directly affects the leaf node output, neither altering the tree structure nor hindering the risk of unstable feature combinations at the output end.

[0118] To test the environmental robustness of the final trained model, an independent test set of 2,000 samples was prepared, with 500 samples sampled from each of the four seasons, while retaining the actual distribution of extreme rainfall and extreme drought. The test results showed an overall accuracy of 88.4%, with 88.0% in spring, 89.1% in summer, 90.0% in autumn, and 86.2% in winter. The accuracy was 84.7% on the subset containing extreme rainfall and 85.5% on the subset containing extreme drought, not significantly different from the 89.6% accuracy of the non-extreme samples. This indicates that the "cross-seasonal sample balancing mechanism" and the "rainfall anomaly compensation mechanism" effectively improved the model's cross-seasonal consistency and tolerance to abnormal climates.

[0119] Step S105: Use the trained model to classify and identify the new multi-source raw data. Without manually setting a fixed threshold, the model achieves dynamic and high-precision user category identification based on the user's sensitivity to changes in rainfall and related environmental factors.

[0120] In step S105, the trained decision tree classification model obtained in the previous step is used to classify and identify the new multi-source raw data. The new multi-source raw data is consistent with the data source used in the previous training, including user electricity consumption data, rainfall data, and user-related environmental attribute data. To ensure the comparability and consistency of the input data, this new data first needs to undergo the same standardization process as in step S101 to ensure that all variables are under uniform dimension and scale constraints. The standardized data is then used to generate new sensitivity feature vectors according to steps S102 and S103, and fused with seasonal features and crop type features to form a new comprehensive sensitivity feature set. This ensures that the data features are consistent between the training and application phases.

[0121] When a new sensitivity feature set is input into the trained decision tree classification model, the model will make judgments layer by layer according to the established splitting rules. Each sample starts from the root node of the tree and moves down the path indicated by the splitting condition based on its sensitivity index, seasonality features, and crop type features, until it falls into a leaf node. In the leaf node, the model has recorded the main class distribution of that type of sample during training, so the new sample can be automatically assigned to the corresponding user category. Unlike traditional classification methods that rely on fixed thresholds, the threshold of each splitting node here is optimized through the Gini coefficient during training, thus dynamically adapting to the feature performance of different users under different environmental conditions.

[0122] In practical implementation, the model can distinguish between agricultural users who are highly sensitive to rainfall and residential or commercial users who are not sensitive to rainfall changes. For example, if a user's electricity consumption significantly decreases during periods of light to moderate rain compared to rainless periods, and this user is identified as a rice-growing user based on crop type characteristics, and is in the summer high-water-consumption period, the model will comprehensively classify them as an agricultural irrigation user. On the other hand, if a user's sensitivity index under different rainfall conditions is close to zero, and neither their seasonal characteristics nor crop type characteristics show obvious agricultural attributes, the model will classify this user as a general residential user. This classification no longer relies on a single, manually set threshold, but is accomplished through multi-dimensional rules automatically learned by the training model, thus achieving more refined and dynamic discrimination.

[0123] In application, this classification can be performed on a daily, weekly, or monthly basis, depending on the time granularity of the input data. Those skilled in the art can adjust the data time window according to the power company's business needs and repeatedly perform the standardization, sensitivity calculation, and feature fusion processes before inputting the generated feature set into the trained decision tree classification model. The final classification result can serve as the basis for the power company's user management, differentiated electricity pricing design, and grid dispatching decisions. In this way, this step not only ensures the automation and efficiency of the classification process but also significantly improves the accuracy and robustness of the classification results, ensuring that the model can operate stably under different environmental and climatic conditions and produce reliable classification conclusions.

[0124] Furthermore, the step of using the trained model to classify and identify new multi-source raw data, without the need for manually setting fixed thresholds, achieves dynamic and high-precision user category discrimination based on the user's sensitivity to changes in rainfall and related environmental factors, including:

[0125] The new multi-source raw data is standardized and mapped in real time. The same dimensional and scale constraints as those in the training phase are adopted. User electricity consumption, rainfall and environmental attribute data are synchronously calibrated to generate a real-time standardized input set, ensuring that the input data and the feature space of the training model remain completely consistent.

[0126] The real-time standardized input set is compared with the user's historical sensitivity features over time. The sensitivity offset of the new data under the same season and crop type conditions is extracted, and the real-time input is corrected using the offset so that it can dynamically reflect the difference between the new data and the historical pattern, thus obtaining the environmentally corrected input set.

[0127] The environmentally corrected input set is input into the cross-seasonal robust splitting structure of the training model. Cross-seasonal discriminant weights are introduced into each splitting node to dynamically reconcile the splitting results of different seasons and output a preliminary classification path to ensure that the prediction results remain consistent in different seasons.

[0128] The stability of the leaf node output of the preliminary classification path is tested. A stability index is introduced from the sensitivity comprehensive feature set. The confidence of sample categories with large sensitivity fluctuations is reduced, and the confidence of categories with stable sensitivity is increased, generating a prediction result after stability correction.

[0129] The stability-corrected prediction results are calculated together with the cross-seasonal smoothing coefficient, which is used to quantify the continuity of user response at seasonal transitions and smooth out prediction results with excessive differences between seasons, to obtain the final classification result. This allows for dynamic and high-precision user category discrimination without the need for manually setting fixed thresholds.

[0130] When using a trained model to classify and identify new multi-source raw data, it is essential to ensure that the input data maintains a strictly consistent scale and dimension with the data from the training phase. Real-time standardization mapping involves recalibrating the collected user electricity consumption data, rainfall data, and environmental attribute data according to the mean and variance established during training, transforming all data into a uniform numerical range. For example, if electricity consumption was specified in kilowatt-hours and processed with zero mean and unit variance during training, then new electricity consumption data must undergo the same processing before being input into the model. The output of this process is the real-time standardized input set, which ensures that the new data is completely consistent with the feature space used by the training model, preventing model bias due to input distribution mismatch.

[0131] After generating the real-time standardized input set, it needs to be compared with the user's historical sensitivity features over time to extract the so-called sensitivity offset. The sensitivity offset is defined as the difference between the new input data and the user's typical sensitivity under the same season and crop type in the past. For example, if a user's historical sensitivity to heavy rain under summer rice growing conditions is -0.20, while the newly collected data shows a sensitivity of -0.28 under the same conditions, the offset is -0.08. The calculation method for the offset is intuitive: subtract the historical value from the new value. This calculation quantifies the difference between the new data and historical patterns. Subsequently, these offsets are used to correct the real-time input set, allowing the input data to dynamically reflect the trend of user sensitivity changes over time. The resulting corrected input set is more closely related to the user's actual state in the current environment than the raw standardized input.

[0132] When training the model with the environment-corrected input set, a cross-seasonal robust splitting structure built during training is required. In this structure, each splitting node not only relies on sensitivity-integrated features but also introduces cross-seasonal discriminative weights. The role of these weights is to dynamically reconcile the splitting results at that node across different seasons, preventing the model from over-relying on features from a single season. For example, if a node exhibits high classification purity in summer but is almost ineffective in winter, introducing weights can reduce the dominance of summer data at that node and increase the proportion of winter data, thus making the output classification path more universal. Ultimately, this process generates a preliminary classification path, ensuring that the model's discrimination remains consistent across different seasons.

[0133] After obtaining the initial classification path, it is necessary to perform a stability test on the output of the leaf nodes. Here, a stability index from the sensitivity feature set is introduced. The stability index is a quantitative value that measures the degree of fluctuation in a user's sensitivity value across multiple adjacent time windows. If the sensitivity value hardly changes between adjacent time windows, the stability index is low, indicating stable user behavior; conversely, if the fluctuation is significant, the index is high. The method for correcting the leaf node output using this index is to decrease the confidence of the prediction for that category when the stability index is high, and increase the confidence when the stability index is low. For example, if a leaf node gives the classification result of "agricultural user" with a confidence of 0.85, but the corresponding stability index shows that the user has fluctuated significantly recently, the confidence might be lowered to 0.70, thereby reducing the risk of the model making incorrect judgments due to random fluctuations.

[0134] After stability correction, the cross-seasonal smoothing coefficient needs to be calculated in conjunction with the data. The cross-seasonal smoothing coefficient quantifies the continuity of a user's response at seasonal transitions; it indicates whether a user's sensitivity changes drastically between adjacent seasons. If a user's sensitivity in summer is very similar to that in autumn, the cross-seasonal continuity is good, and the smoothing coefficient is close to 1; if the difference is large, the smoothing coefficient is low. This coefficient is used to correct the prediction results, reducing confidence or adjusting class boundaries in cases of excessive seasonal differences, ensuring that the final judgment is not distorted by sudden seasonal changes.

[0135] For example, in a real-world scenario, a power company performs online classification of a user's latest week's data during the summer. Since the units and scale parameters were fixed during the training phase, new data undergoes a standardization transformation before entering the model, based on the mean and variance from the training phase: daily electricity consumption is measured in kilowatt-hours and mapped to the training standard scale; rainfall is measured in millimeters and mapped to the training standard scale; temperature, humidity, crop type, and seasonal coding are also standardized using the parameters recorded during training. After this mapping, a real-time standardized input set is obtained. For example, if the user's daily electricity consumption is about one and a half standard units higher than the training average, rainfall is about one standard unit higher than the training average, and temperature and humidity are slightly higher than the training average but less than one standard unit, these values ​​are represented using the training scale, without retaining the original physical units, ensuring complete consistency with the feature space of the training model.

[0136] To ensure the input reflects its distance from historical patterns, a time-series comparison is performed between the real-time standardized input set and the user's historical sensitivity features. Historical records show that during summer, when the crop is rice, the user's typical sensitivity for light rain, moderate rain, heavy rain, and extreme rainfall levels is approximately -0.10, -0.18, -0.24, and -0.30, respectively (all are dimensionless values ​​of relative rate of change; the negative sign indicates a decrease in electricity consumption due to increased rainfall). In the latest weekly window, the new sensitivities obtained by comparing with a rainless baseline are approximately -0.14, -0.25, -0.33, and -0.39, respectively. Comparing these new values ​​with historical typical values ​​yields the sensitivity offset, with offsets of approximately -0.04, -0.07, -0.09, and -0.09 for the four levels. The offset indicates that, under the same seasonal and crop conditions, the user's response to rainfall this week is more sensitive than historically. Using these offsets, the relevant features in the real-time standardized input set are quantitatively corrected: the sensitivity components corresponding to the four rainfall levels are adjusted to be more sensitive, while the components of "crop is rice × heavy rain sensitivity" and "crop is rice × extreme rainfall sensitivity" are amplified in the interactive features, enabling the input to carry information that is "more sensitive than historical data". The corrected data is the environmentally corrected input set.

[0137] The environmentally corrected input set is fed into the cross-seasonal robust splitting structure of the training model. This structure dynamically reconciles the splitting results from different seasons using cross-seasonal discriminative weights at each splitting node. Since the current sample is from summer, the system also considers the historical performance of winter and autumn at the same nodes, slightly downweighting summer splitting evidence and slightly upweighting winter and autumn splitting evidence, preventing the discrimination from overly relying on highly sensitive summer samples. At the root node, the interaction feature "crop is rice × heavy rain sensitivity" is triggered as the primary splitting criterion, guiding the sample to the agriculture-related path. In the second layer, after the cross-seasonal discriminative weights reconcile the effectiveness of the splitting threshold, the sample continues to be guided to the agricultural irrigation cluster branch, forming a preliminary classification path. At this point, the original class distribution given by the leaf nodes shows that the confidence level for classifying the sample as agricultural irrigation type is approximately 0.78, while the combined confidence level for residential and industrial / commercial types is approximately 0.22.

[0138] To avoid overconfidence due to short-term fluctuations, the stability of the leaf node outputs of the initial classification path needs to be checked. The stability index is derived from the sensitivity comprehensive feature set and is defined as a quantitative value of the degree of fluctuation in sensitivity values ​​within adjacent time windows. The more windows and the smaller the numerical changes, the lower the stability index. Calculating the stability index for this user's most recent three-week windows yields a value of approximately 0.11, indicating a certain degree of recent fluctuation. According to the business rules determined in the training and validation phases, when the stability index exceeds 0.10, the confidence level of the leaf node outputs is lowered; the confidence level for agricultural irrigation is lowered from 0.78 to 0.62, while the confidence levels for other categories are increased proportionally, resulting in a stability-corrected prediction result.

[0139] After stability correction, a cross-seasonal smoothing coefficient is introduced and calculated in conjunction with the current forecast results. The cross-seasonal smoothing coefficient quantifies the continuity of a user's response at seasonal transitions. Its calculation involves comparing the average sensitivity differences between adjacent seasons: if the average difference between summer and autumn at various rainfall levels is generally small, the coefficient is close to 1; if the difference is generally large, the coefficient is low. In this example, the difference between light and moderate rain conditions in summer and autumn is small, while the difference between heavy and extreme rainfall conditions is moderate. A comprehensive evaluation yields a cross-seasonal smoothing coefficient of approximately 0.85, indicating that the user's behavior at seasonal transitions is generally smooth. During the joint calculation, when the smoothing coefficient is high and the current forecast is consistent with historical classifications of adjacent seasons, a slight increase in confidence level is allowed. Therefore, the confidence level for agricultural irrigation is increased from 0.62 to 0.66, while other categories are fine-tuned accordingly to obtain the final classification result.

[0140] The example above illustrates the complete process from real-time standardized mapping, to extracting sensitivity offsets and generating an environment-corrected input set, to deriving a preliminary classification path with cross-seasonal discriminative weights within a cross-seasonal robust splitting structure, to obtaining stability-corrected prediction results through stability indices, and finally, combining cross-seasonal smoothing coefficients to output the final classification result.

[0141] Furthermore, the real-time standardized mapping is achieved by invoking the mean and variance parameters from the training phase.

[0142] Furthermore, the sensitivity offset is the difference between the sensitivity value of the new input data and the historical sensitivity value.

[0143] Furthermore, the cross-seasonal discrimination weight is used to proportionally harmonize the classification results of samples from different seasons at the split node.

[0144] Furthermore, the cross-seasonal smoothing coefficient is determined by comparing the differences in the mean sensitivity values ​​of adjacent seasons.

[0145] In practical implementation, to achieve real-time standardized mapping, it is necessary to ensure that the newly input data remains completely consistent with the feature space used by the training model. This invention achieves this by calling the mean and variance parameters saved during the training phase. During the training phase, each feature undergoes statistical calculation to obtain its historical mean and variance, such as the average and standard deviation of daily electricity consumption, the average level and fluctuation range of daily rainfall, and the mean and standard deviation of ambient temperature and humidity. These parameters are recorded as benchmarks. When new data enters the system, the mean and variance are not recalculated; instead, the values ​​stored during the training phase are directly called, and these benchmark values ​​are used to transform the new data, mapping it to the exact same scale as the training data. The result of this process is that even if the distribution of the new data itself differs slightly from the distribution during training, after the same standardization transformation, the input still falls within a consistent feature space, ensuring that the training model can correctly understand and process this data. For example, if the average daily electricity consumption of users during the training phase is 100 kWh and the standard deviation is 20 kWh, while the electricity consumption on a certain day in the new data is 140 kWh, then after standardization, the conversion result is (140-100) / 20=2, which means it falls on the uniform scale of the training space.

[0146] After completing the real-time standardized mapping, the data needs to be corrected using a sensitivity offset to more accurately reflect the changes in the new data relative to historical patterns. The sensitivity offset is defined as the difference between the sensitivity value of the new input data and the historical sensitivity value. The calculation is straightforward: a simple subtraction. For example, historically, a user's sensitivity value under moderate rain conditions was -0.15, indicating that electricity consumption decreased by 15% under moderate rain conditions compared to no rain. However, in the new input data, the calculated sensitivity value is -0.20, indicating that the user's electricity consumption decreased by 20% under moderate rain conditions. The difference between these two values ​​results in a sensitivity offset of -0.05, indicating that the user is currently more sensitive than historically. In subsequent steps, the system uses this sensitivity offset to correct the input data, enabling the model to not only read real-time standardized features but also identify the differences between the user's current behavior and their historical behavior. This approach avoids the shortcomings of relying solely on historical averages and ignoring the dynamic nature of user behavior.

[0147] When the corrected input set is fed into the decision tree structure of the training model, cross-seasonal discriminative weights are introduced to ensure consistency in classification results across different seasons. Cross-seasonal discriminative weights are a mechanism at split nodes to proportionally adjust the classification results of samples from different seasons. Their purpose is to prevent data from one season from dominating the split and masking the features of other seasons. For example, at a certain node, if summer data significantly improves purity under that split while winter data shows almost no improvement, a traditional decision tree would prioritize summer features as the splitting criterion, leading to misclassification of winter data. By introducing cross-seasonal discriminative weights, the weight of winter data at that node can be artificially increased while moderately reducing the influence of summer data. Thus, during the selection of splitting conditions, the system comprehensively considers the performance of all seasons, avoiding overfitting of the classifier to any one season, thereby obtaining a more robust classification path throughout the year. These weights are typically set during training based on the sample size and performance differences of different seasons and dynamically invoked during actual runtime to reconcile splitting decisions.

[0148] After the model outputs its predictions, a final correction is needed using a cross-seasonal smoothing coefficient. The core of the cross-seasonal smoothing coefficient lies in quantifying the continuity of user sensitivity between adjacent seasons. Specifically, it compares the mean sensitivity values ​​of adjacent seasons. If the difference is small, it indicates stable user behavior during seasonal transitions, and the smoothing coefficient is close to 1, maintaining the original confidence level of the prediction. If the difference is large, it indicates significant fluctuations in user behavior during seasonal transitions, resulting in a lower smoothing coefficient and reduced confidence in the prediction. For example, if the mean sensitivity value for users is -0.25 in summer and -0.28 in autumn, the difference is only 0.03, indicating good continuity of behavior across seasons, and the smoothing coefficient can be set above 0.95. However, if the mean sensitivity value is -0.10 in winter, a difference of 0.18 from autumn, it indicates significant differences in user behavior during the autumn-winter transition, and the smoothing coefficient may drop to 0.70. The system will automatically smooth and correct the prediction to avoid misjudgments caused by sudden seasonal changes.

[0149] Through the above four steps, this invention achieves a complete chain from ensuring the consistency of input data, to dynamically capturing changes in sensitivity, to balancing the classification and splitting process across seasons, and finally smoothing the output across seasons.

[0150] A second embodiment of this application provides an electronic device, the electronic device comprising:

[0151] processor;

[0152] The memory is used to store a program, which, when read and executed by the processor, executes an intelligent classification method based on rainfall sensitivity and decision tree provided in the first embodiment of this application.

[0153] The third embodiment of this application provides a computer-readable storage medium storing a computer program thereon. When the program is executed by a processor, it executes an intelligent classification method based on rainfall sensitivity and decision tree provided in the first embodiment of this application.

[0154] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.

Claims

1. An intelligent classification method based on rainfall sensitivity and decision tree, characterized in that, include: Standardize the raw data from multiple sources to generate a standardized dataset; Based on the standardized dataset, a sensitivity index for user electricity consumption relative to changes in rainfall is calculated, and a sensitivity feature vector is constructed with the sensitivity index as the core. The sensitivity feature vector is fused with the seasonality and crop type features in the environmental attribute data to obtain a comprehensive sensitivity feature set; The sensitivity comprehensive feature set is input into a decision tree classification model constructed based on the Gini coefficient splitting criterion. During the training process, a cross-seasonal sample balancing mechanism and a rainfall anomaly compensation mechanism are introduced to form a training model with environmental robustness. The trained model is used to classify and identify new multi-source raw data. The step of calculating a sensitivity index of user electricity consumption relative to changes in rainfall based on the standardized dataset and constructing a sensitivity feature vector with the sensitivity index as the core includes: The rainfall data in the standardized dataset is segmented into multiple levels of intensity, dividing the rainfall conditions into four levels: light rain, moderate rain, heavy rain, and extreme rainfall. The average rainfall level is calculated for each level interval to generate multi-level rainfall intensity reference values. The user electricity consumption data in the standardized dataset is grouped and statistically analyzed accordingly. The average electricity consumption under different rainfall intensity reference values ​​is calculated and compared with the average electricity consumption under no-rain conditions to obtain a multi-level electricity consumption difference rate set. The multi-level electricity consumption difference rate set is subjected to time-series windowing processing. The mean and variance of the difference rate are calculated within a set continuous time window to capture the stability of users' response to rainfall under short-term fluctuations and long-term trends, and to generate a windowed sensitivity curve. The windowed sensitivity curve is normalized and offset corrected by introducing temperature and humidity data from environmental attributes as auxiliary variables and eliminating response bias caused by non-rainfall factors to obtain the corrected sensitivity sequence. The corrected sensitivity sequence is vectorized and combined, and the corrected values ​​under light rain, moderate rain, heavy rain and extreme rainfall conditions are arranged in sequence to form a complete sensitivity feature vector. A stability index obtained by variance measurement is added to this vector to characterize the user's overall response to rainfall. The step of fusing the sensitivity feature vector with seasonal and crop type features in environmental attribute data to obtain a comprehensive sensitivity feature set includes: The sensitivity feature vector and the seasonal feature are time-series aligned, and the response values ​​of different rainfall levels in the sensitivity feature vector are mapped to the corresponding seasonal windows to generate a sensitivity distribution table indexed by season. Crop type weights are introduced into the sensitivity distribution table. These crop type weights are proportionally allocated based on the user's main planting structure at different times. They are used to weight and correct the rainfall response values ​​under different crop types, generating a crop sensitivity correction matrix. The crop sensitivity correction matrix is ​​interactively extended to construct a three-dimensional interactive feature tensor including rainfall level, season, and crop type, so that each feature unit can reflect the response characteristics of a specific season and a specific crop under specific rainfall conditions, thereby obtaining a complete interactive feature tensor. The interaction feature tensor is normalized and compressed. The normalization coefficient is calculated based on the variability of each unit in the time dimension. Units with high information redundancy are dimensionality reduced, and units with high variability are retained and amplified to obtain the compressed core feature set. The core feature set is concatenated with the stability index in the sensitivity feature vector, and a cross-seasonal smoothing coefficient is added to form the final comprehensive sensitivity feature set, which is used to comprehensively characterize the multi-dimensional dynamic response relationship of user electricity consumption behavior to rainfall, season and crop type. The cross-seasonal smoothing coefficient is determined by comparing the differences in the mean sensitivity values ​​of adjacent seasons. The process involves inputting the sensitivity-integrated feature set into a decision tree classification model constructed based on the Gini coefficient splitting criterion. During training, a cross-seasonal sample balancing mechanism and a rainfall anomaly compensation mechanism are introduced to form an environmentally robust training model, including: Cross-seasonal stratified sampling is performed on the sensitivity comprehensive feature set, and data in each season are extracted in the same proportion to generate a cross-seasonal balanced training set with a balanced number of samples, in order to avoid the bias caused to the training model by the excessive proportion of samples in a certain season. Rainfall anomaly detection is performed on the cross-seasonal balanced training set. A reference interval is established using historical rainfall distribution. Extreme rainfall samples that exceed the interval threshold are individually labeled and their impact on split node calculation is reduced through a weight decay strategy to generate an anomaly-compensated cross-seasonal balanced training set. The cross-seasonal balanced training set after anomaly compensation is input into the decision tree construction module. The degree of classification purity reduction is calculated on each candidate feature dimension according to the Gini coefficient splitting criterion. The feature that can reduce impurity to the greatest extent is selected as the current splitting condition to generate the initial split tree structure. The cross-seasonal robustness of the split tree structure is evaluated by comparing the classification purity differences of sub-samples from different seasons at each split node. When the difference exceeds a set threshold, the influence of seasonal differences is reduced by adjusting the sample weights or introducing additional regularization factors, thereby obtaining a cross-seasonal robust split tree structure. The cross-seasonal robust split tree structure is subjected to multiple rounds of pruning operations. The error rate is minimized based on the validation set. Leaf nodes and branches that do not contribute enough information are deleted while retaining the environmental sensitivity features, resulting in a simplified cross-seasonal robust split tree structure. The simplified cross-seasonal robust split tree structure is jointly corrected with the stability index of the sensitivity integrated feature set. When the output category prediction is performed at each leaf node, the stability index is introduced as a correction factor. The confidence of feature combinations with large fluctuations is reduced to obtain the final training model, which is used to maintain stable classification performance in the face of cross-seasonal imbalance and rainfall anomalies.

2. The intelligent classification method based on rainfall sensitivity and decision tree according to claim 1, characterized in that, The standardization process for multi-source raw data to generate a standardized dataset includes: A unified time-series index is established for the multi-source raw data, and user electricity consumption data, rainfall data and environmental attribute data are synchronously aligned within the same time window to obtain a multi-source aligned dataset with time-series consistency. The multi-source aligned dataset is subjected to climate pattern-based interval normalization. The interval normalization is dynamically adjusted according to the joint probability of historical rainfall distribution and user electricity consumption distribution, so that the feature values ​​under different seasons and different climate conditions are mapped to a unified scale, resulting in a normalized dataset with dynamic adaptability. The normalized dataset with dynamic adaptability is subjected to feature decoupling based on mutual information constraints. The mutual information strength between user electricity consumption data, rainfall data and environmental attribute data is calculated. Feature dimensions with redundancy exceeding the threshold are removed, and independent contributing variables are retained through orthogonal mapping to generate a denoised key feature set. The scale of the denoised key feature set is recalibrated based on environmental sensitivity factors. A scale coupling factor is constructed by combining rainfall intensity fluctuation rate and electricity consumption sensitivity. Key features from different sources are uniformly measured and corrected to generate a standardized dataset with uniform dimensions and scale constraints.

3. The intelligent classification method based on rainfall sensitivity and decision tree according to claim 1, characterized in that, The process of classifying and identifying new multi-source raw data using the trained model includes: The new multi-source raw data is standardized and mapped in real time. The same dimensional and scale constraints as those in the training phase are adopted. User electricity consumption, rainfall and environmental attribute data are synchronously calibrated to generate a real-time standardized input set, ensuring that the input data and the feature space of the training model remain completely consistent. The real-time standardized input set is compared with the user's historical sensitivity features in a time series, and the sensitivity offset of the new data under the same season and the same crop type is extracted. The offset is then used to correct the real-time input so that it can dynamically reflect the difference between the new data and the historical pattern, thus obtaining the environmentally corrected input set. The environmentally corrected input set is input into the cross-seasonal robust splitting structure of the training model. Cross-seasonal discriminant weights are introduced into each splitting node to dynamically reconcile the splitting results of different seasons and output a preliminary classification path to ensure that the prediction results remain consistent in different seasons. The stability of the leaf node output of the preliminary classification path is tested. A stability index is introduced from the sensitivity comprehensive feature set. The confidence of sample categories with large sensitivity fluctuations is reduced, and the confidence of categories with stable sensitivity is increased, generating a prediction result after stability correction. The stability-corrected prediction results are combined with the cross-seasonal smoothing coefficient to calculate the final classification result. The cross-seasonal smoothing coefficient is used to quantify the continuity of user response at seasonal transitions and to smooth and correct prediction results with excessive differences between seasons. This results in dynamic and high-precision user category discrimination without the need for manually setting fixed thresholds.

4. The intelligent classification method based on rainfall sensitivity and decision tree according to claim 3, characterized in that, The real-time standardized mapping is achieved by calling the mean and variance parameters from the training phase.

5. The intelligent classification method based on rainfall sensitivity and decision tree according to claim 3, characterized in that, The sensitivity offset is the difference between the sensitivity value of the new input data and the historical sensitivity value.

6. The intelligent classification method based on rainfall sensitivity and decision tree according to claim 3, characterized in that, The cross-seasonal discrimination weight is used to proportionally harmonize the classification results of samples from different seasons at the split node.

Citation Information

Patent Citations

  • User electricity consumption prediction method and system

    CN116777049A

  • Country business management method based on big data

    CN117933946A