Charging load prediction method and system based on comprehensive similarity similar day screening

By combining the similarity-based similar day selection method and the XGBoost model with multi-dimensional features for charging load prediction, the problem of inconsistent similar day selection in existing technologies is solved, and high-precision and stable load prediction results are achieved.

CN121076794BActive Publication Date: 2026-01-09SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511630556.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2026-01-09
Estimated Expiration
2045-11-10

AI Technical Summary

Technical Problem

Existing charging load prediction methods rely on a few external features and ignore historical load patterns, resulting in inconsistent screening results for similar days and making it difficult to maintain high accuracy and generalization ability across different regions and time periods.

Method used

A similar day selection method based on comprehensive similarity is adopted. Through multi-feature comprehensive similarity calculation and optimization, a set of similar days with multi-dimensional features is constructed. The XGBoost model is used for supervised learning to select the most matching similar days. Load prediction is then performed by combining meteorological features and contextual features.

Benefits of technology

It improves the accuracy and stability of load forecasting, enhances adaptability to different regions and time periods, and ensures the consistency and representativeness of similar day screening results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121076794B_ABST
    Figure CN121076794B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of load prediction, and proposes a charging load prediction method and system based on comprehensive similarity similar day screening, which comprises similarity calculation and normalization of the obtained historical load characteristics, meteorological characteristics and context characteristics; taking the mean and standard deviation of the fusion similarity score as the target, the fusion weight of each similarity is solved to construct a standard similar day set; the day-to-feature of the meteorological characteristics and context characteristics similarity of the target day to be predicted and the candidate day is calculated as input, and the most matched similar day identification set is screened out through the trained model; and the load prediction result is obtained based on the meteorological characteristics and context characteristics of the target day to be predicted and the screened similar day identification set. Through the multi-feature comprehensive similarity calculation and optimized similar day screening, the present application realizes the construction of the similar day set based on multi-dimensional features, and provides a high-precision input feature selection framework for the charging load prediction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of load prediction, in particular to a charging load prediction method and system based on comprehensive similarity similar day screening. BACKGROUND

[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute prior art.

[0003] In time series prediction tasks, the similar day screening method has become an important means to improve the prediction accuracy and model generalization ability. The core idea is to find several sample days with similar characteristics to the prediction target day in the historical data, construct a similar day set, and then provide more representative and targeted input samples for the prediction model. Compared with the traditional method of directly using all historical data, this method can effectively reduce input redundancy and highlight typical patterns related to the target, thereby improving prediction efficiency and accuracy. In particular in the scenarios of electric vehicle charging load prediction, traffic flow prediction and the like, which are significantly influenced by external factors such as weather conditions, day type, etc., the traditional unified modeling method often fails to cover the diversity and complexity of the data, so the similar day method has become a key link in many current researches and practices.

[0004] At present, the similar day screening method for charging load prediction is mostly based on external features (such as temperature, weather type, etc.) or statistical features (such as mean, standard deviation) for similarity measurement and screening. Common techniques include feature vector distance-based measurement methods and clustering analysis to construct similar day sets. However, such methods generally have the following key problems: first, they rely too much on a small number of external variables such as weather and day type, ignoring the evolution law of historical load patterns, resulting in that in the case of consistent features but significant differences in load curves, the selected similar days cannot accurately reflect the load characteristics of the target day; second, the screening process is sensitive to feature weight settings, and the similar day screening results differ greatly between different regions or different target days, making it difficult to ensure the consistency and generalization ability of the method in actual deployment. SUMMARY

[0005] In order to solve the above problems, the present application provides a charging load prediction method and system based on comprehensive similarity similar day screening, which realizes the construction of similar day set based on multi-dimensional features through multi-feature comprehensive similarity calculation and optimized similar day screening, and further provides a high-precision input feature selection framework for electric vehicle charging load prediction.

[0006] In order to achieve the above purpose, the present application adopts the following technical solutions:

[0007] The first aspect of the present application provides a charging load prediction method based on comprehensive similarity similar day screening, comprising the following steps:

[0008] Similarity calculations were performed on the acquired historical load characteristics, meteorological characteristics, and contextual characteristics, and then normalized.

[0009] With the goal of minimizing the mean and standard deviation of the fusion similarity score, a multi-objective optimization model is constructed to solve for the fusion weight of each similarity. The calculated similarity of historical load data, meteorological data, and contextual features is then fused, and a standard set of similar days is constructed based on the fusion similarity value.

[0010] The daily pair features, which calculate the similarity between meteorological features and contextual features of the target day and the candidate day, are used as input. The trained XGBoost model is then used to select the set of most matching similar days. The XGBoost model is trained using supervised learning methods, which construct daily pair features from the standard set of similar days as training samples.

[0011] The data, based on the meteorological and contextual characteristics of the target day to be predicted and the data of the selected similar day identification set, is transmitted to the load prediction model for identification to obtain the load prediction result.

[0012] A second aspect of the present invention provides a charging load prediction system based on comprehensive similarity-based similarity day screening, comprising:

[0013] The similarity calculation module is configured to perform similarity calculations on the acquired historical load features, meteorological features, and contextual features, and then normalize them.

[0014] The optimal weight calculation module is configured to construct a multi-objective optimization model with the objective of minimizing the mean and standard deviation of the fused similarity scores. It solves for the fusion weights of each similarity score, fuses the calculated similarities of historical load data, meteorological data, and contextual features, and constructs a standard set of similar days based on the fused similarity values. ;

[0015] The similar day matching and filtering module is configured to take the day-to-day features, which calculate the similarity between the meteorological features and contextual features of the target day and candidate days, as input, and then use the trained XGBoost model to filter out the set of most matching similar days. The XGBoost model, through a standard set of similar days... Daily features are constructed as training samples, and a supervised learning method is used for training.

[0016] The prediction module is configured to identify similar days based on the meteorological and contextual features of the target day to be predicted and a filtered set of similar days. The data is transmitted to a load prediction model as input data to obtain a load prediction result.

[0017] Compared with the prior art, the present application has the following advantages:

[0018] The method introduces a comprehensive similarity calculation mechanism in the similar day screening process, effectively integrates load characteristics, weather information and context variables, and avoids the matching failure problem caused by relying only on a small number of external features in the traditional method. Through normalization and multi-objective optimization, the relative balance and adaptability of various feature similarities in the fusion process are ensured, and the consistency and generalization ability of the similar day screening result are improved. The XGBoost model is used to further mine the nonlinear mapping relationship between the weather and context features and the load mode, to accurately screen high-correlation samples from the preliminary similar day set, and to enhance the representativeness of the final input data. The overall process improves the accuracy, model stability and adaptability to different regions and different time periods of load prediction.

[0019] The advantages of the present application and the advantages of the additional aspects will be described in detail in the following specific embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0020] The drawings accompanying the specification of the present application form a part thereof and serve to provide further understanding of the present application, the illustrative embodiments of the present application and its description serve to explain the present application, and do not constitute limitations thereof.

[0021] Figure 1 is a flowchart of the charging load prediction method of embodiment 1 of the present application;

[0022] Figure 2 is a prediction result graph under the same prediction model structure based on different feature combinations in the experimental example of embodiment 1 of the present application;

[0023] Figure 3(a) is a first prediction comparison graph of different similar day screening schemes in the experimental example of embodiment 1 of the present application;

[0024] Figure 3(b) is a second prediction comparison graph of different similar day screening schemes in the experimental example of embodiment 1 of the present application;

[0025] Figure 4 is a first comparison graph of the load curve of the reference day and the similar day set corresponding to the traffic cell randomly selected in the experimental example of embodiment 1 of the present application;

[0026] Figure 5 is a second comparison graph of the load curve of the reference day and the similar day set corresponding to the traffic cell randomly selected in the experimental example of embodiment 1 of the present application;

[0027] Figure 6is a comparison chart of F1 scores when different similar day scales are used in the experimental example of embodiment 1 of the present application;

[0028] Figure 7 is the change range of F1 scores relative to the reference value when different similar day scales are used in the experimental example of embodiment 1 of the present application;

[0029] Figure 8 is a similarity calculation flowchart of embodiment 1 of the present application. DETAILED DESCRIPTION

[0030] The present application will be further described below in conjunction with the drawings and embodiments.

[0031] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise indicated, all technical and scientific terms used herein have the same meaning as would be commonly understood by one of ordinary skill in the art to which the present application belongs.

[0032] It should be noted that the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit exemplary embodiments according to the present application. As used herein, the singular form is intended to include the plural form unless the context clearly indicates otherwise, and it should be understood that when the terms "comprise" and / or "include" are used in the specification, there is a presence of a feature, step, operation, device, component, and / or combinations thereof. It should be noted that the various embodiments and features in the present application can be combined with each other without conflict, and the embodiments will be described in detail below in conjunction with the drawings.

[0033] Embodiment 1

[0034] In the technical solutions disclosed in one or more embodiments, as shown in Figures 1 to 8 The charging load prediction method based on comprehensive similarity similar day screening includes the following steps:

[0035] Step 1, similarity calculation is performed on the obtained historical load characteristics, meteorological characteristics and context characteristics, and normalization processing is performed;

[0036] Step 2, a multi-objective optimization model is constructed to minimize the mean and standard deviation of the fusion similarity score, and the fusion weight of each similarity is obtained by solving, the historical load data similarity, meteorological data similarity and context characteristic similarity are fused, and a standard similar day set is constructed based on the fusion similarity value ;

[0037] Step 3, calculate the day pair features of the meteorological features and context features of the target day to be predicted and the candidate day as input, and select the most matched similar day recognition set through the trained XGBoost model ; the XGBoost model is trained by the standard similar day set The day pair features are used as training samples to train the model by using a supervised learning method.

[0038] Step 4, based on the meteorological features and context features of the target day to be predicted and the data of the selected similar day recognition set , the input data is transmitted to the load prediction model for recognition to obtain the load prediction result;

[0039] The embodiment provides a similar day screening method under multi-source data, and models the similar day screening as a pattern matching task. The daily load pattern is not only affected by meteorological conditions and context factors, but also deeply affected by historical load features. The embodiment fully integrates historical load, meteorological conditions and context features in multi-source information, and automatically learns the weight and action mechanism of different factors in similarity discrimination through supervised learning. Under this framework, the optimal similar day set of the target day can be screened based on known information at the prediction time, which not only overcomes the limitation of traditional methods using only weather and context information, but also avoids the unreasonable assumption of looking at future load data, thereby significantly improving the practicality and accuracy of similar day screening in actual prediction tasks.

[0040] The method of the embodiment introduces a comprehensive similarity calculation mechanism in the similar day screening process, effectively integrates load features, meteorological information and context variables, and avoids the matching failure problem caused by relying only on a small number of external features in traditional methods. Through normalization and multi-objective optimization, the relative balance and adaptability of various feature similarities in the fusion process are ensured, and the consistency and generalization ability of the similar day screening result are improved. The XGBoost model is used to further mine the nonlinear mapping relationship between meteorological and context features and load patterns, to accurately screen high-correlation samples from the preliminary similar day set, and to enhance the representativeness of the final input data. The overall process improves the accuracy, model stability and adaptability to different regions and different time periods of load prediction.

[0041] In the embodiment, historical load, meteorological and context features are comprehensively considered and fusion similarity is calculated, which effectively solves the problem of relying only on external variables and strengthens the consideration of load curve patterns. The fusion weight is automatically solved by multi-objective optimization, and the XGBoost model is introduced to realize supervised matching, which overcomes the problem of manual weight setting and improves the consistency and generalization ability of the method in different scenarios.

[0042] The content related to the embodiment mainly contains four aspects: similarity calculation method; optimal weight solving algorithm; structured reconstruction of pattern matching problem and training sample construction; similar day screening; specific description as follows:

[0043] In step 1, the similarity is calculated for the obtained historical load characteristics, specifically: for the obtained historical load characteristic sequence data, the shape matching based time series distance measurement method (referred to as ShapeDTW algorithm) is used to calculate the historical load data similarity, including the following steps;

[0044] Step 101, for the obtained historical load characteristic time series data, extract the local shape descriptor sequence for each point on each historical load curve;

[0045] Wherein, each historical load curve can be single-day load data as a historical load curve;

[0046] Step 1011, for any time point x in the historical load time series, a neighborhood is constructed with time point x as the center, and the load value set in the neighborhood range of the time point x is extracted; for each point in the neighborhood, difference operation is performed with the center point to obtain a local difference set for describing the shape change of the point.

[0047] Specifically, for any point in the historical load time series , first examine the points in a neighborhood defined by the window size parameter around it, and calculate the difference set of all points in the neighborhood from the center point ;

[0048] (1);

[0049] Step 1012, construct the local difference set into a histogram vector, and normalize the histogram vector to obtain the local shape descriptor sequence corresponding to each point;

[0050] In order to standardize this morphological information, the difference set is constructed into a histogram. This histogram is a vector that counts the frequency of different size differences, thereby quantifying the local shape. In order to make it not affected by the absolute value of the load, the range of the histogram is dynamic and is normalized. The normalized histogram vector is the shape context descriptor of point , which is expressed as follows:

[0051] (2);

[0052] (3);

[0053] wherein: is the number of bins of the histogram, is the point corresponding to the i-th component of the histogram vector, is the original count value; is a set minimum value to prevent the denominator from being zero.

[0054] Similarly, for any point in another historical load time series of the same time step , the same calculation is made to obtain its corresponding shape descriptor . Through the above steps, the one-dimensional load sequence is converted into a multi-dimensional sequence composed of shape descriptors. The two descriptor sequences corresponding to the above two curves are denoted as and .

[0055] Step 102, based on the obtained local shape descriptor sequence corresponding to the historical load feature sequence, DTW similarity calculation is performed to obtain the ShapeDTW distance of any two historical load curves;

[0056] Specifically, the DTW algorithm finds the optimal alignment path between two historical load feature sequences by constructing a cumulative cost matrix . First, the cost (or distance) of aligning two descriptors and is defined. In this embodiment, the Euclidean distance is used as the cost function to measure the distance between two points, and the calculation formula is as follows:

[0057] (4);

[0058] Based on the point-to-point cost function, the historical load cumulative cost matrix is constructed. Each element in the cumulative cost matrix can be calculated by the following recursive relationship:

[0059] (5);

[0060] (6);

[0061] The final is the ShapeDTW distance of the two sequences. The smaller the distance value, the more similar the local shape and fluctuation pattern of the two original curves. ​

[0062] Traditional DTW is sensitive to the absolute value of the sequence. For example, two load curves have the same shape, but one is twice the value of the other, traditional DTW will consider that there is a big difference. The embodiment adopts ShapeDTW algorithm, and the comparison result is robust to the scaling of the absolute value by extracting local morphological features.

[0063] In step 1, for the meteorological characteristics, the time series data of the historical meteorological characteristics obtained is used to calculate the meteorological characteristic similarity by using a multi-dimensional time series distance measurement method (referred to as multi-dimensional DTW algorithm);

[0064] Similar to ShapeDTW, the core of multi-dimensional DTW is also the distance between points, which directly acts on the original multi-dimensional vector, rather than the converted shape descriptor. For historical meteorological characteristics, assume that there are two multi-dimensional time series with the same time step and , wherein is the quantitative value of the meteorological data at each time point, that is, the state vector.

[0065] The method for calculating the meteorological characteristic similarity comprises the following steps:

[0066] Step 111, the Euclidean distance is used to calculate the distance between each data point in the two meteorological characteristic sequence data;

[0067] First, the distance between a state vector in the meteorological sequence and a state vector in the meteorological sequence is calculated. The Euclidean distance is used to measure the straight-line distance of two vectors in a multi-dimensional space, and the calculation formula is as follows:

[0068] (7);

[0069] In the formula, is the dimension of the meteorological vector; and are the kth feature value of the vector and , respectively.

[0070] Step 112, based on the distance of the meteorological data points obtained, the cumulative cost matrix is calculated to obtain the DTW distance of any two meteorological time series data;

[0071] The meteorological cumulative cost matrix is constructed, and the recursive relationship is as follows:

[0072] (8);

[0073] (9);

[0074] The final obtained is the final DTW distance of the two multi-dimensional time series, and the smaller the value is, the more similar the evolution paths of the two sequences in the multi-dimensional space are.

[0075] In step 1, for the obtained historical day context features, the absolute value of the difference value of all time points of different days is calculated by using the binary difference method, and the sum of the absolute value of the difference value is taken as the similarity of the context features of different days.

[0076] Specifically, in the embodiment, the context features are historical day features, such as whether it is a weekend, whether it is a holiday, etc.; the context features are a sequence composed of discrete binary labels, such as 0 or 1.

[0077] Suppose two context sequences are respectively:

[0078] (10);

[0079] In the formula:

[0080] ;

[0081] Wherein, and are multi-dimensional vectors of the two sequences at time point i; is the dimension of the context feature vector; and are the kth feature values of the two sequences at time point i.

[0082] The distance between the two context sequences is calculated , and the formula is as follows.

[0083] (11);

[0084] The absolute difference values of all time points are accumulated to obtain the final similarity , and the distance represents the total distance of all context sequences of the two days.

[0085] (12);

[0086] The smaller the distance value is, the more similar the two groups of data context sequences are.

[0087] In step 1, the RobustScaler+MinMax normalization algorithm is used to normalize the obtained similarity value, including the following steps:

[0088] Step 121, calculate the median of each similarity value, and scale transform each similarity value;

[0089] For each data point in the original similarity sequence , its robustly scaled value is calculated as follows:

[0090] (13);

[0091] where: the median of the entire similarity value sequence ; : the interquartile range of the entire sequence .

[0092] A new sequence is calculated; where, the center is roughly around 0, the scale is defined by IQR, and the influence of extreme values has been significantly weakened.

[0093] Step 122, normalize the scaled similarity values based on the maximum and minimum values;

[0094] Scale transform ; for each data point in the sequence , its final normalized value is calculated as follows:

[0095] (14);

[0096] where: is the minimum value of the sequence . is the maximum value of the sequence .

[0097] Each element in the normalized sequence strictly lies in the interval [0, 1].

[0098] After calculating three independent similarity normalization scores, a key problem arises: how to fuse these three-dimensional scores into a unified index to comprehensively reflect the overall similarity between the candidate day and the benchmark day. If a fixed weight is simply used for weighting, although it is convenient to implement, it is difficult to adapt to the differences in the similarity distribution of candidate days under different benchmark days, which may lead to large fluctuations or lack of concentration in the fusion results.

[0099] In this embodiment, the multi-objective optimization method NSGA-II is introduced to dynamically search for the optimal combination of weight vectors of the three types of features. To make the similar day set have higher consistency and representativeness while being similar as a whole, the optimization objectives are set to simultaneously minimize the mean and standard deviation of the weighted similarity score. This strategy improves the stability and reliability of the fusion result, and helps to provide more representative input samples for subsequent load forecasting models.

[0100] In step 2, for the weight optimization algorithm, specifically, the NSGA-II multi-objective optimization algorithm is used to obtain the fusion weight of each similarity. First, a multi-objective optimization model is constructed by taking the minimization of the mean and standard deviation of the fusion similarity score as the objective.

[0101] For a given reference day and a region , there is a set of similarity score sets of historical days , where each similarity is a three-dimensional score tuple. The goal is to find a weight vector such that the final weighted score sequence calculated by the weight can achieve optimality on two competing objectives at the same time.

[0102] Among them, a certain day is set as the reference day, and the similar days of the reference day are screened, so as to divide the historical days into multiple similar day sets;

[0103] The calculation formula of the final weighted score is as follows:

[0104] (15);

[0105] In the formula: is the final weighted similarity score of the th historical day; is the three-dimensional weight vector to be optimized; is the three-dimensional normalized similarity score tuple of the th historical day; and are the weight and similarity score of the th type of feature, respectively.

[0106] Specifically, the mean and standard deviation of the fusion similarity score are taken as two independent optimization objectives to form a double-objective optimization problem, so as to simultaneously improve the overall representativeness and internal consistency of the similar day set. The double-objective optimization model is defined as follows:

[0107] Objective 1: Minimize the mean score ):

[0108] (16);

[0109] wherein: is the total number of samples in the historical day set; is the weighted score of the th historical day, denotes the sequence of weighted scores.

[0110] This objective aims to maximize the overall similarity between the filtered set of similar days and the benchmark day.

[0111] Objective two: Minimize Standard Deviation, of the weighted scores:

[0112] (17);

[0113] This objective aims to maximize the internal consistency or homogeneity of the filtered set of similar days.

[0114] The constraint condition of the multi-objective optimization model is the weight vector must satisfy the standard simplex constraint, which is as follows:

[0115] (18);

[0116] These two objectives are essentially competitive: a weight scheme that is too biased towards a certain type of score may result in a very low mean, for example, looking only at the load score may result in increased volatility of the score due to the neglect of other factors. This embodiment adopts the NSGA-II algorithm to find a set of optimal trade-off solutions between the two objectives.

[0117] In step 2, the multi-objective optimization model is solved using the NSGA-II-based multi-objective optimization algorithm, including the following steps:

[0118] Step 21, initialization: randomly generate weight vectors to construct an initial population;

[0119] Specifically, an initial population consisting of random weight vectors is generated. Each weight vector is sampled through a Dirichlet distribution to ensure that it automatically satisfies the constraint of

[0120] Step 22, iterative evolution, including population breeding, evaluation and sorting, and environmental selection;

[0121] Specifically, in each generation , the algorithm performs the following operations:

[0122] Step 221, Propagation: based on the inheritance operator, a child population of the same size is generated from the current population ; ;

[0123] Specifically, the genetic operators include simulated binary tournament selection, simulated binary crossover and polynomial mutation, etc.

[0124] Step 222, Evaluation and Ranking: the parent population and the child population are merged, and the objective values are calculated based on the objective functions of the multi-objective optimization model, the individuals in the merged population are non-dominantly ranked based on the objective values, a plurality of Pareto levels are obtained, and the Pareto level in each level is calculated;

[0125] Specifically, the parent and child are merged into a temporary population of size .For each weight vector in the merged population , the two objective functions and defined above are evaluated. Then, the non-dominant sorting is performed on the merged population , which is divided into a plurality of Pareto levels . At the same time, the crowding distance is calculated for the solutions in each level to maintain the diversity of the solutions.

[0126] Step 223, Environmental Selection: based on the Pareto level, the high level is preferentially selected, and in the same level, the large crowding distance is preferentially selected, from the merged population , the best individuals are selected to form a new population of the next generation .

[0127] Step 23, Selection of Optimal Solution: after the iteration ends, the minimum standard deviation of the similarity weighted score is set as the optimal solution preference, and the optimal solution is selected from the Pareto front composed of a plurality of optimal solutions as the optimal weight vector , that is, as the fusion weight of each similarity;

[0128] After a preset number of generations, the algorithm terminates. At this time, the first Pareto front in the final population contains a plurality of equally excellent and non-dominant weight solutions. These solutions represent different trade-offs between "overall similarity" and "internal consistency". In order to obtain a certain solution, a preference is introduced in this embodiment. In the framework, the solution on the Pareto front is selected as the final optimal weight vector This selection reflects the high value attached to the stability and reliability of the screening results.

[0129] Step 24, Similar Day Set Screening: Similarities are fused based on the fusion weight of each similarity, and the similarity set with the highest similarity to each reference day is obtained based on the fused similarity score;

[0130] Specifically, for each region find its exclusive optimal weight vector After that, the final comprehensive similarity score of all historical days in this region can be calculated. For any historical day , the three-dimensional score tuple , the final score is obtained by the following dot product operation:

[0131] (19);

[0132] In the formula: is the final comprehensive similarity score of the th historical day; is the optimal weight vector found for region .

[0133] It is a single, dynamically optimized scalar value that fairly fuses three types of heterogeneous information. The smaller the value, the more similar the historical day is to the reference day as a whole. After calculating the comprehensive similarity scores of all historical days, the final similar day set is screened according to these scores. Each historical day in each set is a similar day;

[0134] Further, the goal of the similar day set screening process is, for each region , to identify a subset of size from the entire historical day set , which contains the top historical days similar to the reference day. The specific screening process is as follows: arrange the final scores of all historical days in region in ascending order, and select the top historical days. This process can be represented by the following formula:

[0135] (20);

[0136] In the formula: is the similar day set of region ; Operation represents returning the identifiers of the top historical days with the smallest scores.

[0137] For the single feature or the combination of two features, the screening process is the same, only need to replace the single or combined score calculated by the corresponding scheme.

[0138] In the load forecasting, a key premise is to select a set of similar days with similar load patterns in history for the target day to be predicted as the input sample of the model. In this embodiment, a high-quality standard similar day set has been constructed based on the comprehensive method of steps 1 and 2 described above , which includes the load data, weather data and context features of the similar days. However, it should be pointed out that the set inevitably refers to the real load data of the target day to be predicted in the construction process. In the real prediction scenario, the load data of the target day is leaked, which will lead to an idealized evaluation result, so as to fail to truly reflect the generalization ability of the model. Therefore, the core challenge is how to find a set of optimal similar days only by using the context information and weather data of the target day without relying on its load data. In this embodiment, the similar day screening is modeled as a pattern matching task.

[0139] The daily load pattern is comprehensively affected by various situational factors such as holidays and weather conditions, so these factors need to be fully utilized to measure the similarity between the target day and the historical day. In this embodiment, the similar day screening is converted into a supervised learning task, which automatically learns the mechanism of situational factors in similarity discrimination through a classification model, so as to efficiently screen the optimal similar day set of the target day without relying on future load data.

[0140] After the context information and weather conditions of the target day and the candidate historical day are given, it is necessary to judge whether they are similar in load pattern. For this purpose, it is necessary to structure the problem: on the one hand, the context attributes and weather variables of the date are encoded into feature vectors, and the similarity between the target day and the historical day is calculated accordingly; on the other hand, the original "similar day sorting" problem is converted into a binary classification task of "whether to match", so that supervised learning methods can be used to construct training samples and learn the discrimination rules.

[0141] Further, in step 3, the XGBoost model is trained by the day pair features constructed from the standard similar day set , including the following steps:

[0142] Step 31, set the target day, calculate the context similarity and weather feature similarity between the target day and the candidate historical day as day pair similarity features;

[0143] ​In the original feature-driven pattern matching construction, the date attribute and weather conditions need to be converted into quantifiable feature indicators first to characterize the similarity between the target day and the candidate historical day. Specifically, based on the context information of the date and the weather conditions, the context feature similarity and the weather feature similarity between the target day and the candidate historical day are calculated and normalized to obtain and two feature vectors, whose calculation formulas are shown in equations (7), (9) and (11). Finally, the two are combined to form a two-dimensional feature vector , which is used to characterize the overall similarity feature of the day pair, and is expressed as follows:

[0144] (21);

[0145] Step 32, determine whether the candidate historical day belongs to the similar day set of the target day. If yes, construct a positive sample pair; otherwise, construct a negative sample pair. The positive sample pair and the negative sample pair are used as training samples.

[0146] After obtaining the similarity feature of each day pair, the next task is no longer to directly sort the historical dates, but to convert it into a more basic binary classification problem: for a given day pair composed of a target day and a historical day, determine whether the load pattern matches. For this purpose, the known similar day set is used to construct training samples: each sample is composed of input features and corresponding labels , that is . If the historical day belongs to the standard similar day set of the target day , the label is a positive sample, otherwise it is a negative sample, which is shown as follows:

[0147] Positive sample generation:

[0148] (22);

[0149] In the formula, is a target day in the training set, is the standard similar day set of the target day, which is the input of the XGBoost model; is a randomly selected historical day in the set.

[0150] Negative sample generation:

[0151] (23);

[0152] In the formula: is the same target day , and one or more dates are randomly selected from the remaining non-similar days.

[0153] After the determination by equations (22) and (23), the positive sample is defined. tags A value of 1 indicates pattern matching; this defines the input features of negative samples. tags A value of 0 indicates a pattern mismatch.

[0154] By repeating the above construction process for all target days in the training data, a large-scale training sample set containing positive and negative sample pairs can be obtained to characterize the matching relationship between day pairs. These samples provide structured input for subsequent supervised learning, enabling the model to learn the discrimination rules between target days and historical days, thus laying a solid data foundation for the training and optimization of similar day selection.

[0155] Step 33: Construct an XGBoost model. Extract day-to-day features of meteorological features and contextual features similar to the target day and candidate day from the training sample data as input, and train the XGBoost model with the matching relationship as output.

[0156] After obtaining structured day-pair training samples, a supervised learning model is used to complete the similarity discrimination task. In this embodiment, the eXtreme Gradient Boosting (XGBoost) model is selected as the classification model. XGBoost constructs a nonlinear mapping between the meteorological and contextual features of the date pair and the target date through an additive tree model, and simultaneously utilizes first-order and second-order gradient information during the iteration process to achieve efficient optimization. With this mechanism, the model can output matching probabilities for the target date and candidate historical dates, thereby realizing day-pair similarity evaluation and screening based on multi-source features.

[0157] Through multiple rounds of iterative training, XGBoost ultimately generated a dataset composed of... Fusion model composed of decision trees This model learns a complex nonlinear mapping between day-to-day features and matching relationships. In application, this model can be used for any target day. With candidate historical day The similarity between them is quantitatively assessed, and the best matching result is selected accordingly.

[0158] In constructing day-to-day features, the first step is to define the search space for similarity evaluation, which is the set of all available candidate historical days for matching. :

[0159] (twenty four);

[0160] Subsequently, for the candidate set each of the historical days , and build the corresponding day-pair feature vector by calculating the similarity of the meteorological features and the contextual features:

[0161] (25);

[0162] The trained XGBoost model will pass the feature vector to the leaf nodes of each decision tree layer by layer, and calculate the sum of all leaf weights to get an original score .

[0163] (26);

[0164] Since the similarity discrimination is modeled as a binary classification task, the original score needs to be mapped to a probability value in the interval [0, 1] through the Sigmoid function, as a quantitative indicator of the "matching degree" of the date pair, i.e. the similarity matching score:

[0165] (27);

[0166] For the target day and all historical days in the candidate set , the matching scores are calculated, and finally the scores are sorted and filtered according to the scores, which is represented by the formula:

[0167] (28);

[0168] Further, the matching scores of all historical days in the candidate set are sorted in descending order, and the historical days corresponding to the top scores are selected . This set of best matching historical days is the final screening result, which is the similar day recognition set and the input of the load prediction model.

[0169] It should be noted that the standard similar day set is the similar day set used to train the XGBoost model; the similar day recognition set is the similar day set of the target day identified by the trained XGBoost model in the prediction stage, which is the input data of the load prediction model. Among them, the trained XGBoost model uses meteorological data and contextual features to match similar days, without using load data. Through meteorological data and contextual features, it can be concluded that the XGBoost model forms a process from using load data to not using data.​

[0170] In step 4, the load forecasting model can be a MLP, LSTM, etc. The most matching similar day identification set is used to enhance the prediction performance of the MLP, LSTM, etc.

[0171] To evaluate the performance of the proposed model, this embodiment carries out an empirical study based on the electric vehicle charging load data of a certain city S. The data covers the time range from September 1, 2022 to February 28, 2023, a total of 6 months. The data set used in the study contains three different data styles, which are:

[0172] Charging load data: records the electric vehicle charging power of each traffic zone at an hourly interval, used to depict the time sequence variation of regional charging demand;

[0173] Weather feature data: including temperature, air pressure, humidity, precipitation, etc. features, used to reflect the influence of weather on charging load;

[0174] Context feature data: including workday / weekend, holiday, etc. labels, used to reflect the external regularity of travel behavior and charging demand.

[0175] The city S area is further divided into 275 traffic zones to construct a spatial network structure, thereby introducing spatial correlation in the model. The recording interval of the load data is 1 hour, providing continuous time sequence input for model training and prediction.

[0176] To evaluate the effectiveness of the multi-dimensional similar day screening method based on load, weather, and context features, this embodiment carries out analysis from three aspects: empirical verification of prior assumptions, error comparison under different feature combination schemes, and visualization of similar day effect of typical benchmark day.

[0177] Firstly, to verify the proposed prior assumption that the load pattern is driven by multiple features, and that all three types of features have a significant impact on load behavior, the performance of input schemes based on different feature combinations under the same prediction model structure is compared. The experimental results are shown in Figure 2 , where the prediction error of the similar day set selected using only load features (Load) is significantly higher than that after introducing weather (L+W), context (L+C), and fusion of three types of features (L+W+C). Especially when all three types of features are introduced, the prediction error reaches the lowest value, thereby verifying the complementarity and necessity of three types of features for load prediction at the empirical level, supporting the strategy of using multi-source information for similar day screening.

[0178] ​wherein L+W represents the combination of load data and weather data; L+C represents the combination of load data and context data; and L+W+C represents the combination of load data, weather data, and context data.

[0179] On this basis, seven similar day screening schemes are constructed, which are based on single feature, any two feature combinations, and three feature fusion strategies respectively. The performance of each scheme in load forecasting is evaluated. For the two feature combination schemes, a double-objective optimization method based on NSGA-II is used to search for weights, but only two of the three feature dimensions are selected for combination optimization, and the remaining feature weights are set to zero, for example: the weight of the context feature in the similar day set identification obtained by the fusion of load data and weather data is 0. The optimization objective remains unchanged, and the mean and standard deviation of the similarity score are minimized to ensure the overall similarity and consistency of the similar day set. The final optimal weight is used to fuse the similarity scores of the two features to generate a weighted total score, and the similar day set is screened according to the score. When the number of similar days (i.e. the standard similar day set The three-feature fusion scheme performs best in MAE and RMSE indicators when the number of similar days takes different values, significantly better than other single or double feature combinations, indicating that multi-dimensional fusion similarity can more effectively screen out historical samples highly similar to the target day. Figure 3(a) is a comparison chart of load curves of the benchmark day and the similar day set when the number of similar days is Figure 3(b) is a comparison chart of load curves of the benchmark day and the similar day set when the number of similar days is wherein the inventive method represents the charging load forecasting method of the embodiment.

[0180] Finally, Figure 4 and Figure 5 a comparison chart of load curves of the benchmark day and the similar day set is shown. It can be seen that the similar days screened by the three-feature fusion scheme not only have superior prediction accuracy, but also have highly consistent curve shape, peak-valley structure, and change trend with the target day, further verifying the accuracy and interpretability of the proposed method in similarity matching.

[0181] ​​To verify the effectiveness of the XGBoost model in daily similarity and comprehensively evaluate the effectiveness of the method in similar day selection, a performance evaluation system based on set comparison and probabilistic discrimination was constructed. The experimental data covered the complete date range from September 1, 2022 to January 4, 2023. In this system, the core measurement dimensions consist of a set of quantitative indicators, including precision, recall, F1 score, and area under the curve (AUC). The specific calculation results are shown in Table 1.

[0182] Table 1 Evaluation Indicators;

[0183]

[0184] As can be seen from Table 1, for different numbers of similar days Under these conditions, precision, recall, and F1 score remained consistent, a characteristic stemming from the similar day identification set in the evaluation setting. Similar Day Sets The size of the set is This means that all three indicators are determined by the size of the intersection of the two sets. With With the increase in the number of candidate days, the model can incorporate more candidate days, and the proportion of truly similar days covered increases. Therefore, all three metrics show a gradual upward trend, rising from 0.75 to 0.92. Meanwhile, the AUC remains above 0.93 across all scales, and... The model reached a maximum value of 0.9441, indicating that it consistently maintains stability and superiority in global ranking and discrimination capabilities. Overall, this method effectively captures truly similar days at different scales, and achieves high matching performance at medium scales, balancing accuracy and set simplicity.

[0185] In practical applications, how to determine the set of similar dates? scale This presents a trade-off: if the dataset is too large, it may introduce redundant and noisy data, affecting the accuracy of predictions; if the dataset is too small, it may miss key information, reducing the model's coverage. To achieve the optimal balance between coverage and data simplicity, this embodiment further employs the elbow rule to analyze the F1 score curve of the validation set. The specific results are as follows... Figure 6 and Figure 7 As shown in the figure, when When F1 is close to the inflection point, further increasing the candidate day's yield gradually decreases, and reducing the size of the set may miss key similar day information. Therefore, the reference value of the size of the similar day set in the subsequent prediction model input is preferably 35, achieving a balance between coverage and data simplicity. After screening the optimal similar day set, these representative samples can be directly used as one of the input features of the prediction model. Subsequently, the prediction process will combine the load pattern of the similar day set, historical load data, weather information, and contextual features to construct a complete input vector for various models to perform load prediction.

[0186] In this embodiment, the load prediction performance of different prediction models with and without the introduction of the similar day set is compared, and the results are shown in Table 2. For various models such as XGBoost, SVR (Support Vector Regression), multi-layer perception MLP, and long short-term memory network LSTM, the introduction of the similar day set can effectively reduce the prediction error. Taking MLP as an example, after adding the similar day set, the mean absolute error (MAE) decreases from 143.12 to 63.56, and the root mean square error (RMSE) decreases from 161.26 to 77.98, and the prediction accuracy is significantly improved; the error of the LSTM model is also reduced. The traditional machine learning method also shows a trend of error reduction after adding the similar day set.

[0187] Table 2 prediction result table;

[0188]

[0189] The above results show that by introducing the similar day set, the model can more accurately depict the load pattern of the target day, fully utilize historical load and multi-source feature information, and thus significantly improve the prediction accuracy and stability. More importantly, the similar day screening method of this embodiment can achieve optimal similar day selection without relying on future load data in the prediction stage, avoiding the limitation of requiring target day load in the traditional golden set method. This technical scheme not only verifies the effectiveness of similar day screening in actual load prediction, but also provides reliable data support and generalizable technical foundation for heterogeneous charging load dispatching optimization.

[0190] Embodiment 2

[0191] Based on embodiment 1, the charging load prediction system based on comprehensive similarity similar day screening is provided in this embodiment, which includes:

[0192] The similarity calculation module is configured to calculate the similarity for the obtained historical load features, meteorological features, and contextual features, respectively, and perform normalization processing;

[0193] The optimal weight solving module is configured to construct a multi-objective optimization model aiming at minimizing the mean and standard deviation of the fusion similarity score, and solve the fusion weight of each similarity, and fuse the calculated historical load data similarity, meteorological data similarity and context feature similarity, and construct a standard similar day set based on the fusion similarity value .

[0194] The similar day matching and screening module is configured to calculate the day pair feature of the meteorological feature and context feature similarity of the target day to be predicted and the candidate day as input, and screen out the most matched similar day recognition set through the trained XGBoost model ; the XGBoost model constructs the day pair feature as a training sample through the standard similar day set , and is trained by using a supervised learning method;

[0195] The prediction module is configured to transmit the data of the meteorological feature and context feature of the target day to be predicted and the screened similar day recognition set to the load prediction model as input data for identification to obtain a load prediction result.

[0196] Further, a multi-objective optimization algorithm based on NSGA-II is used to solve the multi-objective optimization model, including the following steps:

[0197] Randomly generate a weight vector to construct an initial population;

[0198] Iterative evolution, including population breeding, evaluation and sorting, and environment selection;

[0199] After the iteration is completed, the minimum standard deviation of the similarity weighted score is set as the optimal solution preference, and the optimal solution is selected from the Pareto frontier composed of multiple optimal solutions as the optimal weight vector , that is, as the fusion weight of each similarity;

[0200] The similarities are fused based on the fusion weight of each similarity, and the similarity set with the highest similarity to the set reference day is obtained based on the fusion similarity score.

[0201] Further, the XGBoost model trains the day pair feature as a training sample through the similar day set, including the following steps:

[0202] Set the target day, calculate the context similarity and meteorological feature similarity between the target day and the candidate historical day as the day pair similarity feature;

[0203] If the candidate historical day belongs to the similar day set of the target day, a positive sample pair is constructed, otherwise, a negative sample pair is constructed, and the positive sample pair and the negative sample pair are used as training samples;

[0204] An XGBoost model is constructed, and meteorological features and context features of the target day and the candidate day are extracted as input, and a matching relationship is used as output to train the XGBoost model.

[0205] It should be noted that each module in the embodiment corresponds to each step in Embodiment 1 one by one, and the specific implementation process is the same, which will not be repeated here.

[0206] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

[0207] Although the specific embodiments of the present application are described above in combination with the accompanying drawings, it is not a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications or changes made on the basis of the technical solutions of the present application without creative labor are still within the protection scope of the present application.

Claims

1. A charging load prediction method based on comprehensive similarity similar day screening, characterized in that, The method comprises the following steps: Similarity calculation is performed on the obtained historical load characteristics, meteorological characteristics and context characteristics, and normalization processing is performed; A multi-objective optimization model is constructed to minimize the mean and standard deviation of the fusion similarity score, and the fusion weight of each similarity is obtained by solving the model. The obtained historical load data similarity, meteorological data similarity and context characteristic similarity are fused, and a standard similar day set is constructed based on the fusion similarity value; The day pair features of the meteorological characteristics and context characteristics of the target day to be predicted and the candidate days are calculated as input, and the most matched similar day recognition set is selected by using the trained XGBoost model; the XGBoost model is trained by using the day pair features constructed based on the standard similar day set as training samples by using a supervised learning method; The meteorological characteristics and context characteristics of the target day to be predicted and the data of the selected similar day recognition set are input into a load prediction model to obtain a load prediction result. The calculation formula of the weighted score of the similarity score is as follows: (15); wherein: is the final weighted similarity score for the th historical day; is the three-dimensional weight vector to be optimized; is the three-dimensional normalized similarity score tuple for the th historical day; and are the weight and similarity score, respectively, for the th feature of the th category of historical load, historical load, and contextual feature, respectively. A multi-objective optimization model is constructed to minimize the mean and standard deviation of the fusion similarity score as follows: Objective One: Minimize the mean of the weighted scores : (16); wherein: is the total number of samples in the set of historical days; is the weighted score for the th historical day, denotes the sequence of weighted scores. Objective two: Minimize the standard deviation of the weighted scores : (17); The constraint condition of the multi-objective optimization model is a weight vector The standard simplex constraint is satisfied, and the formula is as follows: (18); A multi-objective optimization algorithm based on NSGA-II is used to solve the multi-objective optimization model, comprising the following steps: An initial population is constructed by randomly generating a weight vector; Iterative evolution is performed, including population breeding, evaluation and sorting, and environment selection; After iteration, the minimum standard deviation of the similarity weighted score is set as the optimal solution preference, and the optimal solution is selected from the Pareto frontier composed of multiple optimal solutions as the optimal weight vector, which is used as the fusion weight of each similarity; The similarities are fused based on the fusion weight of each similarity, and the similarity set with the highest similarity to the set reference day is obtained based on the fusion similarity score.

2. The charge load forecasting method based on the similar day screening according to the comprehensive similarity as claimed in claim 1, characterized in that: For the obtained time series data of historical load characteristics, a time series distance measurement method based on shape matching is used to calculate the historical load data similarity, comprising the following steps: For the obtained time series data of historical load characteristics, a local shape descriptor sequence is extracted for each point on each historical load curve; DTW similarity calculation is performed based on the obtained local shape descriptor sequence of the corresponding historical load characteristic sequence to obtain the ShapeDTW distance of any two historical load curves. 3.The charging load forecasting method based on the similar day screening according to the comprehensive similarity degree according to claim 1, characterized in that: For the meteorological characteristics, the time series data of historical meteorological characteristics are obtained, and a multi-dimensional time series distance measurement method is used to calculate the meteorological characteristic similarity, comprising the following steps: The Euclidean distance is used to calculate the distance between each data point in two meteorological characteristic sequences. Based on the obtained distance of the meteorological data points, a cumulative cost matrix is calculated to obtain the DTW distance of any two meteorological time series data. 4.The charging load forecasting method based on the similar day screening according to the comprehensive similarity degree according to claim 1, wherein: For the obtained historical day context characteristics, a binary difference method is used to calculate the absolute value of the difference of all time points of different days, and the sum of the absolute values of the differences is taken as the similarity of the context characteristics of different days. 5.The charging load forecasting method based on the similar day screening according to the comprehensive similarity degree according to claim 1, wherein: The obtained similarity values are normalized, comprising the following steps: The similarity values of the calculated historical load features, weather features and context features are scaled and transformed; The scaled and transformed similarity values are normalized based on the maximum value and the minimum value. 6.The charging load forecasting method based on the similar day screening according to the comprehensive similarity degree according to claim 1, wherein: The XGBoost model is trained by constructing day pair features of the similarity days as training samples, including the following steps: Set target day, compute target day Contextual similarity and weather feature similarity between the candidate historical day as day pair similarity features; It is judged whether the candidate historical day belongs to the similarity day set of the target day. If yes, it is constructed as a positive sample pair, otherwise as a negative sample pair. The positive sample pair and the negative sample pair are used as training samples; The XGBoost model is constructed, and the day pair features of the similarity of the weather features and the context features of the target day and the candidate day are extracted as inputs to train the XGBoost model with the matching relationship as the output.

7. The charging load forecasting system based on the comprehensive similarity similar day screening, characterized in that, It comprises: The similarity calculation module is configured to calculate the similarity of the obtained historical load features, weather features and context features respectively, and to normalize the processing; The optimal weight solving module is configured to minimize the mean and standard deviation of the fused similarity score as the target, construct a multi-objective optimization model, and solve the fused weight of each similarity. The historical load data similarity, weather data similarity and context feature similarity are fused, and a standard similarity day set is constructed based on the fused similarity value; The similarity day matching and screening module is configured to calculate the day pair features of the similarity of the weather features and the context features of the target day and the candidate day as inputs, and to screen out the most matched similarity day recognition set through the trained XGBoost model. The XGBoost model is trained by constructing day pair features of the standard similarity day set as training samples using a supervised learning method; The prediction module is configured to input the weather features and context features of the target day to be predicted and the data of the screened similarity day recognition set as input data, and to transmit them to the load prediction model for identification to obtain a load prediction result. The calculation formula of the weighted score of the similarity score is as follows: (15); wherein: is the final weighted similarity score for the th historical day; is the three-dimensional weight vector to be optimized; is the three-dimensional normalized similarity score tuple for the th historical day; and are the weight and similarity score, respectively, for the th category of features; are the categories of historical load, historical load, and contextual features, respectively; A multi-objective optimization model is constructed to minimize the mean and standard deviation of the fused similarity score as the target. Objective One: Minimize the mean of the weighted scores : (16); wherein: is the total number of samples in the set of historical days; is the weighted score for the th historical day, denotes the sequence of weighted scores; Objective two: Minimize the standard deviation of the weighted scores : (17); The constraint condition of the multi-objective optimization model is a weight vector The standard simplex constraint is satisfied, and the formula is as follows: (18); A multi-objective optimization algorithm based on NSGA-II is used to solve the multi-objective optimization model, including the following steps: Randomly generate a weight vector to construct an initial population; Iterative evolution is performed, including population breeding, evaluation and sorting, and environment selection; After iteration, the minimum standard deviation of the similarity weighted score is set as the optimal solution preference. From the Pareto frontier composed of multiple optimal solutions, the optimal solution is selected as the optimal weight vector as the fused weight of each similarity; The similarities are fused based on the fused weight of each similarity, and the similarity set with the highest similarity to the set reference day is obtained based on the fused similarity score.

8. The charge load forecasting system based on the similar day screening by integrated similarity according to claim 7, wherein: The XGBoost model is trained by constructing day pair features of the similarity days as training samples, including the following steps: Set target day, compute target day Contextual similarity and weather feature similarity between the candidate historical day as day pair similarity features; It is judged whether the candidate historical day belongs to the similarity day set of the target day. If yes, it is constructed as a positive sample pair, otherwise as a negative sample pair. The positive sample pair and the negative sample pair are used as training samples; The XGBoost model is constructed, the meteorological features and the context feature similarity of the target day and the candidate day are extracted as the input of the day pair features of the training sample data, and the XGBoost model is trained with the matching relationship as the output.

Citation Information

Patent Citations

  • Power system short-term load prediction method, device, equipment, medium and product

    CN119324451A

  • Expressway electric vehicle charging load prediction method and system

    CN119886459A