Method and system for mining data associated with external factors and charging behaviors of electric vehicle
By constructing a data set of electric vehicle charging behavior and external factors, and using a variety of data mining technologies for data expansion, quantification and weight calculation, the problem of low correlation analysis of electric vehicle charging behavior and external factors in the existing technology is solved, and effective identification and analysis of external factors affecting electric vehicle charging is achieved.
Patent Information
- Application Number
- CN202411848119.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to effectively analyze the relationship between electric vehicle charging behavior and external factors, resulting in the lack of data in the power system and low analysis efficiency.
By obtaining charging behavior factors and external factors, the original data set is constructed, and the ADASYN algorithm is used for data expansion, the DBSCAN clustering method is used for quantitative discrete, and the weight is calculated by combining the expert scoring method of LLM large language model simulation, and the Apriori algorithm is improved to mine the association relationship.
The correlation rules mining between external factors of electric vehicles and charging behavior is realized, the main external factors affecting electric vehicle charging are determined, and the efficiency and characteristic representativeness of data analysis are improved.
Smart Images

Figure CN119939284A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a means of associated data mining, belongs to the field of electric vehicles, and in particular to a method and system for mining data associated with external factors of electric vehicles and charging behaviors. Background Art
[0002] The global electric vehicle market is expanding at a rapid pace. This growth trend not only drives the urgent need for electric vehicle charging facilities, but also poses new challenges to the carrying capacity of the power grid. With the popularization of electric vehicles, the concentration of charging demand may lead to a surge in power grid load in specific areas, which in turn may cause problems such as power grid overload and charging congestion. In order to meet these challenges, we can provide solid support for sustainable transportation and energy systems through comprehensive analysis of user behavior, charging facility layout, and power dispatching strategies, thereby promoting the healthy development of the electric vehicle industry.
[0003] At present, the electric vehicle and new energy industries are in a rapid development stage. A large number of charging piles are being built as infrastructure to cope with the surge in charging demand. However, most of the existing research on electric vehicle travel focuses on the site selection and construction of charging piles, or focuses on modeling the travel behavior of electric vehicles or deducing charging demand using the travel chain method. Few works combine data mining technology to analyze the correlation between electric vehicle charging behavior and external factors, resulting in missing data in this part of the power system and incomplete data features, making data mining analysis difficult and inefficient. Summary of the invention
[0004] The purpose of the present invention is to overcome the above-mentioned defects and problems existing in the prior art, and to provide a data mining method and system for the association between external factors of electric vehicles and charging behaviors, which can make up for the missing data of the power system and realize data association mining simply and efficiently.
[0005] To achieve the above objectives, the technical solution of the present invention is: a data mining method for the association between external factors of electric vehicles and charging behavior, comprising:
[0006] S1. Obtain charging behavior factors and external factors and construct the original data set [L i , D i , M i , C i , W i , T i , F i , CT i};
[0007] The charging behavior factors include the charging behavior information of the electric vehicle; the external factors include the charging pile equipment information and the environmental status information;
[0008] The charging behavior information includes: whether charging F i and charging time CT i ; The charging pile equipment information includes: the location L of the charging pile i 、D i 、Charging Type M i 、Electricity Price C i ; The environmental status information includes weather conditions W i With temperature T i ;
[0009] S2, expand the original data set based on the ADASYN algorithm to obtain a new sample data set;
[0010] S3, labeling the dimension items with discrete values in the new sample data set, and quantizing and discretizing the dimension items with continuous values in the new sample data set based on the DBSCAN clustering method, and labeling them with discrete labels to obtain a sample label set;
[0011] S4, the expert scoring method based on the LLM large language model simulation calculates the weights of different dimension items in the sample label set;
[0012] S5. Improve the Apriori algorithm based on the weights of items in different dimensions, and explore the correlation between various charging behavior factors and external factors in the quantized discretized sample label set to determine the main external factors affecting electric vehicle charging.
[0013] The step S2 specifically includes:
[0014] S21, the original data set contains N original sample data, and the n original sample data are represented by ordered number pairs as [time i :{X1(i),…,X n (i)}]; where time i For the moment, X n (i) is the Nth item in the original data set, and i is the i-th group of original sample data in the original data set;
[0015] S22, for each time i , taking external factors as the characteristic dimension items of the original sample data {L i , D i , M i , C i , W i , T i}, with charging behavior item set {F i , CT i} as the category of each original sample data, and construct the category set C = {c1, c2, ..., c m}; where c m Class represents the minority class, and the total number of classes is 1+[(max{CT i} / 4)];
[0016] S23, count the number of samples in each category in the category set C, and obtain the number of samples in the Cth category as Nc; compare Nc with the threshold δ: if Nc<δ, it is marked as the minority class; if Nc>δ, it is marked as the majority class; count the maximum number of samples in each category N max The number and minimum N min the number of
[0017] S24. Calculate each minority class c m ∈S The total number of samples that need to be generated Its expression is as follows:
[0018]
[0019] in: is the number of existing samples of the minority class, β∈[0,1] is the ratio of the number of new samples generated, and S is the set of all minority classes;
[0020] S25. For all existing samples o in each minority class i ∈c m , calculate the scale factor of the i-th sample data and normalize it; o i is a sample data, expressed as [time i :{L i , D i , M i , C i , W i , T i , F i , CT i ];
[0021] The scaling factor r of the i-th sample data i The expression is as follows:
[0022] r i =Ω i / K,i=1,…,n;
[0023] Where: Ω is o i The number of K nearest neighbor samples belonging to the majority class;
[0024] The expression of the normalization process is as follows:
[0025]
[0026] in: Indicates o i The number of samples that need to be generated accounts for the minority class c m The total number of newly generated samples proportion;
[0027] S26, based on the proportional factor and the total number of samples to be generated Calculate i The number of additional samples that need to be generated nearby g i , which is expressed as follows:
[0028]
[0029] S27, for all minority classes c m All existing samples o i , o i ∈c m , c m ∈S, the i-th sample data o i Will generate g i samples, then based on the SMOTE algorithm, from o i Randomly select a minority class sample o of the same category from the K nearest neighbor samples ti , generate new augmented sample data o i ′, and o i ′ and o i Combine to obtain a new sample data set;
[0030] The expression for generating new expanded sample data is as follows:
[0031] o i ′=o i +(o ti -o i )×α;
[0032] Where: α is a scale parameter in the range [0, 1].
[0033] The step S3 specifically includes:
[0034] S31: In the new sample data set, the charging type M i 、Weather conditions i Whether charging F i The dimension items are discrete values, and their variables are labeled M i ={"m1","m2"},W i ={"w1","w2","w3"}, F i ={"f1","f2"}; Location of charging pile Li 、D i 、Electricity Price C i , Charging time CT i 、Temperature conditions T i If the dimension item of is a continuous value, DBSCAN clustering quantization is performed and labeled;
[0035] The steps of DBSCAN clustering quantification and labeling are as follows:
[0036] S32. Calculate the variable X of the nth dimension item n The continuous value range of is as follows:
[0037] Range({X n (i)})=max({X n (i)})-min({X n (i)});
[0038] Where: {X n (i)} is the nth dimension item X n The set of all sample points, max is {X n (i)}, min is the maximum value among {X n (i)}; X n The number of random samples is size({X n (i)});
[0039] S33, initialize the number of quasi-clusters K, set the radius parameter ∈ of the adaptive adjustment, and calculate the minimum point parameter Minpt, which is expressed as follows:
[0040] Minpt=size({X n (i)}) / K;
[0041] S34, based on the minimum point parameter Minpt, check all samples of the nth dimension item, calculate the number of points in the neighborhood of each sample point and compare it with the minimum point parameter Minpt;
[0042] If the number of points of the sample point is greater than or equal to Minpt, the sample point is a core point; if the number of points of the sample point is less than Minpt, the sample point is a non-core point;
[0043] For non-core points, if the sample point is in the neighborhood of the core point, the sample point is marked as a boundary point; if the sample point does not belong to the neighborhood of any core point, the sample point is marked as a noise point;
[0044] For a core point, if there are other core points in its neighborhood, these core points will be classified into the same cluster;
[0045] S35, repeat step S34, traverse all samples until no new clusters are generated, and then assign all samples of the dimension item to the category And label it x i ;
[0046] S36. Cluster and quantify all dimension items to obtain the location L of the charging pile i The labels {"l1", "l2", ..., "l n "}、D i The labels {″d1″, ″d2″, ..., ″d n ″}、Electricity price C i The labels {″c1″, ″c2″, ..., ″c n ″}、Charging time CT i The labels {"ct1", "ct2", ..., "ct n "}、Temperature condition T i The labels {″t1″, ″t2″, ..., ″t n ″}, and obtain the sample label set [time i : {″l i ″,″d i ″,″m i ″,″c i ″,″w i ″,″t i ″,″f i ″,″ct i ″}].
[0047] In step S33, the method for obtaining the adaptively adjusted radius parameter ∈ specifically includes:
[0048] S331, for the i-th sample point X of the n-th dimension item n (i), calculated to X n (i) The distance between the nearest K neighbors Its expression is as follows:
[0049]
[0050] in: For X n The kth nearest neighbor of (i);
[0051] S332, calculating the local density of the sample points, and determining the maximum local density of all sample points under n dimensions;
[0052] The expression of the local density is as follows:
[0053]
[0054] The expression of the local density maximum is as follows:
[0055] R m =max({R(X n (i)});
[0056] S333, preset the initial value of the radius parameter, and determine the n The radius parameter ∈ of (i) is as follows:
[0057]
[0058] Where: R m is the local maximum density, R(X n (i)) is the local density, and ∈0 is the initial value of the preset radius parameter.
[0059] The step S4 specifically includes:
[0060] S41, initialize J LLM individuals, where each LLM individual acts as an expert of an intelligent agent to play a dialogue game, and each round of the game includes a scoring phase, an explanation scoring phase, and an expert dialogue phase in sequence;
[0061] Scoring stage: J experts score the set of n different dimension items in turn; each expert scores each dimension item {″l i ″,″d i ″,″m i ″,″c i ″,″w i ″,″t i ″,″f i ″,″ct i ″} are scored independently, then the Nth expert output score is {s N1 ,s N2 , ..., s NJ} (1) ;
[0062] Explanation and scoring stage: J experts explain the reasons for their scoring in turn. The input for explaining the reasons for scoring is the expert's score in the scoring stage, and the output is the text explaining the reasons for scoring.
[0063] Expert dialogue stage: J experts state the reasons for their ratings in turn, and the output of the previous expert's statement is used as the input of the next expert's statement. After J experts complete a round of dialogue in turn, they score n different dimension items. The output score of the Nth expert is {s N1 ,s N2 , ..., s NJ} (2) ;
[0064] S42, after repeating step S41 for M rounds of game, the scores of all experts will be consistent, and the final score {s N1 ,s N2 , ..., s NJ} (M) As the weights of different dimension items {weight1, weight2, ..., weight N}, and output the sum of the weights
[0065] The step S5 specifically includes:
[0066] S51, input sample label set [time i :{"l i ","d i ","m i ","c i ","w i ","t i ", "f i ", "ct i "}], calculate the weighted support of each dimension item set, and its expression is as follows:
[0067]
[0068] Where: T is the transaction containing each item set, N is the total number of transactions, and X is the item set;
[0069] S52, find all single item sets of items of different dimensions, retain the 1-item set whose weighted support exceeds the minimum weighted support threshold τ, and use the K-1 item set to generate a new item set, and sequentially generate candidate 1 item set to candidate K item set as candidate sets;
[0070] S53, check whether each candidate K item set in the candidate set satisfies the minimum weighted support threshold τ; if it satisfies the threshold τ, it is classified as a frequent item set; if it does not satisfy the threshold τ, it is deleted from the candidate set;
[0071] S54. Based on frequent item sets, find item sets X composed of external factors A and electric vehicle charging behavior item set X B The association rules between items and the weighted confidence are calculated; the weighted confidence represents the item set X A Under the premise that the item set X B Probability of occurrence;
[0072] The calculation formula of the weighted confidence is as follows:
[0073]
[0074] S55. Based on the frequent item sets and weighted confidence, the weighted lift between the item sets of the external factor arrangement combinations and the item sets of the electric vehicle charging behavior is calculated;
[0075] The calculation formula of the weighted lift is as follows:
[0076]
[0077] If the value of weighted lift is greater than 1, it means that item set X A With item set X B There is a positive correlation between them;
[0078] If the value of weighted lift is equal to 1, it means that item set X A With item set X B There is an independent relationship between them;
[0079] If the value of weighted lift is less than 1, it means that item set X A With item set X B There is a negative correlation between them.
[0080] A data mining system for associating external factors of electric vehicles with charging behaviors, the system comprising:
[0081] The dataset construction module is used to obtain charging behavior factors and external factors and construct the original dataset [L i , D i , M i , C i , W i , T i , F i , CT i};
[0082] The charging behavior factors include the charging behavior information of the electric vehicle; the external factors include the charging pile equipment information and the environmental status information;
[0083] The charging behavior information includes: whether charging F i and charging time CT i ; The charging pile equipment information includes: the location L of the charging pile i 、D i 、Charging Type M i 、Electricity Price C i The environmental status information includes weather conditions W i With temperature T i ;
[0084] The data set expansion module is used to expand the original data set based on the ADASYN algorithm to obtain a new sample data set;
[0085] The data set quantization discretization module is used to label the dimension items with discrete values in the new sample data set, and quantize and discretize the dimension items with continuous values in the new sample data set based on the DBSCAN clustering method, and label them with discrete labels to obtain a sample label set;
[0086] The weight calculation module is used to calculate the weights of different dimensional items in the sample label set based on the expert scoring method simulated by the LLM large language model;
[0087] The association analysis module is used to improve the Apriori algorithm based on the weights of items in different dimensions, and to explore the association between various charging behavior factors and external factors in the quantized and discretized sample label set, and to determine the main external factors affecting electric vehicle charging.
[0088] The data set expansion module is used to expand the data according to the following steps:
[0089] S21, the original data set contains N original sample data, and the N original sample data are represented by ordered number pairs as [time i :{X1(i),…,X n (i)}]; where time i For the moment, X n (i) is the nth item in the original data set, and i is the i-th group of original sample data in the original data set;
[0090] S22, for each time i , taking external factors as the characteristic dimension items of the original sample data {L i , D i , M i , C i , W i , T i}, with charging behavior item set {F i , CT i} as the category of each original sample data, and construct the category set C = {c1, c2, ..., c m}; where c m Class represents the minority class, and the total number of classes is 1+[(max{CT i} / 4];
[0091] S23, count the number of samples in each category in the category set C, and obtain the number of samples in the cth category as Nc; compare Nc with the threshold δ: if Nc<δ, it is marked as the minority class; if Nc>δ, it is marked as the majority class; count the maximum number of samples in each category N max The number and minimum N min the number of
[0092] S24. Calculate each minority class c m ∈S The total number of samples that need to be generated Its expression is as follows:
[0093]
[0094] in: is the number of existing samples of the minority class, β∈[0,1] is the ratio of the number of new samples generated, and S is the set of all minority classes;
[0095] S25. For all existing samples o in each minority class i ∈c m , calculate the scale factor of the i-th sample data and normalize it; o i is a sample data, expressed as [time i :{L i , D i , M i , C i , W i , T i , F i , CT i ];
[0096] The scaling factor r of the i-th sample data i The expression is as follows:
[0097] r i =Ω i / K,i=1,…,n;
[0098] Where: Ω is o i The number of K nearest neighbor samples belonging to the majority class;
[0099] The expression of the normalization process is as follows:
[0100]
[0101] in: Indicates o i The number of samples that need to be generated accounts for the minority class c m The total number of newly generated samples proportion;
[0102] S26, based on the proportional factor and the total number of samples to be generated Calculate i The number of additional samples that need to be generated nearby g i , which is expressed as follows:
[0103]
[0104] S27, for all minority classes c m All existing samples o i , o i ∈c m , c m ∈S, the i-th sample data o i Will generate g i samples, then based on the SMOTE algorithm, from o i Randomly select a minority class sample o of the same category from the K nearest neighbor samples ti , generate new augmented sample data o i ′, and o i ′ and o i Combine to obtain a new sample data set;
[0105] The expression for generating new expanded sample data is as follows:
[0106] o i ′=o i +(o ti -o i )×α;
[0107] Where: α is a scale parameter in the range [0, 1].
[0108] The data set quantization discretization module is used to perform quantization discretization according to the following steps:
[0109] S31: In the new sample data set, the charging type M i 、Weather conditions i Whether charging F i The dimension items are discrete values, and their variables are labeled M i ={"m1","m2"},W i ={"w1","w2","w3"}, F i ={"f1","f2"}; Location of charging pile L i 、D i 、Electricity Price C i , Charging time CT i 、Temperature conditions T i If the dimension item of is a continuous value, DBSCAN clustering quantization is performed and labeled;
[0110] The steps of DBSCAN clustering quantification and labeling are as follows:
[0111] S32. Calculate the variable X of the nth dimension item n The continuous value range of is as follows:
[0112] Range({X n (i)})=max({X n (i)})-min({X n (i));
[0113] Where: {X n (i)} is the nth dimension item X n The set of all sample points, max is {X n (i)}, min is the maximum value among {X n (i) The minimum value among {; X n The number of random samples is size({X n (i)});
[0114] S33, initialize the number of quasi-clusters K, set the radius parameter ∈ of the adaptive adjustment, and calculate the minimum point parameter Minpt, which is expressed as follows:
[0115] Minpt=size({X n (i)}) / K;
[0116] S34, based on the minimum point parameter Minpt, check all samples of the nth dimension item, calculate the number of points in the neighborhood of each sample point and compare it with the minimum point parameter Minpt;
[0117] If the number of points of the sample point is greater than or equal to Minpt, the sample point is a core point; if the number of points of the sample point is less than Minpt, the sample point is a non-core point;
[0118] For non-core points, if the sample point is in the neighborhood of the core point, the sample point is marked as a boundary point; if the sample point does not belong to the neighborhood of any core point, the sample point is marked as a noise point;
[0119] For a core point, if there are other core points in its neighborhood, these core points will be classified into the same cluster;
[0120] S35, repeat step S34, traverse all samples until no new clusters are generated, and then assign all samples of the dimension item to the category And label it x i ;
[0121] S36. Cluster and quantify all dimension items to obtain the location L of the charging pile i The labels {″l1″, ″l2″, ..., ″l n ″}、Density around charging stations i The labels {″d1″, ″d2″, ..., ″d n ″}、Electricity price C iThe labels {″c1″, c2″, ..., ″c n ″}、Charging time CT i The labels {″ct1″, ″ct2″, ..., ″ct n ″}、Temperature condition T i The labels {″t1″, ″t2″, ..., ″t n ″}, and obtain the sample label set [time i : {″l i ″,″d i ″,″m i ″,″c i ″,″w i ″,″t i ″,″f i ″,″ct i ″}];
[0122] In step S33, the method for obtaining the adaptively adjusted radius parameter ∈ specifically includes:
[0123] S331, for the i-th sample point X of the n-th dimension item n (i), calculated to X n (i) The distance between the nearest K neighbors Its expression is as follows:
[0124]
[0125] in: For X n The kth nearest neighbor of (i);
[0126] S332, calculating the local density of the sample points, and determining the maximum local density of all sample points under n dimensions;
[0127] The expression of the local density is as follows:
[0128]
[0129] The expression of the local density maximum is as follows:
[0130] R m =max({R(X n (i)});
[0131] S333, preset the initial value of the radius parameter, and determine the n The radius parameter ∈ of (i) is as follows:
[0132]
[0133] Where: R mis the local maximum density, R(X n (i)) is the local density, and ∈0 is the initial value of the preset radius parameter.
[0134] The weight calculation module is used to calculate the weight according to the following steps:
[0135] S41, initialize J LLM individuals, where each LLM individual acts as an expert of an intelligent agent to play a dialogue game, and each round of the game includes a scoring phase, an explanation scoring phase, and an expert dialogue phase in sequence;
[0136] Scoring stage: J experts score the set of n different dimension items in turn; each expert scores each dimension item {″l i ″,″d i ″,″m i ″,″c i ″,″w i ″,″t i ″,″f i ″,″ct i ″} are scored independently, then the Nth expert output score is {s N1 ,s N2 , ..., s NJ} (1) ;
[0137] Explanation and scoring stage: J experts explain the reasons for their scoring in turn. The input for explaining the reasons for scoring is the expert's score in the scoring stage, and the output is the text explaining the reasons for scoring.
[0138] Expert dialogue stage: J experts state the reasons for their ratings in turn, and the output of the previous expert's statement is used as the input of the next expert's statement. After J experts complete a round of dialogue in turn, they score n different dimension items. The output score of the Nth expert is {s N1 ,s N2 , ..., s NJ} (2) ;
[0139] S42, after repeating step S41 for M rounds of game, the scores of all experts will be consistent, and the final score {s N1 ,s N2 , ..., s NJ} (M) As the weights of different dimension items {weight1, weight2, ..., weight N}, and output the sum of the weights
[0140] The association analysis module is used to perform association analysis on charging behavior factors and external factors according to the following steps:
[0141] S51, input sample label set [time i : {″l i ″,″d i ″,″m i ″,″c i ″,″w i ″,″t i ″,″f i ″,″ct i ″}], calculate the weighted support of each dimension item set, and its expression is as follows:
[0142]
[0143] Where: T is the transaction containing each item set, N is the total number of transactions, and X is the item set;
[0144] S52, find all single item sets of items of different dimensions, retain the 1-item set whose weighted support exceeds the minimum weighted support threshold τ, and use the K-1 item set to generate a new item set, and sequentially generate candidate 1 item set to candidate K item set as candidate sets;
[0145] S53, check whether each candidate K item set in the candidate set satisfies the minimum weighted support threshold τ; if it satisfies the threshold τ, it is classified as a frequent item set; if it does not satisfy the threshold τ, it is deleted from the candidate set;
[0146] S54. Based on frequent item sets, find item sets X composed of external factors A and electric vehicle charging behavior item set X B The association rules between items and the weighted confidence are calculated; the weighted confidence represents the item set X A Under the premise that the item set X B Probability of occurrence;
[0147] The calculation formula of the weighted confidence is as follows:
[0148]
[0149] S55. Based on the frequent item sets and weighted confidence, the weighted lift between the item sets of the external factor arrangement combinations and the item sets of the electric vehicle charging behavior is calculated;
[0150] The calculation formula of the weighted lift is as follows:
[0151]
[0152] If the value of weighted lift is greater than 1, it means that item set X A With item set X B There is a positive correlation between them;
[0153] If the value of weighted lift is equal to 1, it means that item set X A With item set X B There is an independent relationship between them;
[0154] If the value of weighted lift is less than 1, it means that item set X A With item set X B There is a negative correlation between them.
[0155] Compared with the prior art, the present invention has the following beneficial effects:
[0156] The invention discloses a data mining method and system for the association between external factors and charging behavior of electric vehicles. The method first obtains charging behavior factors and external factors, constructs an original data set and performs data expansion, then quantizes and discretizes the continuous values therein, and marks them with discrete labels. Then, an expert scoring method based on LLM large language model simulation is used to calculate weights, and an Apriori algorithm is improved. Finally, the association relationship of various factors in the sample label set after quantization and discretization is mined to determine the main external factors affecting the charging of electric vehicles. In the application of the design, the ADASYN algorithm is used to perform oversampling data expansion on the power system data, which makes up for the shortcomings of many missing power system data and incomplete features. Then, the DBSCAN clustering method is used to quantize the feature items of continuous value. Compared with the existing uniform quantization and frequency quantization technologies, the quantization results are more representative of the features and the difficulty of analyzing the original data is reduced. Finally, data mining is realized based on the improved Apriori algorithm. Compared with the original Apriori method, the weights of different influencing factors are changed, and the association rules between external factors and the charging behavior of electric vehicles are realized simply and efficiently, which is of great significance for guiding the power dispatching decision of the power system. BRIEF DESCRIPTION OF THE DRAWINGS
[0157] Figure 1 It is a flow chart of the method of the present invention.
[0158] Figure 2 It is a system structure diagram of the present invention.
[0159] Figure 3 It is a device structure diagram of the present invention.
[0160] In the figure: data set construction module 1, data set expansion module 2, data set quantization and discretization module 3, weight calculation module 4, association analysis module 5, processor 6, memory 7, computer program code 71. DETAILED DESCRIPTION
[0161] The present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0162] Embodiment 1:
[0163] See also Figure 1 , a data mining method for associating external factors of electric vehicles with charging behaviors, comprising:
[0164] S1. Obtain charging behavior factors and external factors and construct the original data set [L i , D i , M i , C i , W i , T i , F i , CT i};
[0165] The charging behavior factors include the charging behavior information of the electric vehicle; the external factors include the charging pile equipment information and the environmental status information;
[0166] The charging behavior information includes: whether charging F i and charging time CT i ; The charging pile equipment information includes: the location L of the charging pile i 、D i 、Charging Type M i (fast charging, slow charging, etc.), electricity price C i The environmental status information includes weather conditions W i (Sunny, cloudy, light rain, etc.) and temperature conditions T i ;
[0167] In this embodiment, the power system collects charging data of electric vehicles at 15-minute intervals, and each charging pile in the system collects external environmental information such as its own equipment status and environmental status information. Then, the ADASYN algorithm is used to expand the existing samples to solve the problems of missing samples and insufficient representativeness of the model. The details are as follows:
[0168] S2, expand the original data set based on the ADASYN algorithm to obtain a new sample data set;
[0169] Furthermore, the step S2 specifically includes:
[0170] S21, the original data set contains N original sample data, and the N original sample data are represented by ordered number pairs as [time i :{X1(i),…,X n (i)}]; where time i is the moment, X n (i) is the nth item in the original data set, and i is the i-th group of original sample data in the original data set;
[0171] S22, for each time i , taking external factors as the characteristic dimension items of the original sample data {L i , D i , M i , C i , W i , T i}, with charging behavior item set {F i , CT i} as the category of each original sample data, and construct the category set C = {c1, c2, ..., c m}; where c m Class represents the minority class, and the total number of classes is 1+[(max{CT i} / 4)]; where charging time is divided into 4 hours per category.
[0172] S23, count the number of samples in each category in the category set C, and obtain the number of samples in the cth category as Nc; compare Nc with the threshold δ: if Nc<δ, it is marked as the minority class; if Nc>δ, it is marked as the majority class; count the maximum number of samples in each category N max The number and minimum N min ; except for the minority class, the other classes exceeding the threshold δ are the majority class.
[0173] S24. Calculate each minority class c m ∈S The total number of samples that need to be generated Its expression is as follows:
[0174]
[0175] in: is the number of existing samples of the minority class; β∈[0,1] is the ratio of the number of new samples generated, which can be generated according to actual needs. The larger the value of β, the more new samples are added; S is the set of all minority classes;
[0176] S25. For all existing samples o in each minority class i ∈c m , calculate the scale factor of the i-th sample data and normalize it; o i is a sample data, expressed as [time i :{L i , D i , M i , C i , W i , T i , F i , CT i ];
[0177] The scaling factor r of the i-th sample data i The expression is as follows:
[0178] r i =Ω i / K,i=1,…,n;
[0179] Where: Ω is o i The number of K nearest neighbor samples belonging to the majority class;
[0180] The expression of the normalization process is as follows:
[0181]
[0182] in: Indicates o i The number of samples that need to be generated accounts for the minority class c m The total number of newly generated samples proportion;
[0183] S26, based on the proportional factor and the total number of samples to be generated Calculate i The number of additional samples that need to be generated nearby g i , which is expressed as follows:
[0184]
[0185] S27, for all minority classes c m All existing samples o i , o i ∈c m , c m ∈S, the i-th sample data o i Will generate g i samples, then based on the SMOTE algorithm, from o i Randomly select a minority class sample o of the same category from the K nearest neighbor samples ti , generate new augmented sample data o i ′, and o i ′ and o i Combine to obtain a new sample data set;
[0186] The expression for generating new expanded sample data is as follows:
[0187] o i ′=o i +(o ti -o i )×α;
[0188] Where: α is a scale parameter in the range [0, 1].
[0189] In this embodiment, after obtaining the expanded samples, it is necessary to quantify the external environment information and the electric vehicle charging information, discretize the continuous value variables, reduce the complexity of the model analysis, and use the DBSCAN clustering algorithm to determine the category to which the continuous value feature dimension item value belongs, and mark it with a discrete label, as follows:
[0190] S3, labeling the dimension items with discrete values in the new sample data set, and quantizing and discretizing the dimension items with continuous values in the new sample data set based on the DBSCAN clustering method, and labeling them with discrete labels to obtain a sample label set;
[0191] Furthermore, the step S3 specifically includes:
[0192] S31: In the new sample data set, the charging type M i 、Weather conditions i Whether charging F i The dimension items are discrete values, and no quantization is required. Then the variables are labeled M respectively. i ={"m1","m2"},W i ={"w1","w2","w3"}, F i ={"f1","f2"}; Location of charging pile L i 、D i 、Electricity Price C i , Charging time CT i 、Temperature conditions T i If the dimension items are continuous values, the DBSCAN clustering algorithm is used to obtain the classification of each group of variables, quantify them and label them, and each value of the same category will be labeled with the same label.
[0193] The steps of DBSCAN clustering quantification and labeling are as follows:
[0194] S32. Calculate the variable X of the nth dimension item n The continuous value range of is as follows:
[0195] Range({X n (i)})=max({X n (i)})-min({X n (i)))
[0196] Where: {X n (i)} is the nth dimension item X n The set of all sample points, max is {X n (i)}, min is the maximum value among {Xn (i)}; X n The number of random samples is size({X n (i)});
[0197] S33, initialize the number of quasi-clusters K, set the radius parameter ∈ of the adaptive adjustment, and calculate the minimum point parameter Minpt, which is expressed as follows:
[0198] Minpt=size({X n (i)}) / K;
[0199] S34, based on the minimum point parameter Minpt, check all samples of the nth dimension item, calculate the number of points in the neighborhood of each sample point and compare it with the minimum point parameter Minpt;
[0200] If the number of points of the sample point is greater than or equal to Minpt, the sample point is a core point; if the number of points of the sample point is less than Minpt, the sample point is a non-core point;
[0201] For non-core points, if the sample point is in the neighborhood of the core point, the sample point is marked as a boundary point; if the sample point does not belong to the neighborhood of any core point, the sample point is marked as a noise point;
[0202] For a core point, if there are other core points in its neighborhood, these core points will be classified into the same cluster;
[0203] S35, repeat step S34, traverse all samples until no new clusters are generated, and then assign all samples of the dimension item to the category And label it x i ;
[0204] S36. Cluster and quantify all dimension items to obtain the location L of the charging pile i The labels {″l1″, ″l2″, ..., ″l n ″}、Density around charging stations i The labels {″d1″, ″d2″, ..., ″d n ″}、Electricity price C i The labels {″c1″, ″c2″, ..., ″c n ″}、Charging time CT i The labels {″ct1″, ″ct2″, ..., ″ct n ″}、Temperature condition T i The labels {"t1", "t2", ..., "t n "}, and obtain the sample label set [time i :{"l i ","di ","m i ","c i ","w i ","t i ", "f i ", "ct i "}.
[0205] Furthermore, in step S33, the method for obtaining the adaptively adjusted radius parameter ∈ specifically includes:
[0206] S331, for the i-th sample point X of the n-th dimension item n (i), calculated to X n (i) The distance between the nearest K neighbors Its expression is as follows:
[0207]
[0208] in: For X n The kth nearest neighbor of (i);
[0209] S332, calculating the local density of the sample points, and determining the maximum local density of all sample points under n dimensions;
[0210] The expression of the local density is as follows:
[0211]
[0212] The expression of the local density maximum is as follows:
[0213] R m =max({R(X n (i)});
[0214] S333, preset the initial value of the radius parameter, and determine the n The radius parameter ∈ of (i) is as follows:
[0215]
[0216] Where: R m is the local maximum density, R(X n (i)) is the local density, and ∈0 is the initial value of the preset radius parameter.
[0217] After clustering quantization, all values become discrete label values. Then, the improved Apriori algorithm is used to mine the associations between various types of discretized external environmental information and the charging behavior of electric vehicles, find the associations, and obtain the main external factors that affect the charging behavior of electric vehicles at the current moment; the details are as follows:
[0218] S4, the expert scoring method based on the LLM large language model simulation calculates the weights of different dimension items in the sample label set;
[0219] Furthermore, the step S4 specifically includes:
[0220] The LLM large language model is an artificial intelligence model based on natural language processing technology. It is trained on massive texts and can complete language interactions between the model and humans, and between models. During the interaction process, language and text are input, and then the model outputs language and text after calculation.
[0221] S41, initialize J LLM individuals, where each LLM individual acts as an expert of an intelligent agent to play a dialogue game, and each round of the game includes a scoring phase, an explanation scoring phase, and an expert dialogue phase in sequence;
[0222] Scoring stage: J experts score the set of n different dimension items in turn; each expert scores each dimension item {″l i ″,″d i ″,″m i ″,″c i ″,″w i ″,″t i ″,″f i ″,″ct i ″} are scored independently, then the Nth expert output score is {s N1 ,s N2 , ..., s NJ} (1) ;
[0223] Explanation and scoring stage: J experts explain the reasons for their scores in turn. The input for explaining the reasons for the scores is the scores given by the experts in the scoring stage, and the output is the text explaining the reasons for the scores. Each expert is limited to one sentence.
[0224] Expert dialogue stage: J experts state the reasons for their ratings in turn, and the output of the previous expert's statement is used as the input of the next expert's statement. After J experts complete a round of dialogue in turn, they score n different dimension items. The output score of the Nth expert is {s N1 ,s N2 , ..., s NJ} (2) ;
[0225] S42, after repeating step S41 for M rounds of game, the scores of all experts will be consistent, and the final score {s N1 ,s N2 , ..., s NJ}(M) As the weights of different dimension items {weight1, weight2, ..., weight N}, and output the sum of the weights
[0226] In this embodiment, several conversations are conducted. In each round, J experts speak one paragraph in turn. The text spoken by the previous expert will serve as the input of the next few LLM experts, triggering the text language output by the next expert. After all N LLM experts have finished speaking in each conversation, the next conversation will be conducted. The conversation between them will affect the score of each agent. After the end of this round of conversation, J experts will score n different weighted items. After the above steps, the next round of game will be conducted. After M rounds of game, the scores of all agents will tend to be consistent, and the final score of the LLM agent will be used as the weight of the different dimensional items. This solution adopts the conversation method of the LLM large language model to generate weights to distinguish and replace the general expert scoring method, improve the consistency of the results, and make the results more consistent and reliable.
[0227] S5. Improve the Apriori algorithm based on the weights of items in different dimensions, and explore the correlation between various charging behavior factors and external factors in the quantized discretized sample label set to determine the main external factors affecting electric vehicle charging.
[0228] Furthermore, the step S5 specifically includes:
[0229] S51, input sample label set [time i : {″l i ″,″d i ″,″m i ″,″c i ″,″w i ″,″t i ″,″f i ″,″ct i ″}], calculate the weighted support of each dimension item set, and its expression is as follows:
[0230]
[0231] Where: T is the transaction containing each item set, N is the total number of transactions, and X is the item set;
[0232] S52, find all single item sets of items of different dimensions, retain the 1-item set whose weighted support exceeds the minimum weighted support threshold τ, and use the K-1 item set to generate a new item set, and sequentially generate candidate 1 item set to candidate K item set as candidate sets;
[0233] S53, check whether each candidate K item set in the candidate set satisfies the minimum weighted support threshold τ; if it satisfies the threshold τ, it is classified as a frequent item set; if it does not satisfy the threshold τ, it is deleted from the candidate set;
[0234] S54. Based on frequent item sets, find item sets X composed of external factors A and electric vehicle charging behavior item set X B The association rules between items and the weighted confidence are calculated; the weighted confidence represents the item set X A Under the premise that the item set X B Probability of occurrence;
[0235] The calculation formula of the weighted confidence is as follows:
[0236]
[0237] S55. Based on the frequent item sets and weighted confidence, calculate the weighted lift between the item sets of the external factor arrangement and combination and the item sets of the electric vehicle charging behavior;
[0238] The calculation formula of the weighted lift is as follows:
[0239]
[0240] If the value of weighted lift is greater than 1, it means that item set X A With item set X B There is a positive correlation between them;
[0241] If the value of weighted lift is equal to 1, it means that item set X A With item set X B There is an independent relationship between them;
[0242] If the value of weighted lift is less than 1, it means that item set X A With item set X B There is a negative correlation between them.
[0243] In this embodiment, when there is a positive correlation between the external factor and the charging behavior, it means that as the external factor increases, the frequency or intensity of the charging behavior will also increase, or as the external factor decreases, the frequency or intensity of the charging behavior will also decrease. For example, if the weather gets colder and the battery performance decreases, the user may need to charge more frequently, which is manifested as a positive correlation between the external factor (temperature) and the charging behavior (charging frequency).
[0244] When there is an independent relationship between external factors and charging behavior, it means that changes in external factors have no effect on charging behavior, or the effect is so small that it can be ignored. For example, if users are always used to charging at the same time every day, their charging behavior may be unrelated to external factors such as weather or traffic conditions.
[0245] When there is a negative correlation between external factors and charging behavior, it means that as the external factors increase, the frequency or intensity of charging behavior decreases, or as the external factors decrease, the frequency or intensity of charging behavior increases. For example, if the electricity price is higher during peak hours, users may choose to charge during off-peak hours when the electricity price is lower, which is manifested as a negative correlation between external factors (electricity prices) and charging behavior (charging time choice).
[0246] By analyzing and exploring the relationship between external factors of electric vehicles and charging behavior, it is helpful to better predict and optimize the charging behavior of electric vehicles, thereby improving charging efficiency and the stability of the power grid.
[0247] Embodiment 2:
[0248] See also Figure 2 , a data mining system for associating external factors of electric vehicles with charging behaviors, the system comprising:
[0249] Dataset construction module 1 is used to obtain charging behavior factors and external factors and construct the original dataset [L i , D i , M i , C i , W i , T i , F i , CT i};
[0250] The charging behavior factors include the charging behavior information of the electric vehicle; the external factors include the charging pile equipment information and the environmental status information;
[0251] The charging behavior information includes: whether charging F i and charging time CT i ; The charging pile equipment information includes: the location L of the charging pile i 、D i 、Charging Type M i 、Electricity Price C i The environmental status information includes weather conditions W i With temperature T i ;
[0252] The data set expansion module 2 is used to expand the original data set based on the ADASYN algorithm to obtain a new sample data set;
[0253] Furthermore, the data set expansion module 2 is used to expand the data according to the following steps:
[0254] S21, the original data set contains N original sample data, and the N original sample data are represented by ordered number pairs as [time i :{X1(i),…,X n (i)}]; where time i is the moment, X n (i) is the nth item in the original data set, and i is the i-th group of original sample data in the original data set;
[0255] S22, for each time i , taking external factors as the characteristic dimension items of the original sample data {L i , D i , M i , C i , W i , T i}, with charging behavior item set {F i , CT i} as the category of each original sample data, and construct the category set C = {c1, c2, ..., c m}; where c m Class represents the minority class, and the total number of classes is 1+[(max{CT i} / 4)];
[0256] S23. Count the number of samples in each category in the category set C, and obtain the number of samples in the cth category as Nc; compare Nc with the threshold δ: if Nc<δ, it is marked as the minority class; if Nc>δ, it is marked as the majority class; count the maximum number of samples in each category N max The number and minimum N min the number of
[0257] S24. Calculate each minority class c m ∈S The total number of samples that need to be generated Its expression is as follows:
[0258]
[0259] in: is the number of existing samples of the minority class, β∈[0,1] is the ratio of the number of new samples generated, and S is the set of all minority classes;
[0260] S25. For all existing samples o in each minority class i ∈c m , calculate the scale factor of the i-th sample data and normalize it; o i is a sample data, expressed as [time i :{L i , D i , M i , C i , W i , T i , F i , CT i ];
[0261] The scaling factor r of the i-th sample data i The expression is as follows:
[0262] r i =Ω i / K,i=1,…n;
[0263] Where: Ω is o i The number of K nearest neighbor samples belonging to the majority class;
[0264] The expression of the normalization process is as follows:
[0265]
[0266] in: Indicates o i The number of samples that need to be generated accounts for the minority class c m The total number of newly generated samples proportion;
[0267] S26, based on the proportional factor and the total number of samples to be generated Calculate i The number of additional samples that need to be generated nearby g i , which is expressed as follows:
[0268]
[0269] S27, for all minority classes c m All existing samples o i , o i ∈c m , c m ∈S, the i-th sample data o i Will generate g i samples, then based on the SMOTE algorithm, from o i Randomly select a minority class sample o of the same category from the K nearest neighbor samples ti, generate new augmented sample data o i ′, and o i ′ and o i Combine to obtain a new sample data set;
[0270] The expression for generating new expanded sample data is as follows:
[0271] o i ′=o i +(o ti -o i )×α;
[0272] Where: α is a scale parameter in the range [0, 1].
[0273] The data set quantization discretization module 3 is used to label the dimension items with discrete values in the new sample data set, and quantize and discretize the dimension items with continuous values in the new sample data set based on the DBSCAN clustering method, and label them with discrete labels to obtain a sample label set;
[0274] Furthermore, the data set quantization discretization module 3 is used to perform quantization discretization according to the following steps:
[0275] S31: In the new sample data set, the charging type M i 、Weather conditions i Whether charging F i The dimension items are discrete values, and their variables are labeled M i ={"m1","m2"},W i ={"w1","w2","w3"}, F i ={"f1","f2"}; Location of charging pile L i 、D i 、Electricity Price C i , Charging time CT i 、Temperature conditions T i If the dimension item of is a continuous value, DBSCAN clustering quantization is performed and labeled;
[0276] The steps of DBSCAN clustering quantification and labeling are as follows:
[0277] S32. Calculate the variable X of the nth dimension item n The continuous value range of is as follows:
[0278] Range({X n (i)})=max({X n (i)})-min({X n (i)});
[0279] Where: {X n (i)} is the nth dimension item X n The set of all sample points, max is {X n (i)}, min is the maximum value among {X n (i)}; X n The number of random samples is size({X n (i)});
[0280] S33, initialize the number of quasi-clusters K, set the radius parameter ∈ of the adaptive adjustment, and calculate the minimum point parameter Minpt, which is expressed as follows:
[0281] Minpt=sizee({X n (i)}) / K;
[0282] S34, based on the minimum point parameter Minpt, check all samples of the nth dimension item, calculate the number of points in the neighborhood of each sample point and compare it with the minimum point parameter Minpt;
[0283] If the number of points of the sample point is greater than or equal to Minpt, the sample point is a core point; if the number of points of the sample point is less than Minpt, the sample point is a non-core point;
[0284] For non-core points, if the sample point is in the neighborhood of the core point, the sample point is marked as a boundary point; if the sample point does not belong to the neighborhood of any core point, the sample point is marked as a noise point;
[0285] For a core point, if there are other core points in its neighborhood, these core points will be classified into the same cluster;
[0286] S35, repeat step S34, traverse all samples until no new clusters are generated, and then assign all samples of the dimension item to the category And label it x i ;
[0287] S36. Cluster and quantify all dimension items to obtain the location L of the charging pile i The labels {″l1″, ″l2″, ..., ″l n ″}、Density around charging stations i The labels {″d1″, ″d2″, ..., ″d n ″}、Electricity price C i The labels {″c1″, ″c2″, ..., ″c n ″}、Charging time CT i The labels {″ct1″, ″ct2″, ..., ″ct n″}、Temperature condition T i The labels {″t1″, ″t2″, ..., ″t n ″}, and obtain the sample label set [time i : {″l i ″,″d i ″,″m i ″,″c i ″,″w i ″,″t i ″,″f i ″,″ct i ″}];
[0288] In step S33, the method for obtaining the adaptively adjusted radius parameter ∈ specifically includes:
[0289] S331, for the i-th sample point X of the n-th dimension item n (i), calculated to X n (i) The distance between the nearest K neighbors Its expression is as follows:
[0290]
[0291] in: For X n The kth nearest neighbor of (i);
[0292] S332, calculating the local density of the sample points, and determining the maximum local density of all sample points under n dimensions;
[0293] The expression of the local density is as follows:
[0294]
[0295] The expression of the local density maximum is as follows:
[0296] R m =max({R(X n (i)});
[0297] S333, preset the initial value of the radius parameter, and determine the n The radius parameter ∈ of (i) is as follows:
[0298]
[0299] Where: R m is the local maximum density, R(X n (i)) is the local density, and ∈0 is the initial value of the preset radius parameter.
[0300] The weight calculation module 4 is used to calculate the weights of different dimension items in the sample label set based on the expert scoring method simulated by the LLM large language model;
[0301] Furthermore, the weight calculation module 4 is used to calculate the weight according to the following steps:
[0302] S41, initialize J LLM individuals, where each LLM individual acts as an expert of an intelligent agent to play a dialogue game, and each round of the game includes a scoring phase, an explanation scoring phase, and an expert dialogue phase in sequence;
[0303] Scoring stage: J experts score the set of n different dimension items in turn; each expert scores each dimension item {″l i ″,″d i ″,″m i ″,″c i ″,″w i ″,″t i ″,″f i ″,″ct i ″} are scored independently, then the Nth expert output score is {s N1 ,s N2 , ..., s NJ} (1) ;
[0304] Explanation and scoring stage: H experts explain the reasons for their scoring in turn. The input for explaining the reasons for scoring is the expert's score in the scoring stage, and the output is the text explaining the reasons for scoring.
[0305] Expert dialogue stage: H experts state the reasons for their ratings in turn. The output of the previous expert's statement is used as the input of the next expert's statement. After H experts complete a round of dialogue in turn, they score n different dimension items. The output score of the Nth expert is {s N1 , S N2 , ..., s NJ} (2) ;
[0306] S42, after repeating step S41 for M rounds of game, the scores of all experts will be consistent, and the final score {s N1 , S N2 , ..., s NJ} (M) As the weights of different dimension items {weight1, weight2, ..., weight N}, and output the sum of the weights
[0307] Association analysis module 5 is used to improve the Apriori algorithm based on the weights of different dimensional items, and to mine the association between various charging behavior factors and external factors in the quantized discretized sample label set, and determine the main external factors affecting electric vehicle charging;
[0308] Furthermore, the correlation analysis module 5 is used to perform correlation analysis on the charging behavior factors and the external factors according to the following steps:
[0309] S51, input sample label set [time i : {″l i ″,″d i ″,″m i ″,″c i ″,″w i ″,″t i ″,″f i ″,″ct i ″}], calculate the weighted support of each dimension item set, and its expression is as follows:
[0310]
[0311] Where: T is the transaction containing each item set, N is the total number of transactions, and X is the item set;
[0312] S52, find all single item sets of items of different dimensions, retain the 1-item set whose weighted support exceeds the minimum weighted support threshold τ, and use the K-1 item set to generate a new item set, and sequentially generate candidate 1 item set to candidate K item set as candidate sets;
[0313] S53, check whether each candidate K item set in the candidate set satisfies the minimum weighted support threshold τ; if it satisfies the threshold τ, it is classified as a frequent item set; if it does not satisfy the threshold τ, it is deleted from the candidate set;
[0314] S54. Based on frequent item sets, find item sets X composed of external factors A and electric vehicle charging behavior item set X B The association rules between items and the weighted confidence are calculated; the weighted confidence represents the item set X A Under the premise that the item set X B Probability of occurrence;
[0315] The calculation formula of the weighted confidence is as follows:
[0316]
[0317] S55. Based on the frequent item sets and weighted confidence, calculate the weighted lift between the item sets of the external factor arrangement and combination and the item sets of the electric vehicle charging behavior;
[0318] The calculation formula of the weighted lift is as follows:
[0319]
[0320] If the value of weighted lift is greater than 1, it means that item set X A With item set X B There is a positive correlation between them;
[0321] If the value of weighted lift is equal to 1, it means that item set X A With item set X B There is an independent relationship between them;
[0322] If the value of weighted lift is less than 1, it means that item set X A With item set X B There is a negative correlation between them.
[0323] Embodiment 3:
[0324] See also Figure 3 , a data mining device for association between external factors of electric vehicles and charging behaviors, the device comprising a processor 6 and a memory 7;
[0325] The memory 7 is used to store computer program code 71 and transmit the computer program code 71 to the processor 6;
[0326] The processor 6 is used to execute the data mining method for associating external factors of electric vehicles with charging behaviors described in Example 1 according to the instructions in the computer program code 71.
[0327] In this embodiment, a computer-readable storage medium is also included, in which computer-executable instructions are stored. When the computer-executable instructions are executed on a computer, the data mining method for associating external factors of electric vehicles with charging behaviors described in Example 1 is implemented.
[0328] Generally speaking, the computer instructions for implementing the method of the present invention may be carried in any combination of one or more computer-readable storage media. Non-transitory computer-readable storage media may include any computer-readable media, except for the signal itself that is temporarily propagating.
[0329] The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EKROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, device, or device.
[0330] Computer program code for performing the operation of the present invention can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, SMalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages, in particular, Python suitable for neural network computing and platform frameworks based on TensorFlow, PyTorch, etc. can be used. The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer or to an external computer (for example, using an Internet service provider to connect via the Internet) through any type of network, including a local area network (LAN) or a wide area network (WAN).
[0331] The above-mentioned device and non-temporary computer-readable storage medium can refer to the specific description of a data mining method and beneficial effects associated with external factors and charging behavior of electric vehicles, which will not be repeated here.
[0332] Although the embodiments of the present invention have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.
Claims
1. A data mining method for the association between external factors and charging behavior of electric vehicles, characterized in that: include: S1. Obtain charging behavior factors and external factors and construct the original data set [L i , D i , M i , C i , W i , T i , F i , CT i }; The charging behavior factor includes the charging behavior information of the electric vehicle; The external factors include charging pile equipment information and environmental status information; The charging behavior information includes: whether charging F i and charging time CT i ; The charging pile equipment information includes: the location L of the charging pile i 、D i 、Charging Type M i 、Electricity Price C i ; The environmental status information includes weather conditions W i With temperature T i ; S2, expand the original data set based on the ADASYN algorithm to obtain a new sample data set; S3, labeling the dimension items with discrete values in the new sample data set, and quantizing and discretizing the dimension items with continuous values in the new sample data set based on the DBSCAN clustering method, and labeling them with discrete labels to obtain a sample label set; S4, the expert scoring method based on the LLM large language model simulation calculates the weights of different dimension items in the sample label set; S5. Improve the Apriori algorithm based on the weights of items in different dimensions, and explore the correlation between various charging behavior factors and external factors in the quantized discretized sample label set to determine the main external factors affecting electric vehicle charging.
2. The data mining method for associating external factors of electric vehicles with charging behaviors according to claim 1 is characterized by: The step S2 specifically includes: S21, the original data set contains N original sample data, and the N original sample data are represented by ordered number pairs as [time i :{X1(i),…,X n (i)}]; where time i is the moment, X n (i) is the nth item in the original data set, and i is the i-th group of original sample data in the original data set; S22, for each time i , taking external factors as the characteristic dimension items of the original sample data {L i , D i , M i , C i , W i , T i }, with charging behavior item set {F i , CT i } as the category of each original sample data, and construct the category set C = {c1, c2, ..., c m }; where c m Class represents the minority class, and the total number of classes is 1+[(max{CT i } / 4)]; S23, count the number of samples in each category in the category set C, and obtain the number of samples in the cth category as Nc; compare Nc with the threshold δ: if Nc<δ, it is marked as the minority class; if Nc>δ, it is marked as the majority class; count the maximum number of samples in each category N max The number and minimum N min the number of S24. Calculate each minority class c m ∈S The total number of samples that need to be generated Its expression is as follows: in: is the number of existing samples of the minority class, β∈[0,1] is the ratio of the number of new samples generated, and S is the set of all minority classes; S25. For all existing samples o in each minority class i ∈c m , calculate the scale factor of the i-th sample data and normalize it; o i is a sample data, expressed as [time i :{L i , D i , M i , C i , W i , T i , F i , CT i ]; The scaling factor r of the i-th sample data i The expression is as follows: r i =Ω i / K,i=1,…,n; Where: Ω is o i The number of K nearest neighbor samples belonging to the majority class; The expression of the normalization process is as follows: in: Indicates o i The number of samples that need to be generated accounts for the minority class c m The total number of newly generated samples proportion; S26, based on the scaling factor and the total number of samples to be generated Calculate i The number of additional samples that need to be generated nearby g i , which is expressed as follows: S27, for all minority classes c m All existing samples o i , o i ∈c m , c m ∈S, the i-th sample data o i Will generate g i samples, then based on the SMOTE algorithm, from o i Randomly select a minority class sample o of the same category from the K nearest neighbor samples ti , generate new augmented sample data o i ′, and o i ′ and o i Combine to obtain a new sample data set; The expression for generating new expanded sample data is as follows: the i ′=o i +(o ti -o i )×a; Where: α is a scale parameter in the range [0, 1].
3. The data mining method for associating external factors of electric vehicles with charging behaviors according to claim 2 is characterized by: The step S3 specifically includes: S31: In the new sample data set, the charging type M i 、Weather conditions i Whether charging F i The dimension items are discrete values, and their variables are labeled M i ={"m1","m2"},W i ={"w1","w2","w3"}, F i ={"f1","f2"}; Location of charging pile L i 、D i 、Electricity Price C i , Charging time CT i , Temperature conditions T i If the dimension item of is a continuous value, DBSCAN clustering quantization is performed and labeled; The steps of DBSCAN clustering quantification and labeling are as follows: S32. Calculate the variable X of the nth dimension item n The continuous value range of is as follows: Range({X n (i)})=max({X n (i)})-min({X n (i)}); Where: {X n (i)} is the nth dimension item X n The set of all sample points, max is {X n (i)}, min is the maximum value among {X n (i)}; X n The number of random samples is size({X n (i)}); S33, initialize the number of quasi-clusters K, set the radius parameter ∈ of the adaptive adjustment, and calculate the minimum point parameter Minpt, which is expressed as follows: Minpt=size({X n (i)}) / K; S34, based on the minimum point parameter Minpt, check all samples of the nth dimension item, calculate the number of points in the neighborhood of each sample point and compare it with the minimum point parameter Minpt; If the number of points of the sample point is greater than or equal to Minpt, the sample point is a core point; if the number of points of the sample point is less than Minpt, the sample point is a non-core point; For non-core points, if the sample point is in the neighborhood of the core point, the sample point is marked as a boundary point; if the sample point does not belong to the neighborhood of any core point, the sample point is marked as a noise point; For a core point, if there are other core points in its neighborhood, these core points will be classified into the same cluster; S35, repeat step S34, traverse all samples until no new clusters are generated, and then assign all samples of the dimension item to the category And label it x i ; S36. Cluster and quantify all dimension items to obtain the location L of the charging pile i The labels {″l1″, ″l2″, ..., ″l n ″}、Density around charging stations i The labels {″d1″, ″d2″, ..., ″d n ″}、Electricity price C i The labels {″c1″, ″c2″, ..., ″c n ″}、Charging time CT i The labels {″ct1″, ″ct2″, ..., ″ct n ″}、Temperature condition T i The labels {″t1″, ″t2″, ..., ″t n ″}, and obtain the sample label set [time i : {″l i ″,″d i ″,″m i ″,″c i ″,″w i ″,″t i ″,″f i ″,″ct i ″}].
4. The data mining method for associating external factors of electric vehicles with charging behaviors according to claim 3 is characterized by: In step S33, the method for obtaining the adaptively adjusted radius parameter ∈ specifically includes: S331, for the i-th sample point X of the n-th dimension item n (i), calculated to X n (i) The distance between the nearest K neighbors Its expression is as follows: in: For X n The kth nearest neighbor of (i); S332, calculating the local density of the sample points, and determining the maximum local density of all sample points under n dimensions; The expression of the local density is as follows: The expression of the local density maximum is as follows: R m =max({R(X n (i)}); S333, preset the initial value of the radius parameter, and determine the n The radius parameter ∈ of (i) is as follows: Where: R m is the local maximum density, R(X n (i)) is the local density, and ∈0 is the initial value of the preset radius parameter.
5. The data mining method for associating external factors of electric vehicles with charging behaviors according to claim 4 is characterized in that: The step S4 specifically includes: S41, initialize J LLM individuals, where each LLM individual acts as an expert of an intelligent agent to play a dialogue game, and each round of the game includes a scoring phase, an explanation scoring phase, and an expert dialogue phase in sequence; Scoring stage: J experts score the set of n different dimension items in turn; each expert scores each dimension item {″l i ″,″d i ″,″m i ″,″c i ″,″w i ″,″t i ″,″f i ″,″ct i ″} are scored independently, then the Nth expert output score is {s N1 ,s N2 , ..., s NJ } (1) ; Explanation and scoring stage: J experts explain the reasons for their scoring in turn. The input for explaining the reasons for scoring is the expert's score in the scoring stage, and the output is the text explaining the reasons for scoring. Expert dialogue stage: J experts state the reasons for their ratings in turn, and the output of the previous expert's statement is used as the input of the next expert's statement. After J experts complete a round of dialogue in turn, they score n different dimension items. The output score of the Nth expert is {s N1 ,s N2 , ..., s NJ } (2) ; S42, after repeating step S41 for M rounds of game, the scores of all experts will be consistent, and the final score {s N1 ,s N2 , ..., s NJ } (M) As the weights of different dimension items {weight1, weight2, ..., weight N }, and output the sum of the weights 6. The data mining method for associating external factors of electric vehicles with charging behaviors according to claim 5 is characterized by: The step S5 specifically includes: S51, input sample label set [time i : {″l i ″,″d i ″,″m i ″,″c i ″,″w i ″,″t i ″,″f i ″,″ct i ″}], calculate the weighted support of each dimension item set, and its expression is as follows: Where: T is the transaction containing each item set, N is the total number of transactions, and X is the item set; S52, find all single item sets of items of different dimensions, retain the 1-item set whose weighted support exceeds the minimum weighted support threshold τ, and use the K-1 item set to generate a new item set, and sequentially generate candidate 1 item set to candidate K item set as candidate sets; S53, check whether each candidate K item set in the candidate set satisfies the minimum weighted support threshold τ; if it satisfies the threshold τ, it is classified as a frequent item set; if it does not satisfy the threshold τ, it is deleted from the candidate set; S54. Based on frequent item sets, find item sets X composed of external factors A and electric vehicle charging behavior item set X B The association rules between items and the weighted confidence are calculated; the weighted confidence represents the item set X A Under the premise that the item set X B Probability of occurrence; The calculation formula of the weighted confidence is as follows: S55. Based on the frequent item sets and weighted confidence, the weighted lift between the item sets of the external factor arrangement combinations and the item sets of the electric vehicle charging behavior is calculated; The calculation formula of the weighted lift is as follows: If the value of weighted lift is greater than 1, it means that item set X A With item set X B There is a positive correlation between them; If the value of weighted lift is equal to 1, it means that item set X A With item set X B There is an independent relationship between them; If the value of weighted lift is less than 1, it means that item set X A With item set X B There is a negative correlation between them.
7. A data mining system for the association between external factors and charging behavior of electric vehicles, characterized in that: The system comprises: Dataset construction module (1) is used to obtain charging behavior factors and external factors and construct the original dataset [L i , D i , M i , C i , W i , T i , F i , CT i ]; The charging behavior factors include the charging behavior information of the electric vehicle; the external factors include the charging pile equipment information and the environmental status information; The charging behavior information includes: whether charging F i and charging time CT i ; The charging pile equipment information includes: the location L of the charging pile i 、D i 、Charging Type M i 、Electricity Price C i The environmental status information includes weather conditions W i With temperature T i ; A data set expansion module (2), used to expand the original data set based on the ADASYN algorithm to obtain a new sample data set; The data set quantization discretization module (3) is used to label the dimension items with discrete values in the new sample data set, and to quantize and discretize the dimension items with continuous values in the new sample data set based on the DBSCAN clustering method, and label them with discrete labels to obtain a sample label set; A weight calculation module (4) is used to calculate the weights of different dimension items in the sample label set based on the expert scoring method simulated by the LLM large language model; The association analysis module (5) is used to improve the Apriori algorithm based on the weights of items in different dimensions, and to mine the associations between various charging behavior factors and external factors in the quantized discretized sample label set, so as to determine the main external factors affecting the charging of electric vehicles.
8. The data mining system for associating external factors of electric vehicles with charging behaviors according to claim 7, characterized in that: The data set expansion module (2) is used to expand the data according to the following steps: S21, the original data set contains N original sample data, and the N original sample data are represented by ordered number pairs as [time i :{X1(i),…,X n (i)}]; where time i is the moment, X n (i) is the nth item in the original data set, and i is the i-th group of original sample data in the original data set; S22, for each time i , taking external factors as the characteristic dimension items of the original sample data {L i , D i , M i , C i , W i , T i }, with charging behavior item set {F i , CT i } as the category of each original sample data, and construct the category set C = {c1, c2, ..., c m }; where c m Class represents the minority class, and the total number of classes is 1+[(max{CT i } / 4)]; S23, count the number of samples in each category in the category set C, and obtain the number of samples in the cth category as Nc; compare Nc with the threshold δ: if Nc<δ, it is marked as the minority class; if Nc>δ, it is marked as the majority class; count the maximum number of samples in each category N max The number and minimum N min the number of S24. Calculate each minority class c m ∈S The total number of samples that need to be generated Its expression is as follows: in: is the number of existing samples of the minority class, β∈[0,1] is the ratio of the number of new samples generated, and S is the set of all minority classes; S25. For all existing samples o in each minority class i ∈c m , calculate the scale factor of the i-th sample data and normalize it; o i is a sample data, expressed as [time i :{L i , D i , M i , C i , W i , T i , F i , CT i ]; The scaling factor r of the i-th sample data i The expression is as follows: r i =Ω i / K,i=1,…,n; Where: Ω is o i The number of K nearest neighbor samples belonging to the majority class; The expression of the normalization process is as follows: in: Indicates o i The number of samples that need to be generated accounts for the minority class c m The total number of newly generated samples proportion; S26, based on the proportional factor and the total number of samples to be generated Calculate i The number of additional samples that need to be generated nearby g i , which is expressed as follows: S27, for all minority classes c m All existing samples i , o i ∈c m , c m ∈S, the i-th sample data o i Will generate g i samples, then based on the SMOTE algorithm, from o i Randomly select a minority class sample o of the same category from the K nearest neighbor samples ti , generate new augmented sample data o i ′, and o i ′ and o i Combine to obtain a new sample data set; The expression for generating new expanded sample data is as follows: the i ′=o i +(o ti -o i )×a; Where: α is a scale parameter in the range [0, 1].
9. The data mining system for associating external factors of electric vehicles with charging behaviors according to claim 8, characterized in that: The data set quantization discretization module (3) is used to perform quantization discretization according to the following steps: S31: In the new sample data set, the charging type M i 、Weather conditions i Whether charging F i The dimension items are discrete values, and their variables are labeled M i ={"m1","m2"},W i ={"w1","w2","w3"},F i ={"f1","f2"}; Location of charging pile L i 、D i 、Electricity Price C i , Charging time CT i , Temperature conditions T i If the dimension item of is a continuous value, DBSCAN clustering quantization is performed and labeled; The steps of DBSCAN clustering quantification and labeling are as follows: S32. Calculate the variable X of the nth dimension item n The continuous value range of is as follows: Range({X n (i)})=max({X n (i)})-min({X n (i)}); Where: {X n (i)} is the nth dimension item X n The set of all sample points, max is {X n (i)}, min is the maximum value among {X n (i)}; X n The number of random samples is size({X n (i)}); S33, initialize the number of quasi-clusters K, set the radius parameter ∈ of the adaptive adjustment, and calculate the minimum point parameter Minpt, which is expressed as follows: Minpt=size({X n (i)}) / K; S34, based on the minimum point parameter Minpt, check all samples of the nth dimension item, calculate the number of points in the neighborhood of each sample point and compare it with the minimum point parameter Minpt; If the number of points of the sample point is greater than or equal to Minpt, the sample point is a core point; if the number of points of the sample point is less than Minpt, the sample point is a non-core point; For non-core points, if the sample point is in the neighborhood of the core point, the sample point is marked as a boundary point; if the sample point does not belong to the neighborhood of any core point, the sample point is marked as a noise point; For a core point, if there are other core points in its neighborhood, these core points will be classified into the same cluster; S35, repeat step S34, traverse all samples until no new clusters are generated, and then assign all samples of the dimension item to the category And label it x i ; S36. Cluster and quantify all dimension items to obtain the location L of the charging pile i The labels {"l1", "l2", ..., "l n "}、D i The labels {″d1″, ″d2″, ..., ″d n ″}、Electricity price C i The labels {″c1″, ″c2″, ..., ″c n ″}、Charging time CT i The labels {″ct1″, ″ct2″, ..., ″ct n ″}、Temperature condition T i The labels {″t1″, ″t2″, ..., ″t n ″}, and obtain the sample label set [time i : {″l i ″,″d i ″,″m i ″,″c i ″,″w i ″,″t i ″,″f i ″,″ct i ″}]; In step S33, the method for obtaining the adaptively adjusted radius parameter ∈ specifically includes: S331, for the i-th sample point X of the n-th dimension item n (i), calculated to X n (i) The distance between the nearest K neighbors Its expression is as follows: in: For X n The kth nearest neighbor of (i); S332, calculating the local density of the sample points, and determining the maximum local density of all sample points under n dimensions; The expression of the local density is as follows: The expression of the local density maximum is as follows: R m =max({R(X n (i)}); S333, preset the initial value of the radius parameter, and determine the n The radius parameter ∈ of (i) is as follows: Where: R m is the local maximum density, R(X n (i)) is the local density, and ∈0 is the initial value of the preset radius parameter.
10. The data mining system for associating external factors of electric vehicles with charging behaviors according to claim 9, characterized in that: The weight calculation module (4) is used to calculate the weight according to the following steps: S41, initialize J LLM individuals, where each LLM individual acts as an expert of an intelligent agent to play a dialogue game, and each round of the game includes a scoring phase, an explanation scoring phase, and an expert dialogue phase in sequence; Scoring stage: J experts score the set of n different dimension items in turn; each expert scores each dimension item {″l i ″,″d i ″,″m i ″,″c i ″,″w i ″,″t i ″,″f i ″,″ct i ″} are scored independently, then the Nth expert output score is {s N1 ,s N2 , ..., s NJ } (1) ; Explanation and scoring stage: J experts explain the reasons for their scoring in turn. The input for explaining the reasons for scoring is the expert's score in the scoring stage, and the output is the text explaining the reasons for scoring. Expert dialogue stage: J experts state the reasons for their ratings in turn, and the output of the previous expert's statement is used as the input of the next expert's statement. After J experts complete a round of dialogue in turn, they score n different dimension items. The output score of the Nth expert is {s N1 ,s N2 , ..., s NJ } (2) ; S42, after repeating step S41 for M rounds of game, the scores of all experts will be consistent, and the final score {s N1 ,s N2 , ..., s NJ } (M) As the weights of different dimension items {weight1, weight2, ..., weight N }, and output the sum of the weights The association analysis module (5) is used to perform association analysis on charging behavior factors and external factors according to the following steps: S51, input sample label set [time i : {″l i ″,″d i ″,″m i ″,″c i ″,″w i ″,″t i ″,″f i ″,″ct i ″}], calculate the weighted support of each dimension item set, and its expression is as follows: Where: T is the transaction containing each item set, N is the total number of transactions, and X is the item set; S52, find all single item sets of items of different dimensions, retain the 1-item set whose weighted support exceeds the minimum weighted support threshold τ, and use the K-1 item set to generate a new item set, and sequentially generate candidate 1 item set to candidate K item set as candidate sets; S53, check whether each candidate K item set in the candidate set satisfies the minimum weighted support threshold τ; if it satisfies the threshold τ, it is classified as a frequent item set; if it does not satisfy the threshold τ, it is deleted from the candidate set; S54. Based on frequent item sets, find item sets X composed of external factors A and electric vehicle charging behavior item set X B The association rules between items and the weighted confidence are calculated; the weighted confidence represents the item set X A Under the premise that the item set X B Probability of occurrence; The calculation formula of the weighted confidence is as follows: S55. Based on the frequent item sets and weighted confidence, the weighted lift between the item sets of the external factor arrangement combinations and the item sets of the electric vehicle charging behavior is calculated; The calculation formula of the weighted lift is as follows: If the value of weighted lift is greater than 1, it means that item set X A With item set X B There is a positive correlation between them; If the value of weighted lift is equal to 1, it means that item set X A With item set X B There is an independent relationship between them; If the value of weighted lift is less than 1, it means that item set X A With item set X B There is a negative correlation between them.