Natural gas sales KD-Tree SHAP attribution analysis method and device and electronic equipment
Through the KD-Tree SHAP attribution analysis method, the Shapley value of the influencing factors of natural gas sales was calculated, which solved the problem of insufficient explanatory ability in the existing technology, and achieved efficient and accurate analysis of the influencing factors of natural gas sales.
Patent Information
- Application Number
- CN202311783477.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-24
AI Technical Summary
The existing technology lacks effective attribution analysis and data impact factor in the prediction of natural gas sales volume and selling price prediction, resulting in insufficient interpretability.
The KD-Tree SHAP attribution analysis method was used to obtain the data sequence of influencing factors associated with the user's gas consumption, build KD-Tree, and combine the trained black box prediction model to calculate the Shapley value of each influencing factor to determine its impact on the user's gas consumption.
The efficiency of Shapley value calculation is improved, the problem of excessive calculation volume and unstable results affected by the order of interpreted characteristics is avoided, and efficient and accurate analysis of factors affecting natural gas sales is achieved.
Smart Images

Figure CN120198245A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a KD-Tree SHAP attribution analysis method, device and electronic device for natural gas sales volume. Background Art
[0002] The natural gas sales field is extremely important and is the core of mining the value of customer consumption data and formulating marketing strategies. Although the machine learning and deep learning methods used in technologies such as natural gas sales volume prediction and selling price prediction ensure a certain degree of accuracy, their interpretability has obvious limitations, lacking attribution analysis support and an effective mechanism for inferring data impact factors.
[0003] The attribution method is a feature-based interpretation method that explains the black-box prediction model by attaching a weight to each feature in the sample. Currently, the related technologies establish models based on the marginal effect and expectation of the impact factor for attribution analysis, mainly including five types, namely Partial Dependence Plot (PDP), Individual Conditional Expectation Plot (ICE), Permuted Feature Importance (PFI), Global Surrogate (GS), Local Surrogate (LIME), and Shapley value (SHAP), etc. Summary of the Invention
[0004] The embodiments of this application provide a KD-Tree SHAP attribution analysis method, device and electronic device for natural gas sales volume, which are used to efficiently and accurately analyze the influence degree of various influencing factors on the natural gas sales volume of users.
[0005] One of the embodiments of this application provides a KD-Tree SHAP attribution analysis method for natural gas sales volume, and the method includes:
[0006] Obtain an influence factor data sequence associated with the gas consumption of a user;
[0007] Construct a KD-Tree according to the influence factor data sequence;
[0008] Obtain the Shapley value of each influence factor data in the influence factor data sequence according to the KD-Tree and the trained black-box prediction model;
[0009] Determine the influence degree of each influencing factor data on the gas consumption of the user within a preset time period according to the Shapley value of each influencing factor data.
[0010] In some embodiments, the black-box prediction model is trained in the following manner:
[0011] Obtain a training data set;
[0012] Train an initial black-box prediction model according to the training data set to obtain the trained black-box prediction model.
[0013] In some embodiments, the training data set is obtained in the following manner:
[0014] Obtain an original data set; wherein, the original data set includes multiple groups of first historical influencing factor data sequences associated with the historical gas consumption of the user;
[0015] Sample the original data set to obtain a sampled data set;
[0016] Screen the features in the first historical influencing factor data sequences in the sampled data set to obtain a feature-screened data set composed of second historical influencing factor data sequences;
[0017] Use each group of second historical influencing factor data sequences in the feature-screened data set as a group of training data, and use the historical gas consumption of the user corresponding to the second historical influencing factor data sequence as a label to obtain the training data set.
[0018] In some embodiments, the original data set includes at least one minority class data set;
[0019] The sampling the original data set to obtain a sampled data set includes:
[0020] For any sample in the minority class data set, perform the following operations:
[0021] Calculate the cosine distance between the sample and other samples in the minority class sample set where it is located to obtain a preset number of neighboring samples;
[0022] Randomly select at least one sample from the neighboring samples as a reference sample according to the imbalance ratio of the minority sample set;
[0023] Calculate an interpolation sample according to the sample and the reference sample using a difference algorithm;
[0024] Add the interpolation sample to the original data set to obtain a sampled data set.
[0025] In some embodiments, the original data set further includes a majority class sample set;
[0026] The method further includes:
[0027] Performing the following operations on the majority class sample set:
[0028] Randomly dividing the majority class sample set into a plurality of first sample subsets with a preset number; wherein, the number of samples in each of the first sample subsets is less than or equal to the number of samples in the minority class sample set;
[0029] Combining the first sample subsets and the minority class samples to obtain second sample subsets; the sampled data set includes a plurality of the second sample subsets.
[0030] In some embodiments, screening the influencing factors in the first historical influencing factor data sequence in the sampled data set to obtain a feature-screened data set composed of a second historical influencing factor data sequence, including:
[0031] Performing the following processing on any first historical influencing factor data sequence in the sampled data set to obtain a second historical influencing factor data sequence:
[0032] Calculating the IV values of all the influencing factors in the first historical influencing factor data sequence;
[0033] Removing the influencing factor data with an IV value lower than a preset IV threshold from the first historical influencing factor data sequence;
[0034] Calculating the Pearson correlation coefficient between any two influencing factor data in the first historical influencing factor data sequence;
[0035] Removing the influencing factor data with a smaller IV value from the two influencing factor data with a Pearson correlation coefficient greater than a preset correlation threshold from the first historical influencing factor data sequence to obtain the second historical influencing factor data sequence.
[0036] In some embodiments, the initial black box prediction model includes a plurality of base classifiers; training the initial black box prediction model according to the training data set to obtain the trained black box prediction model, including:
[0037] Training a plurality of base classifiers according to the plurality of second sample subsets in the training data set;
[0038] Combining all the trained base classifiers to obtain the trained black box prediction model.
[0039] In some embodiments, training a plurality of base classifiers according to a plurality of second sample subsets in the training data set includes:
[0040] Extracting combinations of various influencing factor data from each of the second sample subsets to obtain a plurality of sub-training data sets; each of the sub-training data sets includes combinations of a plurality of pieces of the same influencing factor data;
[0041] Training a plurality of base classifiers according to the plurality of sub-training data sets.
[0042] In some embodiments, the black box prediction model is a tree-type black box prediction model;
[0043] After constructing a KD-Tree according to the influencing factor data sequence, it further includes:
[0044] Inserting virtual nodes into the KD-Tree to obtain a KD-Tree storing information of all leaf nodes;
[0045] Obtaining the Shapley value of each influencing factor data in the influencing factor data sequence according to the KD-Tree and the trained black box prediction model includes:
[0046] Performing the following operations for each node in the KD-Tree:
[0047] Obtaining combinations of various influencing factor data from the KD-Tree;
[0048] Obtaining the Shapley value of each influencing factor data in the influencing factor data sequence according to the combinations of the various influencing factor data.
[0049] One embodiment of the present application provides a KD-Tree SHAP attribution analysis device for natural gas sales volume. The device includes: a first acquisition module for acquiring an influencing factor data sequence associated with the gas consumption of a user;
[0050] A construction module for constructing a KD-Tree according to the influencing factor data sequence;
[0051] A second acquisition module for obtaining the Shapley value of each influencing factor data in the influencing factor data sequence according to the KD-Tree and the trained black box prediction model;
[0052] A determination module for determining the influence degree of each influencing factor data on the gas consumption of the user within a preset time period according to the Shapley value of each influencing factor data.
[0053] An embodiment of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor runs the program, it executes the KD-Tree SHAP attribution analysis method for natural gas sales volume as described above.
[0054] An embodiment of the present application provides a storage medium for storing a computer-readable program. When the computer-readable program is run, it executes the KD-Tree SHAP attribution analysis method for natural gas sales volume as described above.
[0055] The above technical solutions provided by the embodiments of the present application have at least the following advantages compared with the prior art:
[0056] The KD-Tree SHAP attribution analysis method for natural gas sales volume provided by the embodiments of the present application obtains a sequence of influencing factor data associated with the gas consumption of a user, constructs a KD-Tree based on the sequence of influencing factor data, and then obtains the Shapley value of each influencing factor data in the sequence of influencing factor data according to the KD-Tree and the trained black-box prediction model. Finally, the contribution degree of each influencing factor data to the gas consumption of the user within a preset time period is determined through the Shapley value of each influencing factor data. This method omits the process of sampling in the sequence of influencing factor data through the KD-Tree, effectively improves the calculation efficiency of the Shapley value, avoids the problem of excessive calculation amount of the Shapley value and unstable calculation results caused by being affected by the order of interpreted features, and can efficiently and accurately analyze the influence degree of various influencing factors on the natural gas sales volume of the user. Description of the Drawings
[0057] The present application will be further described in the form of exemplary embodiments, and these exemplary embodiments will be described in detail through the drawings; these embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where:
[0058] Figure 1 is an exemplary flowchart of the KD-Tree SHAP attribution analysis method for natural gas sales volume according to some embodiments of the present application;
[0059] Figure 2 is an exemplary schematic diagram of a KD-Tree according to some embodiments of the present application;
[0060] Figure 3 is an exemplary schematic diagram of a KD-Tree after adding virtual nodes according to some embodiments of the present application;
[0061] Figure 4 is an exemplary schematic diagram of all combination forms of different influencing factors according to some embodiments of the present application;
[0062] Figure 5 It is an exemplary schematic diagram of the predicted values of the model corresponding to different combinations of influencing factors shown in some embodiments of the present application;
[0063] Figure 6 It is an exemplary flowchart of a method for obtaining a training data set shown in some embodiments of the present application;
[0064] Figure 7 It is an exemplary schematic diagram of downsampling the samples in the minority class sample set shown in some embodiments of the present application;
[0065] Figure 8 It is an exemplary schematic diagram of the execution process of the KD-Tree SHAP attribution analysis method for natural gas sales volume shown in some embodiments of the present application;
[0066] Figure 9 It is an exemplary schematic diagram of the KD-Tree SHAP attribution analysis device for natural gas sales volume shown in some embodiments of the present application;
[0067] Figure 10 It is an exemplary structural schematic diagram of an electronic device shown in some embodiments of the present application. Detailed implementation manners
[0068] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present application. For those of ordinary skill in the art, without creative efforts, the present application can also be applied to other similar scenarios based on these drawings. Unless obvious from the language context or otherwise stated, the same reference numerals in the drawings represent the same structure or operation.
[0069] It should be understood that the "system", "device", "unit" and / or "module" used herein is a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, they can be replaced by other expressions.
[0070] As shown in the present application and the claims, unless the context clearly indicates an exceptional situation, words such as "a", "an", "one" and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0071] In this application, flowcharts are used to illustrate the operations performed by the systems according to the embodiments of this application. It should be understood that the previous or subsequent operations are not necessarily executed precisely in sequence. On the contrary, the steps can be processed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or more steps can be removed from these processes.
[0072] For ease of understanding, the technical solutions of this application are introduced below in conjunction with the accompanying drawings and embodiments.
[0073] The inventors found that in the prior art, when performing attribution analysis based on the marginal effect and expectation of the impact factor to establish a model, the PDP cannot reflect the distribution of the feature variables themselves, which will lead to too many invalid samples in the calculation process, and there is a non-uniform effect of the overall sample, which can only reflect the average level of the feature variables, ignoring the impact of data heterogeneity on the results; ICE can only reflect the relationship between a single feature variable and the target, and the ICE image often looks too redundant due to too many individuals, making it difficult to obtain explanatory information; the permutation importance method has no room to tolerate randomness, is unstable on small sample datasets, and cannot handle tree models; global surrogate models and LIME approximate black box prediction models will introduce additional errors and can only explain black box prediction models, but not data; the Shapley value is the optimal solution for solving the credible fair distribution problem in coalition games, but the SHAP calculation is extremely time-consuming and may be affected by the order of explanatory features, resulting in unstable results.
[0074] Therefore, how to efficiently and accurately analyze and obtain the influence degree of various influencing factors on the natural gas sales volume of users is a technical problem to be solved urgently. Based on this, after further research and development, the inventors made this application and provided a KD-Tree SHAP attribution analysis method, device, and electronic device for natural gas sales volume.
[0075] Embodiment 1
[0076] An embodiment of this application provides a KD-Tree SHAP attribution analysis method for natural gas sales volume. Referring to Figure 1 as shown, the method includes:
[0077] Step S110, obtaining an influence factor data sequence associated with the gas consumption of the user.
[0078] In the embodiment of this application, the user is a user who needs to predict and analyze their natural gas consumption. For example, the user can be a power plant A.
[0079] The influencing factor data sequence includes influencing factor data for multiple historical time periods. The influencing factor data (also known as influencing factors) are factors that affect the gas consumption of users, including but not limited to: real estate prosperity index, user product sales volume, product inventory, and product price index, etc.
[0080] In the specific implementation process, the influencing factor data sequence can be part or all of the sample data in the training data set obtained in step S610. Regarding the training data set and the process of obtaining the training data set, refer to the relevant content in step S610, which will not be elaborated here.
[0081] Step S120, construct a KD-Tree according to the influencing factor data sequence.
[0082] In the embodiment of the present application, after constructing a KD-Tree according to the influencing factor data sequence, it further includes:
[0083] Insert virtual nodes into the KD-Tree to obtain a KD-Tree in which all leaf nodes store information.
[0084] In the specific implementation process, by constructing a KD-Tree data structure, each influencing factor data in the influencing factor data sequence can be processed into a tree structure. As Figure 2 The given example is a KD-Tree data structure constructed according to the influencing factor data sequence, including nodes (3, 9), (1, 7), (2, 8), (5, 11), (4, 10), and (6, 12), and (3, 9), (1, 7), (2, 8), (5, 11), (4, 10), and (6, 12) are each influencing factor combination.
[0085] In the embodiment of the present application, the black box prediction model is a tree-type black box prediction model. Since all nodes in the KD-Tree contain sample prediction result information, while the sample prediction result information in the tree-type black box prediction model to be explained only exists in the leaf nodes, virtual nodes are inserted into the KD-Tree so that all leaf nodes of the KD-Tree store information, thereby solving the applicability problem. As Figure 3 The given example shows that the modified KD-Tree includes virtual nodes and sample nodes. The virtual nodes act as intermediate nodes of the KD-Tree, and their function is to select left and right branches and save information such as the number of samples flowing through the node; the sample nodes are the leaf nodes, and one leaf node corresponds to one sample, that is, one influencing factor combination, and the leaf nodes store sample data and the prediction result of the model for this sample.
[0086] Step S130, obtain the Shapley value of each influencing factor data in the influencing factor data sequence according to the KD-Tree and the trained black box prediction model.
[0087] In the embodiments of the present application, obtaining the Shapley value of each influencing factor data in the influencing factor data sequence according to the KD-Tree and the trained black-box prediction model specifically includes:
[0088] For each node in the KD-Tree, perform the following operations:
[0089] Obtain combinations of various influencing factor data from the KD-Tree;
[0090] According to the combinations of the various influencing factor data and the trained black-box prediction model, obtain the Shapley value of each influencing factor data in the influencing factor data sequence.
[0091] SHAP (Shapley additive explanation) is a method for explaining models proposed by Lundberg et al. Shapley interprets this value as an additive feature attribution method, interpreting the predicted value of the model as a linear function of multivariate variables. The function formula is as follows in Formula 1:
[0092]
[0093] In the formula, M represents the interpretation model; φ(M, x i ) is the predicted value; φ0 is interpreted as the constant of the model; is the predicted mean of all samples to be explained; d′ represents the number of input features; φ k represents the Shapley value of feature k.
[0094] Among them, the Shapley value calculation formula is as follows in Formula 2:
[0095]
[0096] In the formula, φ i (v) represents the Shapley value of influencing factor i; S represents the feature set; N represents the subset of {S\x i}; the fraction represents the probability corresponding to different feature combinations; v(S∪{i}) represents the prediction result when x i is added to the model under different feature combinations; v(S) represents the prediction result when x i is not added to the model under different feature combinations.
[0097] In the specific implementation process, an influencing factor data sequence contains N influencing factors. Set the black-box prediction model as the function v, which gives the value of any subset S of the N influencing factors, and v(S) gives the value of this subset, that is, x under different subsets iThe prediction result without adding the model. Use the above formula 2 to calculate the contribution of influencing factor i, that is, the Shapley value of influencing factor i. For example, for the influencing factor sequence N = {A, B, C, D}, when i = D, the Shapley value can calculate how much of the gas demand can be attributed to influencing factor D, that is, calculate the Shapley value of D. Consider the results of each possible combination of influencing factors in the influencing factor sequence to determine the contribution of a single influencing factor. Excluding the factor of concern, all possible subsets of the remaining influencing factors need to be considered, that is, the subsets that can be formed by {A, B, C}, such as Figure 4 as shown, including: {A}, {B}, {C}, {AB}, {BC}, {CA}, and {ABC}, a total of 8 subsets. The empty set in the figure is where v represents the number of elements in the set.
[0098] In the specific implementation process, a black-box prediction model can be trained for each subset formed by different combinations of influencing factor data sequences. The hyperparameters and training data of these black-box prediction models are exactly the same. Assuming that 8 regression models (i.e., black-box prediction models) are trained, as Figure 5 shown, a new observation sample x0 can be used to see the prediction results of 8 black-box prediction models on this observation sample (pred is the predicted value of the model for x0), and the prediction effects of these black-box prediction models can be tested.
[0099] Regarding the training process of the black-box prediction model, specifically refer to the relevant content of the training of the black-box prediction model in steps S610 - S620, which will not be elaborated here.
[0100] v(SU{i}) - v(S) is to calculate the marginal contribution of influencing factor i, that is, the marginal value of adding influencing factor i to the model. For example, for the influencing factor sequence N = {A, B, C, D}, when i = D, adding D to the 8 subsets, the marginal values are respectively expressed as in formula 2 Calculate the weights of the marginal values (i.e., the probabilities corresponding to different feature combinations), and finally the weighted average of the marginal values is the Shapley value of D.
[0101] In the specific implementation process, subsets of the influencing factor data sequence can be obtained along the nodes of the KD-Tree established in step S120. Only as an example, for Figure 2For the KD-Tree shown, sample sets such as {(3, 9)}, {(3, 9), (1, 7)}, {(3, 9), (1, 7), (2, 8)} can be obtained in sequence; each sample set is processed using the trained black-box prediction model (for the detailed description, refer to the relevant content in step S620, which will not be elaborated here), and the Shapley value of each influencing factor data in the influencing factor data sequence is calculated using formula 2.
[0102] In the embodiments provided by this application, the KD-Tree is used to sort various influencing factor data in the influencing factor data sequence; then, each combination of the influencing factor data sequence is obtained from the KD-Tree, and the Shapley value of each influencing factor data in the influencing factor data sequence is calculated. By using the KD-Tree, the process of sampling in the influencing factor data sequence is omitted, effectively improving the calculation efficiency of the Shapley value, and avoiding the problems in the related art, such as the excessive calculation amount of the Shapley value and the instability of the calculation result caused by the influence of the order of the interpreted features.
[0103] For a tree-type black-box prediction model and a given KD-Tree, for a feature subset T of a certain sample s, if the subset T contains all features, the predicted value of the leaf node is the output value of the model; if the subset T is empty, the weighted average of the predicted values of all leaf nodes is used as the output value; if the subset T contains some features, the leaf nodes that cannot be reached after removing some features are deleted, and the weighted average is taken among the remaining nodes as the output value.
[0104] Step S140: Determine the influence degree of each influencing factor data on the gas consumption of the user within a preset time period according to the Shapley value of each influencing factor data.
[0105] In the specific implementation process, the influence degree of each influencing factor data on the gas consumption of the user within the target time period can be determined according to the magnitude of the Shapley value of each influencing factor data calculated in step S130. According to the influence degree of each influencing factor data on the gas consumption of the user within the target time period, the reason for the magnitude of the gas consumption of the user within the preset time period can be learned, and then, based on this reason, the sales strategy of natural gas can be guided.
[0106] In the embodiments of this application, as Figure 6 shown, the black-box prediction model is trained through the following steps S610 - S620:
[0107] Step S610: Obtain a training data set.
[0108] In the embodiments of this application, the training data set is obtained through the following method:
[0109] Obtain the original data set; wherein, the original data set includes multiple groups of first historical influencing factor data sequences associated with the historical gas consumption of users.
[0110] Sample the original data set to obtain the sampled data set.
[0111] Screen the features in the first historical influencing factor data sequences in the sampled data set to obtain the feature-screened data set composed of second historical influencing factor data sequences.
[0112] Take each group of second historical influencing factor data sequences in the feature-screened data set as a group of training data, and take the historical gas consumption of the user corresponding to the second historical influencing factor data sequence as a label to obtain the training data set.
[0113] Only as an example, the original data includes the influencing factor data sequences of different natural gas users. It is assumed that there are n user samples in total, and there are d influencing factors in each sample, that is, there are d features in total. Then each sample in the original data set has d dimensions, and the original data set can be expressed as X = {x1, x2,..., x n}, where any sample x i can be expressed as
[0114] Since the number of data in the obtained original data set is large and the sample distribution is uneven, in order to improve the training efficiency of the model, the original data set can be sampled to obtain the sampled data set.
[0115] In the embodiment of the present application, the original data set includes at least one minority class data set.
[0116] The sampling of the original data set to obtain the sampled data set specifically includes:
[0117] For any sample in the minority class data set, perform the following operations:
[0118] Calculate the cosine distance between the sample and other samples in the minority class sample set where it is located to obtain a preset number of neighboring samples.
[0119] According to the imbalance ratio of the minority sample set, randomly select at least one sample from the neighboring samples as a reference sample.
[0120] According to the sample and the reference sample, use the difference algorithm to calculate and obtain an interpolation sample.
[0121] Add the interpolation sample to the original data set to obtain the sampled data set.
[0122] For example, as Figure 7 shown, the original data set contains class 1 samples and class 2 samples. According to the sample imbalance ratio in the original data set, a sampling ratio is set to determine the sampling magnification N. For each sample x in the minority class data set i : Calculate its distances to all other samples in the minority class sample set to obtain k (for example, 5) nearest neighbors; randomly select N (for example, 3) nearest neighbor samples from its k nearest neighbors. For any selected nearest neighbor sample o, use the following formula 3 to construct a new sample:
[0123] o(new) = o + rand(0, 1) * (x i - o) Formula 3;
[0124] In the formula, o(new) represents the generated new sample, that is, the sample synthesized by SMOTE; o represents any selected nearest neighbor sample of x i ; x i represents any sample in the minority class data set.
[0125] In the embodiments of the present application, for the minority class sample set, new samples are generated by means of data interpolation to perform data augmentation on the minority class samples, increasing the diversity of the training data set.
[0126] In the embodiments of the present application, the original data set further includes a majority class sample set;
[0127] For the majority class sample set, the following operations are performed:
[0128] Randomly divide the majority class sample set into multiple first sample subsets with a preset number; where the number of samples in each first sample subset is less than or equal to the number of samples in the minority class sample set;
[0129] Combine the first sample subset and the minority class samples to obtain a second sample subset; the sampled data set contains multiple second sample subsets.
[0130] In the embodiments of the present application, undersampling is performed by the Easy Ensemble method. Finally, each second sample subset has fewer samples than the total number of samples in the data set. However, by integrating with the minority class data, the total information volume of the second sample subset does not decrease compared to the total number of samples in the data set. When training the base classifier with the second sample subset later, multiple base classifiers can fully utilize the original data for training, enabling the finally trained classifier to have good generalization ability.
[0131] After the above data sampling process is performed on the original data set, a sampled data set containing multiple second sample subsets is obtained. The number of samples in the original data set X = {x1, x2,..., x n} changes from n to n'. The sampled data set can be expressed as X' = {x1, x2,..., x n′}. The sample (the first historical influencing factor data sequence) x i can be expressed as
[0132] In the specific implementation process, there may be redundant features (influencing factors) in the sample data. To further improve the efficiency of model training, the features in the first historical influencing factor data sequence in the sampled data set can be screened to obtain a feature-screened data set composed of the second historical influencing factor data sequence.
[0133] In the embodiment of the present application, screening the influencing factors in the first historical influencing factor data sequence in the sampled data set to obtain a feature-screened data set composed of the second historical influencing factor data sequence specifically includes:
[0134] For any first historical influencing factor data sequence in the sampled data set, perform the following processing to obtain the second historical influencing factor data sequence:
[0135] Calculate the IV values of all influencing factors in the first historical influencing factor data sequence;
[0136] Remove the influencing factor data with an IV value lower than the preset IV threshold from the first historical influencing factor data sequence;
[0137] Calculate the Pearson correlation coefficient between any two influencing factor data in the first historical influencing factor data sequence;
[0138] Remove the influencing factor data with a smaller IV value from the two influencing factor data with a Pearson correlation coefficient greater than the preset correlation threshold from the first historical influencing factor data sequence to obtain the second historical influencing factor data sequence.
[0139] In the embodiments of the present application, for any first historical influencing factor data sequence in the sampled data set, the following processing can be performed to obtain a second historical influencing factor data sequence: Calculate the information content (Information Value, IV) of all influencing factors in the first historical influencing factor data sequence; Remove the influencing factors with IV values lower than the preset IV threshold (for example, 0.02) from the first historical influencing factor data sequence; Calculate the Pearson correlation coefficient between any two influencing factors in the first historical influencing factor data sequence; Remove the influencing factor with a smaller IV value from the two influencing factors with a Pearson correlation coefficient greater than the preset correlation threshold (for example, 0.7), thereby obtaining the second historical influencing factor data sequence.
[0140] Among them, the IV value can represent the prediction ability of the input feature for the target feature. The formula for calculating the IV value is as follows, Formula 4:
[0141]
[0142] In the formula, p represents the number of groups; q represents the total number of groups for sample division on this feature; respectively represent the proportions of the samples of this category and their complementary samples in the p-th group.
[0143] The Pearson correlation coefficient formula is as follows, Formula 5:
[0144]
[0145] In the formula, Cov(x, y) represents the covariance of features x and y; σ x and σ y represent the standard deviations of the two features.
[0146] After screening the features in the first historical influencing factor data sequence in the sampled data set using the above method, the number of features in the second historical influencing factor data sequence in the sampled data set becomes d'. The data set after feature screening can be represented as X' = {x1, x2,..., x n′}, where any sample x i can be represented as
[0147] In the embodiments provided by the present application, by combining the IV value and the Pearson correlation coefficient, redundant items and weak influencing factor items in the data set are removed, effectively simplifying the calculation amount when continuously calculating the SHAP feature combination, and effectively preventing the overfitting problem in model training.
[0148] Taking each group of the second historical influencing factor data sequences in the dataset after the above-mentioned feature screening as a group of training data, and taking the historical gas consumption of the user corresponding to the second historical influencing factor data sequence as a label, a training dataset is obtained.
[0149] Step S620: Train the initial black-box prediction model according to the training dataset to obtain the trained black-box prediction model.
[0150] In the embodiment of the present application, the initial black-box prediction model includes multiple base classifiers; the training of the initial black-box prediction model according to the training dataset to obtain the trained black-box prediction model specifically includes:
[0151] Training multiple base classifiers according to multiple second sample subsets in the training dataset;
[0152] Combining all the trained base classifiers to obtain the trained black-box prediction model.
[0153] In the embodiment of the present application, the training of multiple base classifiers according to multiple second sample subsets in the training dataset specifically includes:
[0154] Extracting combinations of multiple influencing factor data from each second sample subset to obtain multiple sub-training datasets; each sub-training dataset includes combinations of multiple same influencing factor data;
[0155] Training multiple base classifiers according to the multiple sub-training datasets.
[0156] In the specific implementation process, corresponding to the calculation process of the Shapley value in step S130, since multiple feature combinations are involved in the calculation process of the Shapley value, it is necessary to train a black-box prediction model corresponding to each feature combination.
[0157] Only as an example, one feature combination (for example, {A, B} or {A, C}, etc.) can be extracted from each second historical influencing factor data sequence (for example, {A, B, C, D}) of a second sample subset to obtain a sub-training dataset; using the sub-training datasets corresponding to the same feature combination extracted from all second sample subsets to train a base classifier respectively. For example, for {A, B, C, D}, there can be 16 feature combinations, so 16 sub-training datasets can be obtained, and one sub-training dataset trains one base classifier, and a total of 16 base classifiers are trained; adding the prediction probabilities of the trained base classifiers, and finally determining the classification output through the sign function to obtain the black-box prediction model for this feature combination (for example, {A, B}).
[0158] In the embodiment of the present application, such asFigure 8 As shown in the figure, it is an exemplary schematic diagram of the execution process of the KD-Tree SHAP attribution analysis method for natural gas sales volume. First, the obtained original data is preprocessed, including missing value processing and outlier processing. The preprocessed original data set is sequentially subjected to data sampling and feature selection. Among them, data sampling can be SMOTE sampling for minority class samples in the original data set and Easy Ensemble undersampling for majority class samples, resulting in a sampled data set containing multiple second sample subsets; feature selection is performed on the sampled data set, and features with an IV value lower than a preset IV threshold (for example, 0.02) are screened through the IV value, and among two features with a Pearson correlation coefficient greater than a preset correlation threshold (for example, 0.7), the feature with a smaller IV value is removed through the Pearson correlation coefficient, obtaining a second historical influencing factor data sequence. A KD-Tree is constructed according to the second historical influencing factor sequence, and virtual nodes are inserted into the constructed KD-Tree to obtain a KD-Tree storing information of all leaf nodes. Finally, Tree SHAP values are calculated based on the KD-Tree, mainly calculating the marginal contribution of features and the weighting of each marginal value, for explaining the black box model (i.e., the black box prediction model), and determining the influence degree of each influencing factor data on the gas consumption of the user within a preset time period.
[0159] The KD-Tree SHAP attribution analysis method for natural gas sales volume provided by the embodiments of the present application obtains an influencing factor data sequence associated with the gas consumption of the user, constructs a KD-Tree according to the influencing factor data sequence, and then obtains the Shapley value of each influencing factor data in the influencing factor data sequence according to the KD-Tree and the trained black box prediction model. Finally, the contribution degree of each influencing factor data to the gas consumption of the user within a preset time period is determined through the Shapley value of each influencing factor data. This method omits the process of sampling in the influencing factor data sequence through the KD-Tree, effectively improves the calculation efficiency of the Shapley value, avoids the problems of excessive calculation amount of the Shapley value and unstable calculation results caused by being affected by the order of interpreted features, and can efficiently and accurately analyze the influence degree of various influencing factors on the natural gas sales volume of the user.
[0160] Embodiment 2
[0161] Based on the same inventive concept, the embodiments of the present application also provide a KD-Tree SHAP attribution analysis device 900 for natural gas sales volume. Referring to Figure 9 As shown in the figure, the device includes:
[0162] A first acquisition module 910, configured to acquire an influencing factor data sequence associated with the gas consumption of the user.
[0163] A building module 920 is configured to construct a KD-Tree according to the influence factor data sequence.
[0164] A second acquisition module 930 is configured to obtain the Shapley value of each influence factor data in the influence factor data sequence according to the KD-Tree and the trained black-box prediction model.
[0165] A determination module 940 is configured to determine the influence degree of each influence factor data on the gas consumption of the user within a preset time period according to the Shapley value of each influence factor data.
[0166] In the embodiments of the above KD-Tree SHAP attribution analysis device for natural gas sales volume, the specific processing of each module and the technical effects brought thereby can be respectively referred to the relevant descriptions in the corresponding method embodiments, which will not be elaborated herein.
[0167] Embodiment III
[0168] Based on the same inventive concept, an embodiment of the present application further provides an electronic device, as Figure 10 shown. The electronic device includes: at least one processor 1001, at least one communication interface 1002, at least one memory 1003, and at least one communication bus 1004; optionally, the communication interface 802 may be an interface of a communication module, such as an interface of a GSM module; the processor 1001 may be a processor CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. The memory 1003 may include a high-speed RAM memory, and may also include a non-volatile memory, for example, at least one disk memory. Among them, the memory 1003 stores a program, and the processor 1001 calls the program stored in the memory 1003 to execute some or all of the above method embodiments.
[0169] Embodiment IV
[0170] Based on the same inventive concept, an embodiment of the present application further provides a storage medium for storing a computer-readable program, and when the computer-readable program is run, it executes some or all of the above method embodiments.
[0171] Optionally, the storage medium may be a non-temporary computer-readable storage medium. For example, the non-temporary computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0172] Embodiment 5
[0173] Based on the same inventive concept, an embodiment of the present application further provides a computer program product containing instructions. When the computer program product runs on a computer device, the computer device is caused to execute some or all of the above method embodiments.
[0174] Embodiment 6
[0175] Based on the same inventive concept, an embodiment of the present application further provides a chip. The chip includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a computer program or instructions to implement some or all of the above method embodiments.
[0176] The basic concepts have been described above. Obviously, for those skilled in the art, the above detailed disclosure is only an example and does not constitute a limitation to the present application. Although not explicitly stated here, those skilled in the art may make various modifications, improvements, and corrections to the present application. Such modifications, improvements, and corrections are proposed in the present application, so such modifications, improvements, and corrections still belong to the spirit and scope of the exemplary embodiments of the present application.
[0177] At the same time, the present application uses specific terms to describe the embodiments of the present application. Such as "one embodiment", "an embodiment", and / or "some embodiments" mean a certain feature, structure, or characteristic related to at least one embodiment of the present application. Therefore, it should be emphasized and noted that the "one embodiment" or "an embodiment" or "an alternative embodiment" mentioned twice or more at different positions in the present application does not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of the present application can be appropriately combined.
[0178] In addition, unless clearly stated in the claims, the order of the processing elements and sequences, the use of numbers, letters, or other names in the present application is not used to limit the order of the processes and methods of the present application. Although some currently considered useful inventive embodiments are discussed through various examples in the above disclosure, it should be understood that such details only serve the purpose of illustration. The appended claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that conform to the essence and scope of the embodiments of the present application. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only through software solutions, such as installing the described system on existing servers or mobile devices.
[0179] Similarly, it should be noted that, in order to simplify the description disclosed in this application and thus help the understanding of one or more embodiments of the invention, in the foregoing description of the embodiments of this application, multiple features are sometimes grouped into one embodiment, drawing, or description thereof. However, this disclosure method does not mean that the features required by the subject matter of this application are more than those mentioned in the claims. In fact, the features of the embodiments are fewer than all the features of the individual embodiments disclosed above.
[0180] In some embodiments, numbers are used to describe components and the quantity of attributes. It should be understood that such numbers used for the description of embodiments are modified by the modifiers "about", "approximate", or "substantially" in some examples. Unless otherwise stated, "about", "approximate", or "substantially" indicate that the said numbers allow a variation of ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, and such approximate values may vary according to the characteristics required by individual embodiments. In some embodiments, the numerical parameters should consider the specified significant digits and adopt the method of retaining the general number of digits. Although the numerical ranges and parameters used to confirm the breadth of the scope in some embodiments of this application are approximate values, in specific embodiments, such numerical settings are as precise as possible within the feasible range.
[0181] For each patent, patent application, patent application publication, and other materials cited in this application, such as articles, books, specifications, publications, documents, etc., their entire contents are hereby incorporated into this application by reference. Except for the application history documents that are inconsistent with or conflict with the content of this application, and except for the documents that limit the broadest scope of the claims of this application (currently or subsequently attached to this application). It should be noted that if there are inconsistencies or conflicts between the descriptions, definitions, and / or uses of terms in the supplementary materials of this application and the content described in this application, the descriptions, definitions, and / or uses of terms in this application shall prevail.
[0182] Finally, it should be understood that the embodiments described in this application are only used to illustrate the principles of the embodiments of this application. Other variations may also fall within the scope of this application. Therefore, by way of example and not limitation, alternative configurations of the embodiments of this application may be considered to be consistent with the teachings of this application. Accordingly, the embodiments of this application are not limited to the embodiments explicitly introduced and described in this application.
Claims
1. A KD-Tree SHAP attribution analysis method for natural gas sales volume, characterized in that, The method includes: Obtaining a sequence of influencing factor data associated with the user's gas consumption; Constructing a KD-Tree according to the sequence of influencing factor data; Obtaining the Shapley value of each influencing factor data in the sequence of influencing factor data according to the KD-Tree and the trained black-box prediction model; Determining the influence degree of each influencing factor data on the user's gas consumption within a preset time period according to the Shapley value of each influencing factor data.
2. The method according to claim 1, wherein The black-box prediction model is trained in the following manner: Obtaining a training data set; Training an initial black-box prediction model according to the training data set to obtain the trained black-box prediction model.
3. The method according to claim 2, wherein The training data set is obtained in the following manner: Obtaining an original data set; wherein, the original data set includes multiple groups of first historical influencing factor data sequences associated with the user's historical gas consumption; Sampling the original data set to obtain a sampled data set; Screening the features in the first historical influencing factor data sequences in the sampled data set to obtain a feature-screened data set composed of second historical influencing factor data sequences; Taking each group of second historical influencing factor data sequences in the feature-screened data set as a group of training data, and taking the user's historical gas consumption corresponding to the second historical influencing factor data sequence as a label to obtain the training data set.
4. The method according to claim 3, characterized in that The original data set includes at least one minority class data set; The sampling the original data set to obtain a sampled data set includes: For any sample in the minority class data set, perform the following operations: Calculating the cosine distance between the sample and other samples in the minority class sample set where it is located to obtain a preset number of neighboring samples; Randomly selecting at least one sample from the neighboring samples as a reference sample according to the imbalance ratio of the minority sample set; Calculating an interpolation sample according to the sample and the reference sample by using a difference algorithm; Adding the interpolation sample to the original data set to obtain a sampled data set.
5. The method according to claim 4, characterized in that The original data set further includes a majority class sample set; The method further includes: For the majority class sample set, perform the following operations: Randomly dividing the majority class sample set into multiple first sample subsets with a preset number; wherein, the number of samples in each first sample subset is less than or equal to the number of samples in the minority class sample set; Combining the first sample subset and the minority class samples to obtain a second sample subset; the sampled data set contains multiple second sample subsets.
6. The method according to claim 3, characterized in that, The screening the influencing factors in the first historical influencing factor data sequences in the sampled data set to obtain a feature-screened data set composed of second historical influencing factor data sequences includes: For any first historical influencing factor data sequence in the sampled data set, perform the following processing to obtain a second historical influencing factor data sequence: Calculating the IV value of all influencing factors in the first historical influencing factor data sequence; Remove the influencing factor data with an IV value lower than the preset IV threshold from the first historical influencing factor data sequence; Calculate the Pearson correlation coefficient between any two influencing factor data in the first historical influencing factor data sequence; Remove the influencing factor data with a smaller IV value from the two influencing factor data with a Pearson correlation coefficient greater than the preset correlation threshold from the first historical influencing factor data sequence to obtain the second historical influencing factor data sequence.
7. The method according to claim 6, wherein The initial black box prediction model includes multiple base classifiers; the training of the initial black box prediction model according to the training data set to obtain the trained black box prediction model includes: Train multiple base classifiers according to multiple second sample subsets in the training data set; Combine all the trained base classifiers to obtain the trained black box prediction model.
8. The method according to claim 7, wherein The training of multiple base classifiers according to multiple second sample subsets in the training data set includes: Extract combinations of various influencing factor data from each of the second sample subsets to obtain multiple sub-training data sets; each sub-training data set includes combinations of multiple influencing factor data of the same type; Train multiple base classifiers according to the multiple sub-training data sets.
9. The method according to any one of claims 1 to 8, characterized in that, The black box prediction model is a tree-type black box prediction model; After constructing the KD-Tree according to the influencing factor data sequence, it further includes: Insert virtual nodes into the KD-Tree to obtain the KD-Tree storing information of all leaf nodes; The obtaining of the Shapley value of each influencing factor data in the influencing factor data sequence according to the KD-Tree and the trained black box prediction model includes: For each node in the KD-Tree, perform the following operations: Obtain combinations of various influencing factor data from the KD-Tree; According to the combinations of various influencing factor data and the trained black box prediction model, obtain the Shapley value of each influencing factor data in the influencing factor data sequence.
10. A KD-Tree SHAP attribution analysis device for natural gas sales volume, characterized in that The device includes: A first acquisition module for acquiring an influencing factor data sequence associated with the gas consumption of a user; A building module for constructing a KD-Tree according to the influencing factor data sequence; A second acquisition module for obtaining the Shapley value of each influencing factor data in the influencing factor data sequence according to the KD-Tree and the trained black box prediction model; A determination module for determining the influence degree of each influencing factor data on the gas consumption of the user within a preset time period according to the Shapley value of each influencing factor data.
11. An electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor runs the program, it executes the KD-Tree SHAP attribution analysis method for natural gas sales volume according to any one of claims 1 to 9.
12. A storage medium for storing a computer-readable program, which, when run, executes the KD-Tree SHAP attribution analysis method for natural gas sales volume according to any one of claims 1 to 9.
Citation Information
Cited By
Public transport passenger flow evaluation method, device and equipment and storage medium
CN121982901A