Decision tree data model establishment method
By using distributed data acquisition, dynamic sliding window, graph embedding algorithm, multi-stage feature selection, adaptive splitting criterion, multi-objective optimization, incremental pruning and differential privacy protection in the decision tree model establishment method, the problems of low processing efficiency of heterogeneous data source, difficulty in optimizing model performance, insufficient online update capability and data privacy protection are solved, and an efficient, accurate and secure decision tree model construction is achieved.
Patent Information
- Application Number
- CN202510379177.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-13
AI Technical Summary
The existing decision tree model establishment methods have inefficiency, difficulty in optimizing model performance, insufficient online update capabilities, and data privacy protection problems when dealing with heterogeneous data sources.
Heterogeneous data sources are obtained through distributed data acquisition nodes, and graph data vectorization is used to process graph data using dynamic sliding window segmented sampling timing data and graph embedding algorithm. A multi-stage feature selection model is built for feature screening, and a decision tree is generated using adaptive splitting criteria, combining multi-objective optimization and incremental pruning mechanism to optimize the model structure, and an online model update module and a differential privacy protection mechanism are introduced.
It improves the fusion and analysis capabilities of heterogeneous data, optimizes the performance of the decision tree model, enhances the online update capabilities of the model, and effectively protects data privacy.
Smart Images

Figure CN120146160A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and particularly to a method for establishing a decision tree data model. Background Art
[0002] In today's digital age, data has shown explosive growth. Its sources are extensive and its forms are diverse, covering various types such as structured data tables, time-series data streams, and graph-structured data, that is, heterogeneous data sources. These heterogeneous data contain huge value, but how to efficiently process and utilize them has become a key problem.
[0003] Traditional data processing methods have many limitations when dealing with heterogeneous data sources. For structured data tables, although there are relatively mature processing methods, their efficiency is low when performing fusion analysis with other types of data. Time-series data streams have the characteristics of dynamics and real-time nature, and traditional fixed-window sampling methods are difficult to capture their time-varying characteristics. For example, in the stock price data of the financial market, prices fluctuate frequently, and fixed windows cannot adapt to market changes in a timely manner, resulting in the loss of important information. Due to the complex topological structure of graph-structured data, traditional algorithms are difficult to effectively extract features and cannot fully explore the potential relationships in the data. Like social network data, the relationships between nodes and edges are complex, and traditional methods are difficult to accurately analyze the complex relationships between users.
[0004] As an important data mining and machine learning model, decision trees are widely used in tasks such as classification and prediction. However, existing methods for establishing decision tree models have deficiencies in dealing with complex data and optimizing model performance. In the feature selection stage, common methods often only consider a single factor and cannot comprehensively evaluate the redundancy, correlation, and interaction between features. For example, only screening features based on a simple correlation coefficient may retain a large number of redundant features, increasing the model training time and the risk of overfitting.
[0005] During the decision tree generation process, traditional splitting criteria are relatively single, such as only using information gain or Gini impurity, and it is difficult to adapt to different data distributions. When there are problems such as class imbalance or complex features in the data, a single criterion is likely to lead to an unreasonable decision tree structure, affecting the classification accuracy and generalization ability of the model. In practical application scenarios, such as medical diagnosis data, the problem of class imbalance is prominent, and traditional splitting criteria are difficult to accurately distinguish diseased and non-diseased samples.
[0006] The optimization and pruning of decision tree models also face challenges. Traditional optimization algorithms are usually oriented towards a single goal, such as only pursuing classification accuracy, ignoring the balance between model complexity and generalization error, and are prone to overfitting or underfitting. Traditional pruning methods have high computational complexity, are difficult to execute efficiently on large-scale data, and may misprune some subtrees that contribute significantly to model performance. In addition, with the dynamic change of data, the model needs to have the ability to update online. Existing online model update technologies are not sensitive enough in detecting changes in data distribution and cannot detect concept drift in a timely manner. When the data distribution changes, the model cannot be updated quickly and effectively, resulting in a decline in model performance. In e-commerce user behavior data, users' purchase preferences may change at any time. If the model cannot be updated in a timely manner, the recommendation results will be inaccurate, affecting user experience and business benefits.
[0007] In the context where data privacy protection is increasingly emphasized, existing methods for establishing decision tree models rarely consider privacy protection issues when processing data. During the process of data sharing and analysis, sensitive information may be leaked, bringing potential risks to users and enterprises. For example, medical data contains a large amount of sensitive information of patients. If it is not effectively protected during the model establishment process, it may lead to the leakage of patients' privacy. Therefore, it is urgent to develop a decision tree data model establishment method that can effectively process heterogeneous data sources, optimize the performance of decision tree models, have the ability to update online, and ensure data privacy. Summary of the Invention
[0008] The purpose of the present invention is to provide a method for establishing a decision tree data model to solve the problems raised in the above background technology.
[0009] To achieve the above purpose, the present invention provides the following technical solution: A method for establishing a decision tree data model, the method includes:
[0010] Obtain heterogeneous data sources through distributed data acquisition nodes, where the heterogeneous data sources include structured data tables, time-series data streams, and graph-structured data; perform segmented sampling on the time-series data stream based on a dynamic sliding window, and perform vectorization processing on the graph-structured data through a graph embedding algorithm to generate a unified feature representation;
[0011] Construct a multi-stage feature selection model, where the multi-stage feature selection model includes a redundancy filtering layer, a correlation analysis layer, and an interaction evaluation layer. Among them, the redundancy filtering layer quantitatively evaluates the redundancy between features based on mutual information entropy, the correlation analysis layer screens target relevant features through a partial correlation coefficient matrix, and the interaction evaluation layer calculates the synergy of feature combinations using a game theory cooperation degree index;
[0012] Construct a dynamic decision tree generation framework based on a feature subset. The dynamic decision tree generation framework adopts an adaptive splitting criterion, which fuses the information gain ratio and the change in Gini impurity, and calculates the splitting threshold through weighted harmonic mean;
[0013] Use a multi-objective optimization algorithm to iteratively adjust the decision tree structure. The multi-objective optimization algorithm takes the minimum model complexity, the highest classification accuracy, and the lowest generalization error as optimization objectives, constructs a Pareto front solution set, and selects the optimal tree structure from the Pareto front solution set based on the entropy weight TOPSIS method;
[0014] Introduce an incremental pruning mechanism. The incremental pruning mechanism constructs a pruning cost function based on node importance scoring and subtree replacement loss, generates a pruning path candidate set through the Monte Carlo tree search strategy, and uses the dynamic programming algorithm to select the global optimal pruning scheme;
[0015] Construct an online model update module. The online model update module monitors the change in data distribution through a concept drift detection algorithm. When a significant drift is detected, it triggers the reconstruction of local subtrees, injects noise into the reconstructed nodes based on the differential privacy protection mechanism, and generates an updated decision tree model.
[0016] Preferably, the segment sampling of the time-series data stream based on a dynamic sliding window includes:
[0017] Define an adaptive adjustment rule for the window length. The rule calculates the window expansion or contraction factor according to the variance change rate of the data stream and the KL divergence between adjacent windows;
[0018] Adopt an overlapping segmentation strategy to slice the time-series data, extract the frequency domain features of the segmented data based on wavelet transform, and perform dimensionality reduction and reconstruction on the frequency domain features through a convolutional autoencoder;
[0019] Add timestamp weights to the segmented data. The timestamp weights follow an exponential decay distribution with data timeliness, and construct a weighted feature matrix for subsequent feature selection.
[0020] Preferably, the interaction evaluation layer uses game theory cooperation degree indicators to calculate the synergy of feature combinations, including:
[0021] Map the feature subset to the set of participants in a cooperative game, and define the feature marginal contribution as the improvement in model accuracy before and after adding a single feature;
[0022] Calculate the cooperative contribution degree of each feature based on the Shapley value allocation algorithm, and screen out feature combinations with positive externalities through the coalition stability index;
[0023] Construct a feature interaction graph model, use the community discovery algorithm to identify feature groups with high cohesion and low coupling, and incorporate the central features of the groups as representative features into the decision tree generation framework.
[0024] Preferably, the dynamic decision tree generation framework adopts an adaptive splitting criterion, including:
[0025] Define a hybrid splitting evaluation function, which is composed of the information gain ratio, the change in Gini impurity, and the class distribution skewness, and perform non-linear weighted fusion through the radial basis function kernel;
[0026] Construct a dynamic adjustment mechanism for the splitting threshold, adjust the threshold relaxation coefficient according to the current node depth and data sparsity, and the relaxation coefficient shows a logarithmic growth trend as the depth increases;
[0027] Introduce a splitting termination condition. When the number of node samples is lower than the adaptive lower limit or the class purity is higher than the dynamic threshold, terminate the splitting and mark the current node as a leaf node.
[0028] Preferably, the multi-objective optimization algorithm iteratively adjusts the decision tree structure, including:
[0029] Use the non-dominated sorting genetic algorithm to generate a population of candidate tree structures, and the individuals in the population are encoded as the pre-order traversal sequence of the binary tree;
[0030] Define structure mutation operators, including subtree crossover, node rotation, and branch weight perturbation, and avoid local optima through the tabu search strategy;
[0031] Construct a diversity preservation mechanism, calculate the difference degree of population individuals using the Hamming distance and the tree edit distance, and dynamically adjust the crossover probability and mutation probability.
[0032] Preferably, the incremental pruning mechanism generates a candidate set of pruning paths based on the Monte Carlo tree search strategy, including:
[0033] Model the decision tree pruning process as a Markov decision process, define the state as the current tree structure, the action as the pruning node selection, and the reward as the improvement in the model performance after pruning;
[0034] Use the upper confidence bound UCB algorithm to balance exploration and exploitation, and construct a pruning strategy tree by simulating the expected reward value of the pruning path;
[0035] Use the Bayesian optimization algorithm to adaptively adjust the pruning depth, and predict the confidence interval of the pruning effect based on Gaussian process regression.
[0036] Preferably, the online model update module monitors the change in data distribution through the concept drift detection algorithm, including:
[0037] Construct a sliding window comparison statistic, and calculate the data distribution difference degree between windows based on chi-square test and Wasserstein distance;
[0038] Adopt an integrated detection strategy, and combine the joint decision mechanism of KS test, mean shift detection and the change of the trace of the covariance matrix;
[0039] When drift is detected, trigger local subtree reconstruction, retain the unaffected branches in the original tree, and only perform recursive splitting based on Gini gain on the paths related to the drift.
[0040] Preferably, the differential privacy protection mechanism injects noise into the reconstructed nodes, including:
[0041] Define the node partition sensitivity as the maximum difference in the leaf node category count, and generate noise that satisfies ε-differential privacy based on the Laplace mechanism;
[0042] Adopt an adaptive noise allocation strategy, adjust the noise intensity according to the node hierarchy depth and data sparsity, and the amount of noise injected into deep nodes decays exponentially with depth;
[0043] Normalize and correct the probabilities of leaf nodes after injecting noise to ensure that the sum of category probabilities is 1 and the distribution is smooth.
[0044] Preferably, the graph embedding algorithm vectorizes the graph structure data, including:
[0045] Adopt a heterogeneous information network representation learning method, and generate node sequences based on the meta-path-guided random walk strategy;
[0046] Use a hierarchical attention mechanism to aggregate the features of neighbor nodes, where the structural attention weights are jointly determined by the edge type and the node degree;
[0047] Optimize the embedding representation through negative sampling contrast learning, maximize the mutual information of positive sample pairs and minimize the similarity of negative sample pairs.
[0048] Preferably, the present invention further includes an electronic device, including:
[0049] A processor;
[0050] A memory for storing instructions executable by the processor;
[0051] Wherein, the processor is configured to call the instructions stored in the memory to execute the method for building the decision tree data model described above.
[0052] Compared with the prior art, the beneficial effects of the present invention are:
[0053] Heterogeneous data sources such as structured data tables, time-series data streams, and graph-structured data are obtained through distributed data collection nodes, and specific processing methods are adopted for different types of data. The time-series data stream is segmented and sampled using a dynamic sliding window, which can adaptively adjust the window length according to the variance change rate of the data stream and the KL divergence between adjacent windows, accurately capturing the dynamic characteristics of the time-series data. The graph-structured data is vectorized using a graph embedding algorithm based on the meta-path-guided random walk strategy and the hierarchical attention mechanism, fully mining the complex relationships between nodes in the graph data, generating a unified feature representation, and greatly enhancing the ability to fuse and analyze heterogeneous data.
[0054] The multi-stage feature selection model comprehensively considers the relationships between features from three aspects: redundancy filtering, correlation analysis, and interaction evaluation. Mutual information entropy quantifies feature redundancy, the partial correlation coefficient matrix screens target-related features, and the game theory cooperation degree index calculates the synergistic effect of feature combinations, effectively removing redundant features, retaining key features, reducing the model complexity, and improving the training efficiency. The adaptive splitting criterion of the dynamic decision tree generation framework combines the information gain rate and the change in Gini impurity, combines the class distribution skewness, and through non-linear weighted fusion with the radial basis function kernel, can also dynamically adjust the splitting threshold according to the node depth and data sparsity, introducing a reasonable splitting termination condition, making the decision tree structure more suitable for the data distribution, and significantly improving the classification accuracy.
[0055] The multi-objective optimization algorithm takes the minimum model complexity, the highest classification accuracy, and the lowest generalization error as the optimization objectives, uses the non-dominated sorting genetic algorithm to generate a population of candidate tree structures, cooperates with the structure mutation operator and the diversity preservation mechanism to construct the Pareto front solution set, effectively balancing multiple performance indicators of the model, avoiding overfitting and underfitting, and enhancing the generalization ability of the model. The incremental pruning mechanism constructs a pruning cost function based on the node importance score and the subtree replacement loss, and through the Monte Carlo tree search strategy and the dynamic programming algorithm, accurately selects the globally optimal pruning scheme, greatly reducing the model complexity and improving the model operation efficiency without sacrificing too much classification accuracy.
[0056] The online model update module uses the concept drift detection algorithm, and by comparing statistical quantities and the ensemble detection strategy using a sliding window, can monitor the change in data distribution in a timely and accurate manner. Once a significant drift is detected, it triggers the local subtree reconstruction, retains the unaffected branches, and only recursively splits the drift-related paths based on the Gini gain, ensuring that the model can quickly adapt to data changes and maintain good performance. The differential privacy protection mechanism generates noise that satisfies ε-differential privacy according to the node partition sensitivity when reconstructing nodes, and adopts an adaptive noise allocation strategy to normalize and correct the probabilities of the leaf nodes after injecting noise, effectively protecting data privacy and reducing the risk of data leakage while ensuring the model performance. Description of the Drawings
[0057] Figure 1 This is the working principle diagram of the decision tree data model establishment method described in the present invention;
[0058] Figure 2 This is the flowchart of the dynamic decision tree node splitting based on the adaptive splitting criterion;
[0059] Figure 3 This is the flowchart for generating the candidate set of the decision tree pruning path based on Monte Carlo tree search. Specific embodiments
[0060] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0061] Please refer to Figures 1-3 , the present invention provides a method for establishing a decision tree data model, and the steps of the method include:
[0062] Use distributed data acquisition nodes to obtain heterogeneous data sources including structured data tables, time series data streams, and graph structure data. For time series data streams, perform segmented sampling based on a dynamic sliding window. By defining the window length adaptive adjustment rule, calculate the window expansion or contraction factor according to the variance change rate of the data stream and the KL divergence between adjacent windows; adopt an overlapping segmentation strategy to slice the time series data, extract the frequency domain features of the segmented data based on wavelet transform, and perform dimensionality reduction and reconstruction on the frequency domain features through a convolutional autoencoder; add timestamp weights to the segmented data to construct a weighted feature matrix. For graph structure data, perform vectorization processing using a graph embedding algorithm. Through a heterogeneous information network representation learning method, generate node sequences based on the meta-path-guided random walk strategy; use a hierarchical attention mechanism to aggregate neighbor node features; optimize the embedding representation through negative sampling contrast learning to generate a unified feature representation.
[0063] Construct a multi-stage feature selection model including a redundancy filtering layer, a correlation analysis layer, and an interaction evaluation layer. The redundancy filtering layer quantitatively evaluates the redundancy between features based on mutual information entropy; the correlation analysis layer screens target-related features through a partial correlation coefficient matrix; the interaction evaluation layer maps the feature subset to a set of participants in a cooperative game, calculates the collaborative contribution degree of each feature based on the Shapley value allocation algorithm, screens feature combinations with positive externalities through the coalition stability index, constructs a feature interaction graph model, uses a community discovery algorithm to identify highly cohesive and low-coupling feature groups, and incorporates the central features of the groups as representative features into the decision tree generation framework.
[0064] Construct a dynamic decision tree generation framework based on the feature subset after feature selection. This framework adopts an adaptive splitting criterion, defines a mixed splitting evaluation function composed of information gain ratio, change in Gini impurity, and class distribution skewness, and performs non-linear weighted fusion through a radial basis function kernel; constructs a dynamic adjustment mechanism for splitting thresholds, and adjusts the threshold relaxation coefficient according to the current node depth and data sparsity; introduces a splitting termination condition, and terminates splitting when the number of node samples is lower than the adaptive lower limit or the class purity is higher than the dynamic threshold, and marks the current node as a leaf node.
[0065] Use a multi-objective optimization algorithm to iteratively adjust the decision tree structure, with the minimum model complexity, the highest classification accuracy, and the lowest generalization error as the optimization objectives. Use the non-dominated sorting genetic algorithm to generate a population of candidate tree structures, and the individuals in the population are encoded as the pre-order traversal sequence of a binary tree; define structure mutation operators, including subtree crossover, node rotation, and branch weight perturbation, and avoid local optima through a tabu search strategy; construct a diversity preservation mechanism, calculate the difference degree of population individuals using Hamming distance and tree edit distance, dynamically adjust the crossover probability and mutation probability, construct a Pareto front solution set, and select the optimal tree structure from the solution set based on the entropy weight TOPSIS method.
[0066] Introduce an incremental pruning mechanism, and construct a pruning cost function based on node importance scoring and subtree replacement loss. Model the decision tree pruning process as a Markov decision process, use the upper confidence bound UCB algorithm to balance exploration and exploitation, construct a pruning strategy tree by simulating the expected reward value of the pruning path; use the Bayesian optimization algorithm to adaptively adjust the pruning depth, predict the confidence interval of the pruning effect based on Gaussian process regression, generate a candidate set of pruning paths through the Monte Carlo tree search strategy, and use the dynamic programming algorithm to select the global optimal pruning plan.
[0067] Construct an online model update module, and monitor the change of data distribution through a concept drift detection algorithm. Construct a sliding window comparison statistic, and calculate the data distribution difference degree between windows based on chi-square test and Wasserstein distance; adopt an integrated detection strategy, combined with the joint decision mechanism of KS test, mean shift detection, and change in covariance matrix trace; when a significant drift is detected, trigger local subtree reconstruction, retain the unaffected branches in the original tree, and only perform recursive splitting based on Gini gain on the paths related to the drift. Inject noise into the reconstructed nodes based on the differential privacy protection mechanism, define the node division sensitivity as the maximum difference in leaf node class counts, and generate noise that satisfies ε-differential privacy based on the Laplace mechanism; adopt an adaptive noise allocation strategy, and adjust the noise intensity according to the node hierarchy depth and data sparsity; normalize and correct the probabilities of the leaf nodes after injecting noise to generate an updated decision tree model.
[0068] The following further describes the implementation of the present invention in combination with Embodiments 1 to 5.
[0069] Embodiment 1:
[0070] This embodiment mainly describes the specific process of segmenting and sampling time-series data streams based on a dynamic sliding window, which aims to more effectively process time-series data streams, extract valuable features, and provide a high-quality data basis for subsequent model construction.
[0071] In actual operation, first define the rule for adaptively adjusting the window length. For a time-series data stream , calculate its variance change rate. Let the current window be , the window length be , and the variance of the window be . The variance change rate is calculated by the formula (when , a very small non-zero value can be set for calculation to avoid division-by-zero errors).
[0072] At the same time, calculate the KL divergence between adjacent windows. The KL divergence is used to measure the difference between two probability distributions. For adjacent windows and , their probability distributions are obtained by frequency statistics of the data within the window, denoted as and respectively, then the KL divergence .
[0073] According to the variance change rate and the KL divergence between adjacent windows, calculate the window expansion or contraction factor . For example, the formula can be adopted, where and are weight coefficients set according to experience, and . If , the window is appropriately expanded; if , the window is contracted.
[0074] Next, adopt an overlapping segmentation strategy to slice the time-series data. Assume the window length is , the overlapping length is ( ), starting from the starting position of the time-series data, slice it successively with as the step size to obtain a series of segmented data.
[0075] For each segment of data, extract frequency domain features based on wavelet transform. Wavelet transform is a time-frequency analysis method. By selecting an appropriate wavelet basis function , perform wavelet transform on the segmented data , where is the scale parameter, and is the translation parameter. By adjusting the values of and , frequency domain features at different frequencies and time positions can be obtained.
[0076] Then, use a convolutional autoencoder to perform dimensionality reduction and reconstruction on the frequency domain features. A convolutional autoencoder is a deep learning model composed of an encoder and a decoder. The encoder maps the frequency domain features to a low-dimensional space through convolutional layers, and the decoder then reconstructs the low-dimensional features back to the original dimension. When training the convolutional autoencoder, the goal is to minimize the reconstruction error and continuously adjust the parameters of the model to achieve effective dimensionality reduction of the frequency domain features.
[0077] Finally, add timestamp weights to the segmented data. Let the segmented data be , and its corresponding timestamp be . The timestamp weight follows an exponential decay distribution with data timeliness and can be expressed as , where is the decay coefficient, is the timestamp of the latest data collected currently. Construct the weighted segmented data into a weighted feature matrix for subsequent feature selection.
[0078] Example 2:
[0079] This example elaborates in detail the implementation method of calculating the synergistic effect of feature combinations using the game theory cooperation degree index in the interaction evaluation layer, aiming to more accurately screen out feature combinations that have a positive synergistic effect on the decision tree model and improve the performance and accuracy of the model.
[0080] In practical applications, first map the feature subset to the set of participants in the cooperative game. Suppose there is a feature subset , and each feature is regarded as a participant in the game. Define the feature marginal contribution as the improvement in model accuracy before and after adding a single feature. Let the initial model be . When adding the feature , the model is obtained. The model accuracy is measured by the classification accuracy on the validation dataset, denoted as . Then the marginal contribution of the feature is
[0081] Next, the Shapley value allocation algorithm is used to calculate the collaborative contribution degree of each feature. The Shapley value is a method for fairly allocating benefits in cooperative games. For a feature subset , the feature 's Shapley value 's calculation formula is:
[0082]
[0083] where, is a subset in that does not contain , represents the number of elements in the subset , and respectively represent the improvement in model accuracy when adding and not adding to the subset .
[0084] After calculating the Shapley value of each feature, the coalition stability index is used to screen feature combinations with positive externalities. The coalition stability index can adopt concepts such as the Core and Stable Set. Taking the Core as an example, a feature combination is in the Core if and only if for any other feature combination , there is , that is, the total collaborative contribution degree of this feature combination is not lower than any other combination. By this way, feature combinations with positive externalities are screened out, and these combinations can cooperate with each other in the decision tree model to improve the model performance.
[0085] Then, a feature interaction graph model is constructed. Using features as nodes, if there is a strong synergistic effect between two features (judged by the Shapley value for example), then an edge is connected between them, and the weight of the edge can be set according to the strength of the synergistic effect. The community discovery algorithm, such as the Louvain algorithm, is used to identify feature groups with high cohesion and low coupling. The Louvain algorithm optimizes the modularity function by continuously merging nodes, where is the total number of edges in the graph, represents whether there is an edge connection between nodes and (1 if there is an edge, 0 if there is no edge), is the degree of node , is the community to which node belongs, is 1 when , otherwise 0. Through this algorithm, features can be divided into different groups.
[0086] Finally, the group center features are incorporated into the decision tree generation framework as representative features. The group center features can be the features with the largest degree, or selected according to the importance scores of the features (such as Shapley values). These representative features can reflect the information of other features within the group to a certain extent, reduce the feature dimension while retaining important information, and improve the construction efficiency and performance of the decision tree.
[0087] Example 3:
[0088] This example describes in detail the content of adopting an adaptive splitting criterion for the dynamic decision tree generation framework, whose significance lies in enabling the decision tree to dynamically adjust the splitting strategy according to the characteristics of the data during the construction process, improving the accuracy and adaptability of the decision tree.
[0089] In the specific implementation, first define the mixed splitting evaluation function. Let the information gain ratio be , the change in Gini impurity be , and the class distribution skewness be . The information gain ratio is used to measure the change in information before and after splitting, the change in Gini impurity reflects the change in data purity after splitting, and the class distribution skewness is used to describe the degree of imbalance in the class distribution. The mixed splitting evaluation function is obtained by non-linearly weighted fusion of these three indicators through a radial basis function kernel, that is , where are the weight coefficients, satisfying , is the radial basis function, such as the Gaussian radial basis function , respectively represent , and , are the corresponding reference values, is the bandwidth parameter of the radial basis function.
[0090] Next, construct a dynamic adjustment mechanism for the splitting threshold. During the construction of the decision tree, as the depth of the node increases, the local characteristics of the data will change, and at the same time, the data sparsity will also affect the splitting decision. Let the current node depth be , the data sparsity be , define the threshold relaxation coefficient , increases logarithmically with the depth , for example, it can be expressed as . According to the threshold relaxation coefficient adjust the splitting threshold , , where is the initial splitting threshold. In this way, splitting can be carried out more strictly in the shallow layer of the tree, while the splitting conditions can be appropriately relaxed in the deep layer to avoid overfitting.
[0091] Then, a splitting termination condition is introduced. When the number of samples in a node is lower than the adaptive lower limit or the class purity is higher than the dynamic threshold , the splitting is terminated and the current node is marked as a leaf node. The adaptive lower limit can be set according to the scale and characteristics of the dataset. For example , where is the total number of samples in the dataset, is a small proportionality coefficient. The class purity is measured by calculating the proportion of the most dominant class in the node. When this proportion is greater than the dynamic threshold , it is considered that the class of this node is pure enough and no further splitting is required. The dynamic threshold can be adjusted according to the depth of the decision tree, such as , where is the initial class purity threshold, is the maximum depth allowed for the decision tree. As the depth increases, the class purity threshold gradually increases to ensure the quality of the leaf nodes.
[0092] Example 4:
[0093] This example details the specific implementation steps of the multi-objective optimization algorithm for iteratively adjusting the decision tree structure, aiming to achieve a better balance among the complexity, classification accuracy, and generalization error of the model through optimizing the decision tree structure.
[0094] In actual operation, first, the non-dominated sorting genetic algorithm is used to generate a population of candidate tree structures. The decision tree structure is encoded as the pre-order traversal sequence of a binary tree. For example, for a simple binary tree with the root node , the left child node , and the right child node , its pre-order traversal sequence is . The initial population is obtained by randomly generating a certain number of pre-order traversal sequences of binary trees.
[0095] Then, non-dominated sorting is performed on the population. For two candidate tree structures and , if is not inferior to in terms of the three objectives of model complexity, classification accuracy, and generalization error, and is superior to in at least one objective, then is said to dominate The individuals in the population are stratified according to whether they are dominated by other individuals. The first layer consists of individuals not dominated by any other individual, i.e., non-dominated individuals. The second layer consists of individuals only dominated by the individuals in the first layer, and so on.
[0096] Next, define the structure mutation operators, including subtree crossover, node rotation, and branch weight perturbation. The subtree crossover operation randomly selects subtrees from two parent tree structures for exchange to generate a new offspring tree structure. For example, a certain subtree structure of the parent tree is , and a certain subtree structure of the parent tree is . After exchanging these two subtrees, a new offspring tree structure is obtained. The node rotation operation exchanges the position of a node with its child node to change the structure of the tree. The branch weight perturbation randomly makes a small adjustment to the weights of the branches in the tree to simulate the mutation process in natural evolution.
[0097] To avoid local optima, a tabu search strategy is adopted. After each mutation operation, record the recently used mutation operations to form a tabu list. If a new mutation operation is the same as the operation in the tabu list, then this operation is prohibited from being executed, and other feasible mutation operations are selected instead, thereby broadening the search space and increasing the probability of finding the global optimal solution.
[0098] Construct a diversity preservation mechanism, and calculate the individual difference degree of the population using the Hamming distance and the tree edit distance. The Hamming distance is used to measure the number of different characters between two binary sequences. For the tree structure encoded by the preorder traversal sequence of a binary tree, it can be converted into a binary sequence and then the Hamming distance is calculated. The tree edit distance measures the minimum number of operations required to convert one tree structure into another through node insertion, deletion, and modification operations. Dynamically adjust the crossover probability and the mutation probability according to the individual difference degree. When the individual difference degree of the population is small, increase the crossover probability and the mutation probability to promote the diversity of the population; when the individual difference degree is large, appropriately reduce the crossover probability and the mutation probability to retain the current better solution structure. For example, when the individual difference degree is less than the set threshold , , , where and are the initial crossover probability and mutation probability, is the adjustment coefficient, and .
[0099] By continuously iterating the above process, a series of candidate tree structures are generated to construct the Pareto front solution set. The Pareto front solution set is a set of solutions that cannot be dominated by other solutions in multi-objective optimization. In this solution set, any improvement in one objective will necessarily lead to a decrease in other objectives.
[0100] To select the optimal tree structure from the Pareto front solution set, the entropy weight TOPSIS method is adopted. First, data preprocessing is performed on the three objectives (model complexity, classification accuracy, and generalization error) of the decision tree structure to convert them into unified dimensionless indicators. Suppose there are candidate tree structures in the Pareto front solution set, and each structure has objective indicators, forming the decision matrix , where represents the value of the th objective indicator of the th candidate tree structure.
[0101] Next, calculate the entropy value of each objective indicator. The entropy value reflects the information content of the indicator, and the calculation formula is , where . Then, calculate the weight of each objective indicator according to the entropy value, .
[0102] Determine the positive ideal solution and the negative ideal solution . The positive ideal solution is a vector composed of the optimal values of each objective indicator, and the negative ideal solution is a vector composed of the worst values of each objective indicator. For model complexity, the smaller its value, the better; for classification accuracy and generalization error, the larger their values, the better. Therefore, the positive ideal solution , and the negative ideal solution .
[0103] Calculate the distances and from each candidate tree structure to the positive ideal solution and the negative ideal solution. The distance uses the weighted Euclidean distance, , , where and are the th components of the positive ideal solution and the negative ideal solution respectively.
[0104] Finally, calculate the relative closeness of each candidate tree structure, . The larger the relative closeness, the closer the candidate tree structure is to the positive ideal solution and the farther it is from the negative ideal solution. Select the candidate tree structure with the largest relative closeness as the final optimal tree structure.
[0105] Example 5:
[0106] This example will describe in detail each link of the online model update module, whose function is to enable the decision tree model to adapt to the dynamic changes of the data distribution and maintain the effectiveness and accuracy of the model.
[0107] Concept drift detection: The first step of online model update is to monitor the change of data distribution, which is achieved through the concept drift detection algorithm. Build a sliding window comparison statistic, and use the chi-square test and Wasserstein distance to calculate the data distribution difference degree between windows.
[0108] First, set two sliding windows, one is the reference window , which stores historical data; the other is the test window , which stores the currently newly collected data. For the chi-square test, divide the eigenvalue of the data into several intervals, and count the frequency of the data in each interval. Let the frequency of the th interval in the reference window be , and the frequency of the th interval in the test window be , and the expected frequency . The chi-square statistic . According to the chi-square distribution table, determine the critical value at a given significance level. If is greater than the critical value, it is considered that there is a significant difference in the data distribution between the two windows.
[0109] The Wasserstein distance is used to measure the minimum movement cost between two probability distributions. For the data distributions of the reference window and the test window and , the Wasserstein distance , where ) is the set of all joint distributions with and as the marginal distributions. In actual calculation, an approximate algorithm such as the Sinkhorn algorithm can be used for calculation.
[0110] Adopt an integrated detection strategy, combining the joint decision mechanism of the KS test, mean shift detection, and the change of the trace of the covariance matrix. The KS test calculates the maximum difference value by comparing the empirical distribution functions of the data in the two windows, where and ) are the empirical distribution functions of the data in the reference window and the test window respectively. According to the KS distribution table, determine the critical value. If If it is greater than the critical value, it is considered that concept drift exists.
[0111] Mean shift detection calculates the mean vectors of the data in two windows and , and calculates the change amount of the mean vectors . When exceeds the set threshold, it is determined that mean shift exists.
[0112] Covariance matrix trace change detection calculates the covariance matrices of the data in the reference window and the test window and , and calculates the change amount of the trace . When exceeds the set threshold, it is considered that the covariance structure has changed significantly.
[0113] Only when at least two of the detection methods such as chi-square test, Wasserstein distance, KS test, mean shift detection, and covariance matrix trace change detect significant drift, will local subtree reconstruction be triggered.
[0114] Local subtree reconstruction: When significant drift is detected, retain the unaffected branches in the original tree and perform recursive splitting based on Gini gain only on the paths related to the drift.
[0115] First, determine the paths related to the drift. Starting from the root node of the decision tree, traverse down along the branches of the decision tree according to the feature values of the new data at each node, and record the node paths passed through. This path is the path that may be affected by data drift.
[0116] For each node on the path, calculate the Gini gain. Let the node contain the sample set , and the feature has values. Divide the sample set into subsets . The Gini impurity of the node , where is the proportion of samples belonging to class in the sample set , and is the total number of classes. The Gini gain of the feature for the node .
[0117] Select the feature with the largest Gini gain to split the node, and recursively perform the same operation on the split child nodes until the splitting termination condition is met, such as the number of node samples being lower than the adaptive lower limit or the class purity being higher than the dynamic threshold.
[0118] Differential privacy protection: To protect the privacy of data, noise injection is performed on the reconstructed nodes based on the differential privacy protection mechanism.
[0119] Define the node partition sensitivity as the maximum difference in the class counts of leaf nodes. Let the leaf node have the sample count of class be , and the node partition sensitivity .
[0120] Generate noise that satisfies differential privacy based on the Laplace mechanism. The Laplace noise , where is the differential privacy budget, which controls the degree of privacy protection. For each class count of the reconstructed nodes, add the Laplace noise .
[0121] Adopt an adaptive noise allocation strategy to adjust the noise intensity according to the node hierarchy depth and data sparsity. Let the node hierarchy depth be , the data sparsity be , and the adjustment coefficient . The noise intensity . When generating the Laplace noise, use the adjusted noise intensity , that is, .
[0122] Normalize and correct the probabilities of the leaf nodes after injecting noise. Let the count of class in the leaf nodes after injecting noise be , the total count , and the corrected class probability .
[0123] Through the above steps, the online model is updated to generate an updated decision tree model, enabling it to better adapt to changes in the data distribution.
[0124] In summary, through the above five embodiments, the specific implementation methods of each key link in the decision tree data model establishment method of the present invention are elaborated in detail. From data collection and preprocessing, feature selection, decision tree generation, structure optimization to online model update, each link cooperates with each other to jointly construct an efficient, accurate and adaptive decision tree data model.
[0125] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.
[0126] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for establishing a decision tree data model, characterized in that: The method comprises: Obtain heterogeneous data sources through distributed data collection nodes, wherein the heterogeneous data sources include structured data tables, time series data streams, and graph structure data; perform segmented sampling on the time series data stream based on a dynamic sliding window, perform vectorization processing on the graph structure data through a graph embedding algorithm, and generate a unified feature representation; Constructing a multi-stage feature selection model, the multi-stage feature selection model includes a redundant filtering layer, a correlation analysis layer, and an interaction evaluation layer, wherein the redundant filtering layer quantitatively evaluates the redundancy between features based on mutual information entropy, the correlation analysis layer screens target-related features through a partial correlation coefficient matrix, and the interaction evaluation layer calculates the synergistic effect of feature combinations using a game theory cooperation index; A dynamic decision tree generation framework is constructed based on feature subsets, wherein the dynamic decision tree generation framework adopts an adaptive splitting criterion, wherein the adaptive splitting criterion integrates the information gain rate and the Gini impurity change, and calculates the splitting threshold by weighted harmonic average; A multi-objective optimization algorithm is used to iteratively adjust the decision tree structure. The multi-objective optimization algorithm takes the minimum model complexity, the highest classification accuracy and the lowest generalization error as optimization goals, constructs a Pareto front solution set, and selects the optimal tree structure from the Pareto front solution set based on the entropy weight TOPSIS method; An incremental pruning mechanism is introduced. The incremental pruning mechanism constructs a pruning cost function based on node importance scores and subtree replacement losses, generates a pruning path candidate set through a Monte Carlo tree search strategy, and selects the global optimal pruning solution using a dynamic programming algorithm. An online model updating module is constructed. The online model updating module monitors the changes in data distribution through a concept drift detection algorithm. When a significant drift is detected, the local subtree reconstruction is triggered. Noise is injected into the reconstructed nodes based on the differential privacy protection mechanism to generate an updated decision tree model.
2. The method for establishing a decision tree data model according to claim 1, characterized in that: The segmented sampling of the time series data stream based on the dynamic sliding window includes: Define a window length adaptive adjustment rule, wherein the rule calculates a window expansion or contraction factor according to a variance change rate of a data stream and a KL divergence between adjacent windows; The overlapping segmentation strategy is used to slice the time series data, the frequency domain features of the segmented data are extracted based on wavelet transform, and the frequency domain features are reconstructed by reducing the dimension through convolutional autoencoder; Timestamp weights are added to the segmented data. The timestamp weights decay exponentially with the timeliness of the data, and a weighted feature matrix is constructed for subsequent feature selection.
3. The method for establishing a decision tree data model according to claim 2, characterized in that: The interaction evaluation layer uses game theory cooperation index to calculate the synergistic effect of feature combinations, including: Map feature subsets to a set of participants in the cooperative game, and define feature marginal contribution as the improvement in model accuracy before and after a single feature is added; The collaborative contribution of each feature is calculated based on the Shapley value allocation algorithm, and the feature combination with positive externalities is screened through the alliance stability index; A feature interaction graph model is constructed, and the community discovery algorithm is used to identify feature groups with high cohesion and low coupling. The group center features are incorporated into the decision tree generation framework as representative features.
4. The method for establishing a decision tree data model according to claim 3, characterized in that: The dynamic decision tree generation framework adopts an adaptive splitting criterion including: A hybrid split evaluation function is defined, wherein the function is composed of information gain rate, Gini impurity change and category distribution skewness, and nonlinear weighted fusion is performed through a radial basis function kernel; Construct a dynamic adjustment mechanism for split thresholds, and adjust the threshold relaxation coefficient according to the current node depth and data sparsity. The relaxation coefficient increases logarithmically with increasing depth. A split termination condition is introduced. When the number of node samples is lower than the adaptive lower limit or the category purity is higher than the dynamic threshold, the split is terminated and the current node is marked as a leaf node.
5. The method for establishing a decision tree data model according to claim 4, characterized in that: The multi-objective optimization algorithm iteratively adjusts the decision tree structure including: A non-dominated sorting genetic algorithm is used to generate a candidate tree structure population, wherein the population individuals are encoded as a pre-order traversal sequence of a binary tree; Define structural mutation operators, including subtree crossover, node rotation, and branch weight perturbation, and avoid local optimality through taboo search strategy; Construct a diversity preservation mechanism, use Hamming distance and tree edit distance to calculate the individual differences of the population, and dynamically adjust the crossover probability and mutation probability.
6. The method for establishing a decision tree data model according to claim 5, characterized in that: The incremental pruning mechanism generates a candidate set of pruning paths based on a Monte Carlo tree search strategy, including: The decision tree pruning process is modeled as a Markov decision process, where the state is defined as the current tree structure, the action is the pruning node selection, and the reward is the improvement in model performance after pruning. The upper confidence bound (UCB) algorithm is used to balance exploration and utilization, and the pruning strategy tree is constructed by simulating the expected reward value of the pruning path. The Bayesian optimization algorithm is used to adaptively adjust the pruning depth, and the confidence interval of the pruning effect is predicted based on Gaussian process regression.
7. The method for establishing a decision tree data model according to claim 6, characterized in that: The online model updating module monitors data distribution changes through a concept drift detection algorithm, including: Construct sliding window comparison statistics and calculate the difference in data distribution between windows based on the chi-square test and Wasserstein distance; An integrated detection strategy is adopted, which combines KS test, mean shift detection and covariance matrix trace change joint decision mechanism; When drift is detected, local subtree reconstruction is triggered, which retains the unaffected branches in the original tree and only performs recursive splitting based on the Gini gain on the drift-related paths.
8. The method for establishing a decision tree data model according to claim 7, characterized in that: The differential privacy protection mechanism injects noise into the reconstruction node including: Define node partition sensitivity as the maximum difference in leaf node category counts, and generate noise that satisfies ε-differential privacy based on the Laplace mechanism; Adopting an adaptive noise allocation strategy, the noise intensity is adjusted according to the node level depth and data sparsity. The amount of noise injected into deep nodes decays exponentially with the depth. The leaf node probabilities after noise injection are normalized to ensure that the sum of the category probabilities is 1 and the distribution is smooth.
9. The method for establishing a decision tree data model according to claim 8, characterized in that: The graph embedding algorithm performs vectorization processing on graph structure data including: A heterogeneous information network representation learning method is adopted to generate node sequences based on a random walk strategy guided by meta-paths. A hierarchical attention mechanism is used to aggregate neighbor node features, where the structural attention weight is determined by both edge type and node degree; The embedding representation is optimized through contrastive learning with negative sampling, maximizing the mutual information of positive sample pairs and minimizing the similarity of negative sample pairs.
10. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 9.
Citation Information
Cited By
Compiler collaborative optimization method and system based on Wasserstein distance and Bayesian Thompson sampling
CN120491945A
Online incremental learning method and system based on adaptive B + tree index
CN120763364A
Segmented modeling method and system for intake and exhaust pressure adjusting device
CN121031390A
New and old battery mixing station thermoelectric coordination regulation system
CN121216676A
A new and old battery hybrid station heat and electricity co-regulation system
CN121216676B