Data feature dimension reduction selection method based on dynamic game
Through the dynamic game framework combining Shapley values and hybrid core HSIC, the balance problem of contribution and redundancy in feature selection is solved, and the calculation complexity is reduced through the dynamic bidding-elimination strategy, achieving efficient feature dimensionality reduction selection.
Patent Information
- Application Number
- CN202510294264.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-13
AI Technical Summary
The existing feature selection methods are difficult to balance the individual prediction contribution of features with group redundancy, and the computational complexity in large-scale data scenarios limits its application.
The data feature dimensionality reduction selection method based on dynamic game is used to calculate the individual contribution of the features through Shapley values, and the mixed core HSIC evaluates the linear and nonlinear redundancy between features, and gradually screens out the feature set with high contribution and low redundancy through dynamic bidding-elimination games.
It effectively solves the trade-off between feature contribution and redundancy, reduces the complexity of HSIC computing, and significantly improves the applicability and efficiency of the method in large-scale data.
Smart Images

Figure CN120144986A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of machine learning and data mining, and specifically relates to a method for data feature dimensionality reduction and selection based on dynamic game. Background Art
[0002] Feature selection, as a key preprocessing step in the field of machine learning, aims to improve the generalization ability and computational efficiency of the model by screening out a feature subset with high contribution and low redundancy. However, traditional feature selection methods often face the following two major challenges: First, existing methods often have difficulty balancing the individual prediction contribution and group redundancy of features due to their single metric; Second, in large-scale data scenarios, the bottleneck of existing methods in computational complexity limits their large-scale application.
[0003] Existing research mostly evaluates feature importance based on statistical correlation or information theory, but such methods often ignore the interaction effects between features and are difficult to distinguish linear and nonlinear redundancy relationships. The game theory framework provides a theoretical support for quantifying the marginal contribution of features, but its co-optimization mechanism with redundancy measurement is still imperfect; Although the redundancy evaluation based on kernel methods can capture complex dependence relationships, its application in large-scale data is limited due to its extremely high computational complexity.
[0004] Considering these limitations, an optimized feature selection method is needed, which can simultaneously consider the linear and nonlinear relationships between features and accurately perform feature dimensionality reduction. In addition, problems such as excessively high computational complexity, long computational time, and insufficient memory need to be solved to ensure that this feature selection method is both efficient and practical. Summary of the Invention
[0005] Based on the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a method for data feature dimensionality reduction and selection based on dynamic game to solve the above technical problems.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A method for data feature dimensionality reduction and selection based on dynamic game, comprising:
[0007] S1: Input the initial data set that needs to perform feature dimensionality reduction;
[0008] S2: Based on the initial data set, calculate the individual contribution degree of features to the prediction result based on the Shapley value;
[0009] S3: According to the individual contribution degree, obtain the G1 index after normalization through the SHAP framework, which is used to quantify the relative importance of features;
[0010] S4: Use the mixed kernel HSIC to calculate the linear and nonlinear redundancy degrees between features, and obtain the G2 index after normalization by adjusting the linear and nonlinear attention weights;
[0011] S5: For the HSIC calculation process in S4, subsample the large-scale dataset, combine with the in-place centering algorithm, and incrementally calculate the kernel matrix to reduce the HSIC calculation complexity;
[0012] S6: Construct a bi-objective game model according to S3 - S5 and conduct dynamic bidding-elimination games;
[0013] S7: After multiple rounds of selection, obtain the final feature set.
[0014] The present invention is further configured such that the importance of the features in S2 is obtained by weighted averaging the incremental contributions to the prediction results when any feature subset is added;
[0015] For feature i, its Shapley value φ i is defined as: where φ i is the Shapley value of feature i, F is the complete set of features, S is the feature subset, and f S (x represents the prediction result using subset S, and the weight term in the formula ensures that the contributions of all subset combinations are fairly weighted.
[0016] The present invention is further configured such that S3 specifically includes:
[0017] Calculate the feature contribution stability index according to the individual contribution degree;
[0018] After normalizing the feature contribution stability index through the SHAP framework, obtain the G1 index. The calculation logic of the feature contribution stability index is where Ω i is the feature contribution stability index of feature i, K is the number of subset partitions, is the Shapley value of feature i on the k-th subset, is the average contribution of feature i on all subsets, is the local variance of feature i on the k-th subset, and λ is the non-linear adjustment index.
[0019] The present invention is further configured such that the hybrid kernel HSIC calculation logic is: K mix =γK linear +(1 - γ)K RBF , and at the same time fuse the linear kernel K linear =XX T to capture linear correlations and the RBF kernel K RBF (x,y)=exp(-σ∥x - y∥ 2) is used to capture nonlinear correlations and adjust the linear and nonlinear attention by adjusting the parameter γ.
[0020] The present invention is further configured that, in S3 and S4, feature scaling and normalization processing is performed on the G1 index and the G2 index, specifically: Among them, G is the original data, G max and G min are the maximum and minimum values of the feature in the data set, G norm is the normalized value.
[0021] The present invention is further configured that S5 specifically includes the following steps:
[0022] S51: By randomly selecting min(m, N) samples, a comparison chart of HSIC calculation time and accuracy is drawn, where m is the preset sampling scale and N is the total number of samples;
[0023] S52: Through the in-place centralization algorithm, incremental calculation replaces full matrix storage, releases memory immediately after each calculation, and processes feature pairs in batches to avoid loading all feature data at the same time;
[0024] The centralization operation of the in-place centralization algorithm is defined as K c =HKH, Subtract the mean of each column to eliminate column-wise bias: Subtract the mean of each row to eliminate row-wise bias: Compensating for overcorrections caused by dual centralization: Get K c =HKH.
[0025] The present invention is further configured that S6 specifically includes the following steps:
[0026] S61: First, the feature set F = {i 1 ,i 2 ,...,i p All participate in the initial bidding, and the bidding price of all candidate features in each iteration is defined as Among them, G1 i Represents the G1 index of feature i. The selected feature set is initially an empty set. λ is the penalty coefficient, which is used to adjust the balance between importance and redundancy. The larger λ is, the stronger the redundancy penalty is. It is determined by cross-validation to balance importance and redundancy. It is the cumulative sum of HSIC of feature i and all selected features, indicating the difference between feature i and the selected feature set S selected Total redundancy of
[0027] S62: After calculating the bid prices of all features in the feature set F in each round, a bidding is conducted and the feature with the highest bid price is added to the selected set: Update the set and add this winning feature i * Add selected feature set S selected , and remove this feature from the feature set F. This freezing mechanism allows the selected feature i * The bid price is no longer updated to avoid repeated selection of this feature in subsequent rounds;
[0028] S63: When the highest bid price in a round is not greater than 0, and there is no feature that meets the screening requirements, all features including this feature are eliminated.
[0029] The present invention provides a data feature dimension reduction selection method based on dynamic game, and the method comprises the following steps:
[0030] S1: Input the initial data set that needs to be reduced in dimension;
[0031] S2: Based on the initial data set, calculate the individual contribution of features to the prediction results based on the Shapley value
[0032] S3: According to the individual contribution, the G1 index is obtained after normalization through the SHAP framework to quantify the relative importance of the features;
[0033] S4: The hybrid kernel HSIC is used to calculate the linear and nonlinear redundancy between features, and the linear and nonlinear attention weights are adjusted by parameters. The G2 index is obtained after normalization.
[0034] S5: For the HSIC calculation process of S4, sub-sampling is performed on large-scale data sets, combined with the local centralization algorithm, and the kernel matrix is incrementally calculated to reduce the HSIC calculation complexity;
[0035] S6: Construct a dual-objective game model based on S3-S5 and conduct a dynamic bidding-elimination game;
[0036] S7: After multiple rounds of selection, the final feature set is obtained;
[0037] The beneficial effects include: constructing a dual-objective game framework through Shapley value and hybrid kernel HSIC to solve the trade-off between feature contribution and redundancy. Shapley value is based on cooperative game theory to quantify the individual contribution of features to prediction and ensure fair distribution of importance; hybrid kernel HSIC integrates linear and nonlinear kernel functions to comprehensively evaluate the redundant relationship between features.
[0038] To address the problem of high computational complexity of HSIC, a subsampling technique is proposed: randomly select a fixed-size sample m, which significantly reduces the computational complexity. At the same time, an in-place centering algorithm is introduced to complete the centering of the column and row means of the kernel matrix and error compensation step by step, replacing traditional matrix storage, and the memory occupancy is only one-third. The two techniques cooperate to solve the memory and computational bottlenecks of the kernel method under large-scale data, and significantly improve the applicability of the method in edge devices and high-dimensional scenarios.
[0039] Redundant features are gradually eliminated through a dynamic bidding-elimination strategy. The bid price of each round of features is determined by subtracting the cumulative sum of HSIC redundancy with the selected features from the Shapley value, and the adaptive factor λ balances the weights. Initially, features with high contributions are preferentially selected. Later, the redundancy penalty increases as the number of selected features increases, and the bid price of redundant features continues to decline until they are eliminated. This strategy avoids the deviation of static thresholds and achieves a dynamic balance between contribution and redundancy.
[0040] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the following specifically gives the specific implementation manners of this application. Brief Description of the Drawings
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings. In the drawings:
[0042] Figure 1 It is a flowchart of a method for dimensionality reduction and selection of data features based on dynamic game in an embodiment of the present invention;
[0043] Figure 2 It is a comparison chart of the importance and redundancy of some measured data features in an embodiment of the present invention;
[0044] Figure 3 It is a comparison chart of the HSIC calculation time and accuracy of some measured data in an embodiment of the present invention;
[0045] Figure 4 It is a bubble chart of the feature bidding process of some measured data in an embodiment of the present invention;
[0046] Figure 5 It is a schematic block diagram of the specific process of dynamic game in an embodiment of the present invention;
[0047] Figure 6 It is a line chart comparing the accuracy of cross-model feature sets of some measured data in an embodiment of the present invention. Detailed Implementation Modes
[0048] The following will illustrate the implementation modes of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation modes, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention, rather than for limiting the protection scope of the present invention.
[0049] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0050] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0051] Embodiment 1
[0052] A data feature dimensionality reduction and selection method based on dynamic game, as Figure 1 shown, includes:
[0053] S1: Input the initial data set that needs to be dimensionally reduced.
[0054] The data set used in S1 is the non-intrusive electric bicycle charging load data collected in the field and preprocessed. There are a total of 20,000 samples and 14 initial features (current I1, active power P1, reactive power Q1, voltage U1, 1-9th harmonics x1_n, and total harmonic distortion THDi). Since a non-intrusive load identification system for electric bicycle charging needs to be deployed on edge devices, feature dimensionality reduction is required to save computing power.
[0055] S2: Based on the initial data set, calculate the individual contribution degree of features to the prediction result based on the Shapley value.
[0056] In feature selection, each feature is regarded as a game participant, and its Shapley value reflects the marginal contribution of this feature to the model prediction. The core idea is: The importance of a feature is obtained by weighted averaging the incremental contribution to the prediction result when it is added to any feature subset.
[0057] For feature i, its Shapley value φ i is defined as: where φ i is the Shapley value of feature i, F is the complete set of features, S is a subset of features, and f S (x represents the prediction result using subset S, and the weight term in the formula ensures that the contributions of all subset combinations are fairly weighted.
[0058] The complexity of directly calculating the Shapley value is O(2 |F| ), and an approximation algorithm needs to be used. The present invention adopts the SHAP (SHapley Additive exPlanations) framework. Based on the additive property of the LightGBM model, the Shapley value is efficiently calculated through tree path tracing. After normalization, G1 i ∈[0,1], eliminating the dimension difference and intuitively reflecting the relative importance of features.
[0059] S4: Calculate the linear and non-linear redundancy between features using the hybrid kernel HSIC, and obtain the G2 index after normalization by adjusting the linear and non-linear attention weights;
[0060] HSIC is a statistical method for measuring the independence between two variables. Its core idea is to map the data to a high-dimensional space through a kernel function, so as to capture the non-linear relationship between features. The traditional HSIC is defined as: HSIC(X,Y) = tr(KHLH) / n 2 , where K and L are the kernel matrices of features X and Y respectively, is the centering matrix, e is a vector of all 1s. The embodiment of the present invention adopts a hybrid kernel function to enhance adaptability: K mix = γK linear +(1 - γ)K RBF , and at the same time fuses the linear kernel K linear = XX T to capture linear correlation and the RBF kernel K RBF (x,y) = exp(-σ∥x - y∥ 2 ) to capture non-linear correlation, and adjusts the linear and non-linear attention degrees by adjusting the parameter γ. When γ ∈ [0,1] is larger, it means more attention to linear correlation, and when it is smaller, more attention is paid to non-linear correlation.
[0061] The normalization method used in the embodiments of the present invention is "Min - Max Normalization", also known as "feature scaling". It is a method of scaling the range of data features to a specified maximum and minimum value (usually 1 and 0), which is represented by the following formula: where G is the original data, G max and G min are the maximum and minimum values of the feature in the dataset respectively, and G norm is the value after normalization.
[0062] As Figure 2 shown, the G1 contribution degree and G2 redundancy degree of each feature are calculated. It can be preliminarily observed from the figure the average importance and independence degree of each feature. When the G1 value of feature i is larger, it represents greater importance of the feature, while when the G2 value of feature i is larger, it indicates worse comprehensive independence, but this does not show the redundancy degree for each feature. Therefore, further dynamic bidding and elimination are required.
[0063] S5: For the HSIC calculation process in S4, subsample the large - scale dataset, and combine with the in - place centering algorithm to incrementally calculate the kernel matrix, reducing the HSIC calculation complexity;
[0064] By randomly selecting min(m, N) samples, where m is the preset sampling scale and N is the total number of samples, to balance the calculation accuracy and efficiency. When dealing with a large - scale dataset, since the space complexity of the kernel method is O(N 2 ), for relatively large data points, the matrix will be very large and occupy a large amount of memory. Using the subsampling technique, the calculation complexity is reduced from O(N 2 ) to O(m 2 ), where m is the subsample size, and the memory occupation is reduced by approximately (N / m) 2 times. In the embodiments of the present invention, when N = 2 * 10 4 and m = 2 * 10 3 , it is reduced by 2 orders of magnitude.
[0065] As Figure 3As shown, for the HSIC calculation time and accuracy trade - off, the horizontal axis is the sample size (logarithmic scale), the left vertical axis is the calculation time (logarithmic scale), and the right vertical axis is the accuracy (1 - relative error). The solid line is the time trend, and the dashed line is the accuracy trend. From the figure, we can see that the HSIC calculation time decreases as the number of subsamples decreases, but the finite - sample estimation error of the kernel matrix increases as m decreases, resulting in a systematic deviation of the estimated value from the true value. Therefore, it is necessary to select an appropriate subsampling scale m to balance the sampling time and calculation accuracy. In the embodiment of the present invention, m = 2000 is selected. When the data volume exceeds 10000, under the condition of significantly reducing the calculation cost, the HSIC calculation error rate is not higher than 5%.
[0066] The embodiment of the present invention introduces an in - place centering algorithm. Incremental calculation is used instead of full - matrix storage, and the memory is released immediately after each calculation. Feature pairs are processed in batches to avoid loading all feature data at the same time.
[0067] The centering operation of traditional HSIC is defined as K c = HKH, where this algorithm realizes the equivalent operation through the following three steps:
[0068] 1. Subtract the mean of each column to eliminate the column - direction deviation:
[0069] 2. Subtract the mean of each row to eliminate the row - direction deviation:
[0070] 3. Compensate for the over - correction caused by double centering:
[0071] Finally, we still get K c = HKH. However, the memory occupation is only one - third of the traditional method, showing significant improvement in large - scale kernel matrix processing and memory - sensitive tasks (embedded devices, large - scale parallel computing).
[0072] S6: Construct a bi - objective game model according to S3 - S5 and conduct a dynamic bidding - elimination game;
[0073] As Figure 4 shown, it is the flow block diagram of the dynamic bidding - elimination game. The dynamic bidding - elimination game strategy gradually screens out high - contribution and low - redundancy features by simulating the competition and elimination mechanism among features. Its core idea is: features compete for the qualification to be selected through the bid price (importance - redundancy penalty), and redundant features are gradually eliminated due to cumulative penalties.
[0074] First, all features in the feature set F = {i 1 , i 2 ,..., i p} participate in the initial bidding. The bid price of all candidate features in each round of iteration is defined as: Among them, G1 iThe G1 index representing feature i, and the initially selected feature set is an empty set. λ is the penalty coefficient used to adjust the balance between importance and redundancy. The larger λ is, the stronger the redundancy penalty. It is determined through cross-validation to balance importance and redundancy. is the cumulative sum of HSIC between feature i and all selected features, representing the total redundancy between feature i and the selected feature set S. selected of the total redundancy.
[0075] After calculating the bid prices of all features in the feature set F in each round, a bid is conducted, and the feature with the highest bid price is added to the selected set: Then the set needs to be updated. The winning feature i * is added to the selected feature set S selected , and this feature is removed from the full feature set F. This freezing mechanism can prevent the bid price of the selected feature i * from being updated again, avoiding repeated selection of this feature in subsequent rounds.
[0076] It can be found from the bidding formula that when bidding for the first time at this time, the bid price Bid i is only related to G1 i , that is, the G1 index of feature i. However, as the number of bidding rounds increases, obviously, due to the continuous expansion of the scale of S selected , the bid prices of highly redundant features will also continuously and dynamically decrease, effectively avoiding local optimal solutions. It achieves the purpose of emphasizing contribution degree in the early stage of feature screening and redundancy degree in the later stage, and well realizes the dynamic balance between importance and redundancy.
[0077] During the process of dynamic bidding, there will be a bid price Bid i for each round. The feature with the highest bid is selected, and this feature will be prohibited from participating in the next bid. Through continuous dynamic bidding, the bid prices of features with high redundancy with the selected features will continuously decline.
[0078] S7: After multiple rounds of selection, the final feature set is obtained. As Figure 5 shown, since different features have different redundancy degrees with the features contained in the selected feature set, the ranking will change in each round of bidding until the bid price of the low-contribution feature drops below 0 and is eliminated. During the selection process, X1_3, I1, X1_9, Q1, X1_7, X1_2 are successively selected in Round1 - Round6, while in Round7, the Bid i of U1 is not greater than 0 and is eliminated, and the bidding stops. After the dynamic bidding - elimination game, the selected feature set S selected = {I1, Q1, X1_2, X1_3, X1_7, X1_9}, and the feature dimension is reduced from 14 dimensions to 6 dimensions.
[0079] As shown Figure 6 in the figure, the selected feature set S selected is marked as the feature set Z0 as the experimental group, and the feature set Z1 consisting of only the 6 highest-scoring features screened by the Shapley value ranking is constructed. The optimal 6 features obtained by the random forest-recursive feature elimination method (RF-RFE) form the feature set Z2, and the mutual information method selects the first 6 features to form the feature set Z3. The current four common models, LightGBM, SVM, KNN, and Random Forest, are used to perform cross-validation on the Z0-Z3 feature sets.
[0080] The experimental results show that the Z0 feature set has the highest average accuracy among the four mainstream models (LightGBM: 0.9708, SVM: 0.9632, KNN: 0.9611, RF: 0.9661), significantly superior to the traditional methods: an average increase of 0.42% compared to the Shapley single-index method (Z1), an average increase of 0.31% compared to RF-RFE (Z2), and an average increase of 0.64% compared to the mutual information method (Z3). Z0 reaches the highest accuracy of 97.08% in the gradient boosting framework (LightGBM), verifying the ability of the features screened in the embodiments of the present invention to capture complex non-linear relationships, indicating that the present invention effectively balances feature importance and linear and non-linear redundancy when performing feature dimensionality reduction.
[0081] In addition, the Shapley value measures the contribution degree of individual features, but it does not consider the stability of features. The feature contribution stability index Ω i is introduced, and its calculation logic is where Ω i is the feature contribution stability index of feature i, K is the number of subset partitions, is the Shapley value of feature i on the k-th subset, is the average contribution of feature i on all subsets, is the local variance of feature i on the k-th subset, and λ is the non-linear adjustment index. According to the feature contribution stability index, the G1 index is obtained after normalization through the SHAP framework, which not only considers the contribution degree but also takes into account the stability and redundancy removal ability, thereby improving the rationality of feature selection.
[0082] The data feature dimensionality reduction selection method based on dynamic game proposed by the present invention effectively solves the deficiencies of traditional methods in feature contribution-redundancy trade-off and computational efficiency through a dual-objective optimization and dynamic game mechanism. In the model framework, the Shapley value quantifies the individual prediction contribution of features, and the mixed kernel HSIC evaluates feature redundancy from multiple perspectives. The combination of the two ensures the global optimality of the feature subset. The introduction of the subsampling technique and the in-place centering algorithm reduces the computational complexity of HSIC from O(N2 ) reduced to O(m 2 ), and the memory occupancy is reduced by approximately (N / m) 2 , significantly improving the applicability in large-scale data. The dynamic bidding-elimination strategy realizes the progressive elimination of redundant features through an adaptive penalty mechanism, avoiding the bias caused by static thresholds.
[0083] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0084] It should be understood that the term "and / or" in this article is only a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context.
[0085] In this application, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following items (pieces)" or its similar expressions refer to any combination of these items, including any combination of single items (pieces) or plural items (pieces). For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0086] It should be understood that in various embodiments of the present application, the sequence numbers of the above processes do not imply the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0087] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0088] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0089] In several embodiments provided by the present application, it should be understood that the disclosed system can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0090] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0091] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0092] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0093] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A data feature dimension reduction selection method based on dynamic game, characterized in that: include: S1: Input the initial data set that needs to be reduced in dimension; S2: Based on the initial data set, the individual contribution of the feature to the prediction result is calculated based on the Shapley value; S3: According to the individual contribution, the G1 index is obtained after normalization through the SHAP framework to quantify the relative importance of the features; S4: The hybrid kernel HSIC is used to calculate the linear and nonlinear redundancy between features, and the linear and nonlinear attention weights are adjusted by parameters. The G2 index is obtained after normalization. S5: For the HSIC calculation process of S4, sub-sampling is performed on large-scale data sets, combined with the local centralization algorithm, and the kernel matrix is incrementally calculated to reduce the HSIC calculation complexity; S6: Construct a dual-objective game model based on S3-S5 and conduct a dynamic bidding-elimination game; S7: After multiple rounds of selection, the final feature set is obtained.
2. A data feature dimension reduction selection method based on dynamic game according to claim 1, characterized in that: The importance of features in S2 is obtained by weighted average of their incremental contributions to the prediction results when added to any feature subset; For feature i, its Shapley value φ i Defined as: Among them, φ i is the Shapley value of feature i, F is the full set of features, S is the feature subset, f S (x represents the prediction result using subset S, and the weight term in the formula Ensure that the contributions of all subset combinations are weighted fairly.
3. The method for selecting data feature dimensionality reduction based on dynamic game according to claim 2, characterized in that: S3 specifically includes: Calculate the characteristic contribution stability index based on individual contribution; The feature contribution stability index is normalized by the SHAP framework to obtain the G1 index. The calculation logic of the feature contribution stability index is: Among them, Ω i is the feature contribution stability index of feature i, K is the number of subset divisions, is the Shapley value of feature i on the kth subset, is the average contribution of feature i on all subsets, is the local variance of feature i on the kth subset, and λ is the nonlinear adjustment index.
4. The method for selecting data feature dimensionality reduction based on dynamic game according to claim 1, characterized in that: The hybrid core HSIC calculation logic is: K mix =γK linear +(1-γ)K RBF , while integrating the linear kernel K linear =XX T Used to capture linear correlation and RBF kernel K RBF (x,y)=exp(-σ∥xy∥ 2 ) is used to capture nonlinear correlations and adjust the linear and nonlinear attention by adjusting the parameter γ.
5. The method for selecting data feature dimensionality reduction based on dynamic game according to claim 1, characterized in that: In S3 and S4, feature scaling and normalization processing is performed on the G1 index and the G2 index, specifically: Among them, G is the original data, G max and G min are the maximum and minimum values of the feature in the data set, G norm is the normalized value.
6. The method for selecting data feature dimensionality reduction based on dynamic game according to claim 1, characterized in that: S5 specifically includes the following steps: S51: By randomly selecting min(m, N) samples, a comparison chart of HSIC calculation time and accuracy is drawn, where m is the preset sampling scale and N is the total number of samples; S52: Through the in-place centralization algorithm, incremental calculation replaces full matrix storage, releases memory immediately after each calculation, and processes feature pairs in batches to avoid loading all feature data at the same time; Among them, the centralized operation of HSIC is defined as K c =HKH, Subtract the mean of each column to eliminate column-wise bias: Subtract the mean of each row to eliminate row-wise bias: Compensating for overcorrections caused by dual centralization: Get K c =HKH.
7. The method for selecting data feature dimensionality reduction based on dynamic game according to claim 5, characterized in that: S6 specifically includes the following steps: S61: First, the feature set F = {i1, i2, ..., i p All participate in the initial bidding, and the bidding price of all candidate features in each iteration is defined as Among them, G1 i Represents the G1 index of feature i. The selected feature set is initially an empty set. λ is the penalty coefficient, which is used to adjust the balance between importance and redundancy. The larger λ is, the stronger the redundancy penalty is. It is determined by cross-validation to balance importance and redundancy. It is the cumulative sum of HSIC of feature i and all selected features, indicating the difference between feature i and the selected feature set S selected Total redundancy of S62: After calculating the bid prices of all features in the feature set F in each round, a bidding is conducted and the feature with the highest bid price is added to the selected set: Update the set and add this winning feature i * Add selected feature set S selected , and remove this feature from the feature set F. This freezing mechanism allows the selected feature i * The bid price is no longer updated to avoid repeated selection of this feature in subsequent rounds; S63: When the highest bid price in a round is not greater than 0, and there is no feature that meets the screening requirements, all features including this feature are eliminated.
Citation Information
Cited By
Multi-modal sentiment classification method based on dynamic game strategy
CN120429602A