Ant colony feature selection method and device based on reinforcement learning, equipment and medium

By introducing reinforcement learning into the traditional ant colony optimization feature selection method, dynamically update pheromone information, and comprehensively evaluate the importance of characteristics, the problem of high randomness and easy to fall into local optimality in the traditional method is solved, and efficient and accurate feature selection is achieved.

CN120196920APending Publication Date: 2025-06-24CHINA TELECOM CLOUD TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411783356.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-05
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The traditional ant colony optimization feature selection method has strong randomness in the global feature search space, is prone to falling into local optimality, and is highly complex when evaluating feature importance, and has a long calculation time, making it difficult to maintain a high accuracy while selecting fewer features.

Method used

The ant colony feature selection method based on reinforcement learning is adopted, and the original feature space dimension is reduced by obtaining the target sample set and performing initial feature screening; ants are regarded as agents, and the heuristic information update method is redefined using reinforcement learning algorithms, and state transfer is combined with greed and random strategies to dynamically update pheromone information; in the process of feature selection, the importance of candidate features is comprehensively evaluated from the perspective of multi-objective optimization.

Benefits of technology

Effectively reduce the data dimension, reduce the initial randomness of the ant colony search process, avoid local optimization, and improve the accuracy and efficiency of feature selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196920A_ABST
    Figure CN120196920A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine learning, and discloses an ant colony feature selection method and device based on reinforcement learning, equipment and a medium, and the method comprises the steps: obtaining a target sample set needing feature selection, and carrying out the feature preliminary screening, and obtaining a sample initial feature set; and obtaining a Pareto optimal solution through a heuristic strategy, performing heuristic information updating on feature nodes in the sample initial feature set, and accumulating pheromone information on the feature nodes to obtain optimal heuristic information and optimal pheromone information, thereby performing feature selection on the sample initial feature set to obtain an optimal feature subset. According to the method, feature preliminary screening is utilized, the original feature space dimension is reduced, a static heuristic information updating process in a traditional ant colony feature selection algorithm is converted into a dynamic heuristic learning process in combination with reinforcement learning, the data dimension can be effectively reduced, and high classification accuracy can be kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and particularly to an ant colony feature selection method, device, equipment and medium based on reinforcement learning. Background Art

[0002] In recent years, in machine learning applications, data has shown an explosive growth in both the number of samples and dimensions. The irrelevance and redundancy of data features will not only reduce the accuracy of classification, but also affect the convergence speed of the learning model, thus posing challenges for obtaining effective information. Taking feature selection in the field of text mining as an example, when the feature selection method in the prior art performs feature selection on a text data set, a hybrid feature selection method such as the ant colony algorithm is usually adopted. However, the traditional ant colony optimization feature selection method conducts an initial search in the global feature search space, with strong randomness; in the process of updating heuristic information, a single and static measurement standard is adopted, and only past information is utilized in the state transition rule, making the algorithm prone to falling into local optimality and having a low accuracy; when presetting the next candidate feature, the performance of the classifier is directly used to evaluate the feature importance, with a high complexity and a long calculation time.

[0003] Therefore, there is an urgent need for a feature selection method that can overcome the deficiencies of the traditional ant colony optimization feature selection method and can maintain a high accuracy while selecting fewer features. Summary of the Invention

[0004] In view of this, the present application provides an ant colony feature selection method, device, equipment and medium based on reinforcement learning, which can overcome the deficiencies of the traditional ant colony optimization feature selection method and maintain a high accuracy while selecting fewer features. The technical solution is as follows.

[0005] In the first aspect, the present invention provides a feature selection method based on reinforcement learning, and the method includes:

[0006] Obtain a target sample set that needs to perform feature selection and conduct an initial feature screening to obtain an initial sample feature set;

[0007] Regarding each ant as the current intelligent agent, the initial sample feature set as the state, each ant's selection of the currently unvisited feature as the action, and the similarity between the candidate feature and the category as the reward, and starting from the perspective of multi-objective problems through a heuristic strategy, obtain the Pareto optimal solution;

[0008] Based on the Pareto optimal solution, perform heuristic information update on the feature nodes in the initial sample feature set and accumulate pheromone information on the feature nodes to obtain the optimal heuristic information and the optimal pheromone information;

[0009] Based on the optimal heuristic information and the optimal pheromone information, perform feature selection on the initial sample feature set to obtain an optimal feature subset.

[0010] In an alternative embodiment, obtaining the initial sample feature set includes: calculating the feature weights of the target sample set through the ReliefF algorithm, sorting each feature node in the target sample set according to the weights to obtain a feature sorted sample set; screening the feature sorted sample set according to a set weight threshold to obtain the initial sample feature set.

[0011] In an alternative embodiment, updating the heuristic information of the feature nodes in the initial sample feature set based on the Pareto optimal solution and accumulating pheromone information on the feature nodes includes: randomly placing ants on the feature nodes of the initial sample feature set, and sequentially updating the heuristic information of the feature nodes and accumulating the pheromone information of each feature node.

[0012] In an alternative embodiment, sequentially updating the heuristic information of the feature nodes includes: selecting the next node of the current feature node for heuristic information update according to the state transition rule; the state transition rule is set according to the greedy strategy and the random strategy of the reinforcement learning algorithm.

[0013] In an alternative embodiment, obtaining the optimal heuristic information and the optimal pheromone information includes: when all ants in the ant colony algorithm have completed the update of the heuristic information and the pheromone information, obtaining the respective feature subsets corresponding to each ant; evaluating the respective feature subsets through a fitness function to obtain the optimal feature subset.

[0014] In an alternative embodiment, the constraint condition for obtaining the Pareto optimal solution is:

[0015]

[0016] where X represents the candidate feature vector; Z is the feature vector selected by the last action in the state S t ; S t represents the state of the environment at time t, that is, the selected feature subset at this time; Y represents the sample label vector; n is the number of samples; Correlation(X,Z) is used to calculate the redundancy between two feature vectors; Cosine(X,Y) cosine similarity is used to measure the correlation degree between the feature and the label.

[0017] The ant colony feature selection method based on reinforcement learning provided by the present invention has the following advantages.

[0018] The ant colony feature selection method based on reinforcement learning of the present invention first obtains a target sample set that needs to perform feature selection, and performs feature screening on the target sample set to reduce the dimension of the original feature space, and screens out better features from it, thereby reducing the dimension of the initial search space of the ant colony. Feature screening can use the ReliefF algorithm to sort the features of the target sample set according to weights, and perform feature screening based on the sorted features, so as to screen out a feature sample set, and calculate the similarity matrix between features and the correlation vector between each feature and the label. Subsequently, ants are randomly placed on the features of the feature sample set, and the heuristic information update method of the ant colony algorithm is redefined in combination with reinforcement learning. Specifically, the current selection situation of the feature subset in the ant colony algorithm is used as the state, the selection of the current unvisited feature by each ant is used as the action, and the similarity between the candidate feature and the category is used as the reward, and the original static heuristic function is transformed into a dynamic heuristic learning process, and the candidate features in the feature sample set are updated in turn to complete the update of the heuristic information matrix, and the pheromone information is accumulated during the update process, so as to complete the update of the pheromone matrix. When the entire ant colony completes feature selection, the classification error rate and the length of the selected feature subset are used as comprehensive evaluation criteria, and the fitness function is used to evaluate each obtained feature subset, so as to obtain the optimal feature subset. Using feature pre-screening to reduce the dimension of the original feature space to reduce the initial randomness of the ant colony search process; combining reinforcement learning to transform the static heuristic information update process in the traditional ant colony feature selection algorithm into a dynamic heuristic learning process, and using a state transition rule that combines greed and randomness to balance the relationship between exploration and exploitation in the ant colony search process; in the process of screening candidate features, from the perspective of multi-objective optimization problems, using two constraint indicators, the correlation between the candidate feature and the category and the redundancy with the selected features, to comprehensively evaluate the importance of the candidate feature, so as to complete the selection process of the feature subset, effectively reducing the data dimension and maintaining a high classification accuracy.

[0019] In a second aspect, the present invention provides a feature selection device based on reinforcement learning, and the device includes:

[0020] A sample pre-screening module, configured to obtain a target sample set that needs to perform feature selection and perform feature pre-screening to obtain an initial sample feature set;

[0021] An optimal solution acquisition module, configured to use each ant as the current agent, the initial sample feature set as the state, the selection of the current unvisited feature by each ant as the action, and the similarity between the candidate feature and the category as the reward, and obtain the Pareto optimal solution from the perspective of multi-objective problems through a heuristic strategy;

[0022] An information update module, which heuristically updates the feature nodes in the initial sample feature set based on the Pareto optimal solution and accumulates pheromone information on the feature nodes to obtain optimal heuristic information and optimal pheromone information;

[0023] A feature selection module, which performs feature selection on the initial sample feature set based on the optimal heuristic information and optimal pheromone information to obtain an optimal feature subset.

[0024] In an alternative embodiment, the feature screening module is specifically configured to: update the feature weights of the target sample set through the ReliefF algorithm, sort the feature nodes in the target sample set according to the weights to obtain a feature sorted sample set; and screen the feature sorted sample set according to a set weight threshold to obtain the initial sample feature set.

[0025] In an alternative embodiment, the information update module is specifically configured to: randomly place ants on the feature nodes of the initial sample feature set, sequentially perform heuristic information updates on the feature nodes, and accumulate the pheromone information of each feature node.

[0026] In an alternative embodiment, the information update module is further configured to: select the next node of the current feature node for heuristic information update according to the state transition rule; the state transition rule is set according to the greedy strategy and random strategy of the reinforcement learning algorithm.

[0027] In an alternative embodiment, the information update module is further configured to: when all the ants in the ant colony algorithm have completed the heuristic information and pheromone information updates, obtain the respective feature subsets corresponding to each ant; evaluate the respective feature subsets through a fitness function to obtain the optimal feature subset, and use the heuristic information and pheromone information of the optimal feature subset as the optimal heuristic information and the optimal pheromone information.

[0028] In a third aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the feature selection method based on reinforcement learning according to the first aspect or any corresponding embodiment thereof.

[0029] In a fourth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the feature selection method based on reinforcement learning according to the first aspect or any corresponding embodiment thereof.

[0030] Fifth aspect, the present invention provides a computer program product, including computer instructions for causing a computer to execute the reinforcement learning-based feature selection method according to the first aspect or any corresponding embodiment thereof as described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0032] Figure 1 is a flowchart showing a reinforcement learning-based ant colony feature selection method provided by an embodiment of the present invention.

[0033] Figure 2 is a method flowchart showing a reinforcement learning-based ant colony feature selection method according to an exemplary embodiment.

[0034] Figure 3 is a structural schematic diagram of a reinforcement learning-based ant colony feature selection device provided by an embodiment of the present application.

[0035] Figure 4 is a structural schematic diagram of a computer device provided by an optional embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0037] It should be understood that the "indication" mentioned in the embodiments of the present application can be a direct indication, an indirect indication, or a representation of an associated relationship. For example, A indicates B, which can mean that A directly indicates B. For example, B can be obtained through A; it can also mean that A indirectly indicates B. For example, A indicates C, and B can be obtained through C; it can also mean that there is an associated relationship between A and B.

[0038] In the description of the embodiments of the present application, the term "corresponding" can represent a direct or indirect corresponding relationship between two, can also represent an associated relationship between two, can also be an indication and being indicated, a configuration and being configured, etc.

[0039] In the embodiments of the present application, "pre - definition" can be implemented by pre - saving corresponding codes, tables or other ways that can be used to indicate relevant information in a device (for example, including terminal devices and network devices). The present application does not limit its specific implementation manner.

[0040] In recent years, in machine learning applications such as text mining, computer vision, and biomedicine, data has shown explosive growth in both the number of samples and dimensions. The irrelevance and redundancy of data features will not only reduce the accuracy of classification, but also affect the convergence speed of the learning model, thus posing challenges to the acquisition of effective information. Therefore, using feature selection methods to select highly representative features from the initial feature space and reduce the data dimension is crucial for solving the above problems.

[0041] According to different feature subset evaluation criteria and methods, feature selection methods can be classified into filter - based (Filter), wrapper - based (Wrapper), and embedded (Embedded). In the process of using the filter - based feature selection method for feature selection, only the essential characteristics between data are used to sort the features, and no specific learning algorithm operations are performed. Therefore, the relationships between features are not considered, resulting in error in the results. The wrapper - based feature selection method uses a prediction model to score the feature set. For each feature subset, training and prediction are performed, and a better feature subset is selected according to the prediction error results. Therefore, when verifying the accuracy of the selected feature subset, the wrapper - based method will probably have better performance than the filter - based method, but the calculation speed is slower. Considering the complementarity of the two types of feature selection methods, researchers naturally combine the filter - based and wrapper - based methods to obtain a hybrid feature selection method, which not only ensures the running time of the algorithm but also ensures that the selected feature subset has a small prediction error on the training model. At the same time, more and more scholars comprehensively consider the better global search ability and flexibility of swarm intelligence algorithms and combine them with the filter - based feature selection method. Among them, the ant colony algorithm has been used by many scholars due to its flexible and easy - to - use discrete expression.

[0042] However, the traditional ant colony optimization feature selection method conducts an initial search in the global feature search space, with strong randomness; in the process of updating heuristic information, a single and static measurement standard is adopted, and only past information is used in the state transition rule, making the algorithm easily fall into local optimality; when presetting the next candidate feature, most algorithms directly use the classifier performance to evaluate the feature importance, with high complexity and long calculation time.

[0043] Therefore, the embodiment of the present invention provides an ant colony feature selection method based on reinforcement learning, which combines reinforcement learning to transform the static heuristic information update process in the traditional ant colony feature selection algorithm into a dynamic heuristic learning process, and uses a state transition rule that combines greed and randomness to balance the relationship between exploration and exploitation in the ant colony search process; in the process of screening candidate features, from the perspective of multi-objective optimization problems, two constraint indicators, namely the correlation between candidate features and categories and the redundancy with the selected features, are used to comprehensively evaluate the importance of candidate features, so as to complete the feature subset selection process and reduce the method complexity.

[0044] The ant colony feature selection method based on reinforcement learning provided by the embodiment of the present invention has a method flow as Figure 1 shown, and includes the following steps.

[0045] S101. Obtain a target sample set that needs to perform feature selection and perform preliminary feature screening to obtain an initial sample feature set;

[0046] Specifically, in step S101, the target sample set usually contains a large number of features, most of which contain redundant information, noise or irrelevant features, and there may be correlations or redundancies between features. Therefore, before performing feature selection, it is necessary to perform preliminary screening on these features to reduce the dimension of the original feature space and facilitate subsequent feature selection.

[0047] Optionally, the data of the target sample set in the embodiment of the present invention can be text data in the field of text mining or gene data in the medical field.

[0048] Optionally, in step S101, the ReliefF algorithm is used to update the feature weights of the target sample set, and each feature node in the target sample set is sorted according to the weights to obtain a feature sorted sample set; the feature sorted sample set is screened according to a set weight threshold to obtain the initial sample feature set.

[0049] In the above steps, the main idea of ReliefF is: divide the target sample set into several training samples. In each training sample set, randomly select a sample R, and then find K nearest neighbor samples (near Hits) from other sample data of the same category as R, and find K nearest neighbor samples (near Misses) from different sample sets of R, and then update the feature weights to obtain the feature sorting result. The update formula for the feature weights is:

[0050]

[0051] Where m is the number of sample samplings; diff(A, R1, R2) represents the difference between sample R1 and sample R2 in feature A, that is, diff(A, R, H) represents the difference between sample R and the same-class sample H in feature A; diff(A, R, M) represents the difference between sample R and the different-class sample M in feature A; C represents a certain category; P(C) is the prior probability of this category; P(Class(R)) is the category proportion of randomly selecting a certain sample; W(A) is the weight corresponding to the current feature A, and this value is updated in each iteration to reflect the importance of feature A; W is the updated feature weight; M j (C) represents j nearest neighbors that are different from the reference sample R in class C; R1[A] and R2[A] represent the values of sample R1 and sample R2 in feature A; max(A) and min(A) are the maximum and minimum values of feature A respectively, which are used to normalize the difference value.

[0052] S102. Take each ant as the current agent, the initial feature set of the sample as the state, each ant's selection of the currently unvisited feature as the action, and the similarity between the candidate feature and the category as the reward. Starting from the perspective of the multi-objective problem through the heuristic strategy, obtain the Pareto optimal solution.

[0053] Specifically, in step S103, combine the reinforcement learning Markov decision process to perform feature space transformation on the traditional ant colony feature selection algorithm, so as to transform the original static heuristic function into a dynamic heuristic learning process, perform dynamic heuristic information update and accumulate pheromone information, and combine the greedy strategy and the random strategy to perform state transition to update the next feature in turn, so as to obtain the optimal heuristic information and the optimal pheromone information.

[0054] In the above steps, take the current selection situation of the feature subset in the ant colony algorithm as the state, each ant's selection of the currently unvisited feature as the action, and the similarity between the candidate feature and the category as the reward. The specific contents of each parameter are as follows.

[0055] The state S of the environment: a set composed of all features of the original data. At time t, the state of the environment is the feature subset S t ∈S.

[0056] The action A of the individual: Each ant's selection of the currently unvisited feature represents a set of actions. At time t, the current state of the environment is S t , and the action taken by the individual is to select a certain candidate feature A t ∈A.

[0057] The state transition function T: In the feature selection problem, it represents adding the currently selected feature to the feature subset corresponding to the search path.

[0058] Reward R of the environment: A function, called the return function (reward function). At time t, when the environmental state is S t , the individual takes action A t ; at time t+1, the individual obtains the reward R corresponding to this action t+1 . Calculate the cosine correlation between the candidate feature and the category as the corresponding return value. The cosine correlation expression is:

[0059]

[0060] In the formula, X represents the candidate feature vector, Y represents the sample label vector, and n is the number of samples.

[0061] Policy π: Represents the basis for the individual to take actions in state S t . With the goal of optimizing the feature relationship, starting from the perspective of multi-objective problems, the Pareto optimal solution (non-dominated solution) corresponding to the current problem is obtained, and the next preferred action is determined based on this. The constraint conditions are as follows.

[0062]

[0063] In the formula, X is the candidate feature vector; Z is the feature vector selected by the last action in state S t ; Y represents the sample label vector; Correlation(X,Z) is used to calculate the redundancy between the two feature vectors; Cosine(X,Y) cosine similarity is used to measure the degree of correlation between the feature and the label.

[0064] Value Function: Use the action value function to describe the value of executing a certain action in a certain state, that is, the calculation of the Q value, so as to complete the update of the Q-Table. The Q value expression is:

[0065] Q(S t ,A t )=(1-α)Q(S t ,A t )+α(R(Q(S t ,A t )+γmax(Q(Q(S t+1 ,A t+1 )));

[0066] In the formula, α∈[0,1] is the learning rate, and γ∈[0,1] is the reward decay coefficient.

[0067] Based on the above various parameters, the dynamic heuristic update specifically includes: after the ant completes the selection process of the current candidate feature, comprehensively considering the correlation between the candidate feature and the category and the reward value brought by the next action of the candidate feature, multi-angle dynamic heuristic information update is performed, and the expression is:

[0068] η[k] = Cosine(X, Y)E + Q - Table ′ [k];

[0069] Among them, η is the heuristic information matrix; k is the candidate feature serial number; X represents the candidate feature vector; Y represents the sample label; Q - Table ′ represents the Q - value matrix corresponding to Q - Table; E represents a unit vector with the same column dimension as Q - Table ′ and all values are 1.

[0070] S103. Perform heuristic information update on the feature nodes in the sample initial feature set based on the Pareto optimal solution and accumulate pheromone information on the feature nodes to obtain the optimal heuristic information and the optimal pheromone information.

[0071] Specifically, in step S103, the ants are randomly placed on the feature nodes of the sample initial feature set, and heuristic information update is sequentially performed on the feature nodes, and the pheromone information of each feature node is accumulated. The sequential heuristic information update of the feature nodes includes: selecting the next node of the current feature node for heuristic information update according to the state transition rule; the state transition rule is set according to the greedy strategy and the random strategy of the reinforcement learning algorithm. The obtaining of the optimal heuristic information and the optimal pheromone information includes: when all the ants in the ant colony algorithm have completed the heuristic information and pheromone information update, the respective feature subsets corresponding to each ant are obtained; the respective feature subsets are evaluated through the fitness function to obtain the optimal feature subset, and the heuristic information and pheromone information of the optimal feature subset are used as the optimal heuristic information and the optimal pheromone information.

[0072] S104. Perform feature selection on the sample initial feature set based on the optimal heuristic information and the optimal pheromone information to obtain the optimal feature subset.

[0073] Specifically, in step S104, based on the obtained optimal heuristic information and pheromone information, the ant colony algorithm is initialized and configured, and feature selection is performed on the feature sample set to obtain the optimal feature subset.

[0074] To better illustrate the above - mentioned ant colony feature selection method based on reinforcement learning, a specific example will be used for illustration below.

[0075] In this example, a Python environment is prepared based on Pycharm, and the experimental dataset is randomly divided into a 70% training set and a 30% test set. Subsequently, the ReliefF algorithm is used for initial feature screening to reduce the dimension of the original feature space, and the similarity matrix between features and the correlation vector between each feature and the label are calculated. Finally, variable initialization is performed according to relevant parameter settings, and the current ants are randomly placed on different initial features, and the heuristic information update process is carried out. The next feature is selected according to the state transition rule until all the ant colonies are updated and reach the maximum number of iterations, and the optimal feature subset is output. After feature selection by the hybrid ant colony feature selection algorithm, the data dimension can be effectively reduced and a high classification accuracy can be maintained.

[0076] The specific process of this example is as Figure 2 shown and includes the following steps.

[0077] Construct an experimental dataset and divide the experimental dataset into a 70% training set and a 30% test set. Use the ReliefF algorithm to perform initial feature screening on the experimental dataset, and calculate the similarity matrix between features and the correlation vector between each feature and the label. Initialize the parameters of the ant colony algorithm, including the number of ants, the initial value of pheromone, the pheromone evaporation coefficient, the importance parameter of the heuristic factor, etc. The setting of these parameters will affect the performance and convergence speed of the algorithm.

[0078] Randomly place the ants on different initial features for heuristic information update, and visit different features in the search space through the state transition rule to complete the construction of the feature subset. Considering the problem that ants are prone to fall into local optimal solutions when selecting feature subsets, in the setting of the state transition rule, a combination of greedy strategy and random strategy is selected to balance the relationship between exploration and exploitation. Among them, the greedy strategy enhances the local search ability of ants. At the same time, the random strategy enables ants to find more solutions and expands its solution space. The specific calculation expression of the state transition is:

[0079]

[0080] In the formula, represents the transition probability of the k-th ant from feature i to feature j at time t; represents the set of feasible features that the k-th ant can access but has not accessed starting from feature i; τ i represents the total amount of pheromone assigned to feature i; η ijDenote the heuristic information between feature i and feature j; q is a uniformly random variable distributed in [0, 1]; q0 is a constant (0 ≤ q0 ≤ 1); α and β respectively specify the relative importance of the pheromone amount and the heuristic information (0 ≤ α ≤ 1, 0 ≤ β ≤ 1). If β = 0, that is, the influence of the heuristic information on the ant colony search path will be completely ignored, and the ant colony will completely depend on the pheromone deposited by previous actions for the next state transition. The two parameters q and q0 define the relative importance between exploration and exploitation of the ant colony. When an ant decides to transfer from feature i to feature j, a random number q will be generated. If q > q0, each feature node will have an equal chance of being selected according to the probability. If q ≤ q0, the current ant will access the best feature according to the maximum pheromone accumulation.

[0081] After the ant colony completes the construction process of the feature subset, the fitness function is used to evaluate the corresponding solution set. The purpose of feature selection is to maintain a high accuracy of the original data set with fewer features. Therefore, the classification error rate and the length of the selected feature subset are used as comprehensive evaluation criteria to calculate the fitness function value corresponding to the current ant search path. The smaller its value, the better the performance of the feature subset corresponding to this path, and it can better meet the fundamental requirements corresponding to the data processing process of feature selection. The expression of the fitness function is:

[0082]

[0083] In the formula, error(G) represents the classification error rate corresponding to the current feature subset; N represents the size of the feature dimension of the original data set.

[0084] After the evaluation, update the global pheromone of the ant colony algorithm. The pheromone update method of the traditional ant colony algorithm is as follows:

[0085]

[0086] In the formula, τ i (t + 1) and τ i (t) respectively represent the pheromone concentrations corresponding to feature i at time t and time t + 1; ρ(0 < ρ < 1) is the pheromone evaporation coefficient; Δτ i represents the pheromone increment value obtained by feature i according to certain evaluation criteria; represents the pheromone deposition concentration on feature i during the iteration of the global best ant g; χ is the pheromone evaporation coefficient of the best ant.

[0087] In the pheromone update process of the method in this example, the size of the feature subset selected by the current ant and its corresponding classification performance are comprehensively considered. The update method is as follows:

[0088]

[0089] Among them, Accuracy(G) represents the classification accuracy corresponding to the current feature subset; G represents the currently selected feature subset; |G| represents the size of the current feature subset.

[0090] After completing the global pheromone update, based on the updated pheromone matrix, the best feature subset of the experimental dataset is output, and the feature selection of the experimental dataset is completed.

[0091] The ant colony feature selection method based on reinforcement learning of the present invention first obtains a target sample set that needs to perform feature selection, and performs feature screening on the target sample set to reduce the dimension of the original feature space, and screens out better features from it, thereby reducing the dimension of the initial search space of the ant colony. Feature screening can use the ReliefF algorithm to sort the features of the target sample set according to weights, and perform feature screening based on the sorted features, so as to screen out a feature sample set, and calculate the similarity matrix between features and the correlation vector between each feature and the label. Subsequently, ants are randomly placed on the features of the feature sample set, and the heuristic information update method of the ant colony algorithm is redefined in combination with reinforcement learning. Specifically, the current selection situation of the feature subset in the ant colony algorithm is used as the state, the selection of each ant of the currently unvisited feature is used as the action, and the similarity between the candidate feature and the class is used as the reward, and the original static heuristic function is transformed into a dynamic heuristic learning process. The candidate features in the feature sample set are updated in turn to complete the update of the heuristic information matrix, and the pheromone information is accumulated during the update process, thereby completing the update of the pheromone matrix. When the entire ant colony completes feature selection, the classification error rate and the length of the selected feature subset are used as comprehensive evaluation criteria, and the fitness function is used to evaluate each obtained feature subset, so as to obtain the optimal feature subset. Using feature pre-screening to reduce the dimension of the original feature space to reduce the initial randomness in the ant colony search process; combining reinforcement learning to transform the static heuristic information update process in the traditional ant colony feature selection algorithm into a dynamic heuristic learning process, and using a state transition rule that combines greed and randomness to balance the relationship between exploration and exploitation in the ant colony search process; in the process of screening candidate features, from the perspective of multi-objective optimization problems, using two constraint indicators, the correlation between the candidate feature and the class and the redundancy with the selected features, to comprehensively evaluate the importance of the candidate feature to complete the selection process of the feature subset, effectively reducing the data dimension and maintaining a high classification accuracy.

[0092] In an embodiment of the present application, an ant colony feature selection device based on reinforcement learning is further provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0093] An embodiment of the present application provides an ant colony feature selection device based on reinforcement learning. Figure 3 FIG. 5 is a schematic structural diagram of an ant colony feature selection device based on reinforcement learning provided by an embodiment of the present application. The device includes:

[0094] A sample pre-screening module 301, configured to obtain a target sample set that needs to perform feature selection and perform feature pre-screening to obtain a sample initial feature set;

[0095] An optimal solution acquisition module 302, configured to use each ant as the current agent, the sample initial feature set as the state, each ant's selection of a currently unvisited feature as an action, and the similarity between the candidate feature and the category as a reward, and obtain a Pareto optimal solution from the perspective of a multi-objective problem through a heuristic strategy;

[0096] An information update module 303, based on the Pareto optimal solution, performs heuristic information update on the feature nodes in the sample initial feature set and accumulates pheromone information on the feature nodes to obtain optimal heuristic information and optimal pheromone information;

[0097] A feature selection module 304, based on the optimal heuristic information and optimal pheromone information, performs feature selection on the sample initial feature set to obtain an optimal feature subset.

[0098] In an optional implementation manner, the feature screening module 301 is specifically configured to: update the feature weights of the target sample set through the ReliefF algorithm, sort the respective feature nodes in the target sample set according to the weights to obtain a feature sorted sample set; and screen the feature sorted sample set according to a set weight threshold to obtain the sample initial feature set.

[0099] In an optional implementation manner, the information update module 303 is specifically configured to: randomly place ants on the feature nodes of the sample initial feature set, and sequentially perform heuristic information update on the feature nodes and accumulate the pheromone information of each feature node.

[0100] In an alternative embodiment, the information update module 303 is further configured to: select the next node of the current feature node according to the state transition rule for heuristic information update; the state transition rule is set according to the greedy strategy and the random strategy of the reinforcement learning algorithm.

[0101] In an alternative embodiment, the information update module 303 is further configured to: when all ants in the ant colony algorithm have completed the update of heuristic information and pheromone information, obtain the respective feature subsets corresponding to each ant; evaluate the respective feature subsets through a fitness function to obtain the optimal feature subset, and use the heuristic information and pheromone information of the optimal feature subset as the optimal heuristic information and the optimal pheromone information.

[0102] The further function descriptions of the above respective modules and units are the same as those in the corresponding foregoing embodiments, and will not be elaborated herein.

[0103] The ant colony feature selection device based on reinforcement learning in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0104] An embodiment of the present invention further provides a computer device having the above Figure 3 shown ant colony feature selection device based on reinforcement learning.

[0105] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present invention. As Figure 4 shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting the components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphic information in a graphical user interface on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Figure 4 In

[0106] The processor 10 may be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 may further include a hardware chip. The above-mentioned hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device may be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.

[0107] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.

[0108] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may further include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0109] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 may further include a combination of the above types of memories.

[0110] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 may be connected through a bus or other means, Figure 4 Taking connection through a bus as an example.

[0111] Embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.

[0112] A part of the present invention can be applied as a computer program product, for example, computer program instructions. When executed by a computer, through the operation of the computer, the methods and / or technical solutions according to the present invention can be called or provided. Those skilled in the art should understand that the forms of existence of computer program instructions in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible by the computer.

[0113] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. An ant colony feature selection method based on reinforcement learning, characterized in that: include: Obtain the target sample set that needs feature selection and perform initial feature screening to obtain the initial sample feature set; Each ant is regarded as the current agent, the initial feature set of the sample is regarded as the state, the feature selected by each ant that has not been visited is regarded as the action, and the similarity between the candidate feature and the category is regarded as the reward. Through the heuristic strategy, the Pareto optimal solution is obtained from the perspective of multi-objective problems. Based on the Pareto optimal solution, heuristic information is updated on the feature nodes in the initial feature set of the sample and pheromone information is accumulated on the feature nodes to obtain optimal heuristic information and optimal pheromone information; Based on the optimal heuristic information and the optimal pheromone information, feature selection is performed on the initial feature set of the sample to obtain an optimal feature subset.

2. The method according to claim 1, characterized in that The obtaining of the sample initial feature set comprises: Calculate the feature weight of the target sample set by using the ReliefF algorithm, and sort each feature node in the target sample set according to the weight to obtain a feature sorted sample set; The feature sorting sample set is screened according to a set weight threshold to obtain the sample initial feature set.

3. The method according to claim 2, characterized in that The step of performing heuristic information updating on the feature nodes in the initial feature set of the sample based on the Pareto optimal solution and accumulating pheromone information on the feature nodes includes: Ants are randomly placed on the feature nodes of the initial feature set of the sample, and the feature nodes are updated heuristically in sequence, and the pheromone information of each feature node is accumulated.

4. The method according to claim 3, characterized in that The step of sequentially updating the characteristic nodes by heuristic information includes: The next node of the current feature node is selected according to the state transfer rule for heuristic information update; the state transfer rule is set according to the greedy strategy and random strategy of the reinforcement learning algorithm.

5. The method according to claim 1, characterized in that The obtaining of the optimal heuristic information and the optimal pheromone information includes: When all ants in the ant colony algorithm have completed the update of heuristic information and pheromone information, the feature subsets corresponding to each ant are obtained; The respective feature subsets are evaluated by a fitness function to obtain the optimal feature subset, and the heuristic information and pheromone information of the optimal feature subset are used as the optimal heuristic information and the optimal pheromone information.

6. The method according to claim 1, characterized in that The constraints for obtaining the Pareto optimal solution are: Where X represents the candidate feature vector; Z is the state S t The feature vector selected by the last action; S t Indicates the state of the environment at time t, that is, the feature subset selected at this time; Y represents the sample label vector; n is the number of samples; Correlation(X,Z) is used to calculate the redundancy between two feature vectors; Cosine(X,Y) cosine similarity is used to measure the degree of correlation between features and labels.

7. An ant colony feature selection device based on reinforcement learning, characterized in that: include: The sample initial screening module is used to obtain the target sample set that needs feature selection and perform initial feature screening to obtain the initial sample feature set; The optimal solution acquisition module is used to take each ant as the current agent, the initial feature set of the sample as the state, the feature that each ant selects that has not been visited as the action, and the similarity between the candidate feature and the category as the reward, and obtain the Pareto optimal solution from the perspective of multi-objective problems through a heuristic strategy; An information updating module, which performs heuristic information updating on feature nodes in the initial feature set of the sample based on the Pareto optimal solution and accumulates pheromone information on the feature nodes to obtain optimal heuristic information and optimal pheromone information; The feature selection module performs feature selection on the sample initial feature set based on the optimal heuristic information and the optimal pheromone information to obtain an optimal feature subset.

8. The device according to claim 7, characterized in that The sample initial screening module is specifically used for: The feature weight of the target sample set is calculated by the ReliefF algorithm, and each feature node in the target sample set is sorted according to the weight to obtain a feature sorted sample set; the feature sorted sample set is screened according to a set weight threshold to obtain the sample initial feature set.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the ant colony feature selection method based on reinforcement learning as described in any one of claims 1 to 6 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the ant colony feature selection method based on reinforcement learning according to any one of claims 1 to 6.