A multi-objective optimization classification method for joint design of feature selection and classifier

Through the multi-objective optimization method designed jointly by feature selection and classifier, the redundant feature problem in high-dimensional data sets is solved, efficient feature selection and classifier optimization are achieved, and classification accuracy and system performance are improved.

CN115661546BActive Publication Date: 2025-07-18XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211400699.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-09
Publication Date
2025-07-18
Estimated Expiration
2042-11-09

AI Technical Summary

Technical Problem

The redundancy and unrelated features in high-dimensional datasets lead to degradation of identification system performance, increased computational complexity and increased memory requirements. The existing feature selection methods are cost-effective and inefficient in computing on high-dimensional datasets.

Method used

A multi-objective optimization method designed jointly by feature selection and classifier is adopted, the data set is divided using hierarchical random technology, and the initial population is generated by a hybrid coding scheme. A multi-objective feature selection model is established through NSGA-II and a hybrid operator, and feature selection and classifier optimization are combined with selective neural network integration.

Benefits of technology

It improves the classification accuracy of high-dimensional data sets, reduces the number of features, simplifies the computational complexity, and improves the performance of the identification system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115661546B_ABST
    Figure CN115661546B_ABST
Patent Text Reader

Abstract

A multi-objective optimization classification method for joint design of feature selection and classifier, comprising the following steps: Step 1: Use the hierarchical random technique to divide the high-dimensional data set for medical diagnosis and grayscale image classification into a training set and a test set, and standardize the features of the training set and the test set; Step 2: Adopt a hybrid coding scheme to encode individuals and generate an initial population; Step 3: Use NSGA-II and a hybrid operator to establish and solve the multi-objective feature selection model; Step 4: Solve to obtain multiple Pareto optimal solutions, each optimal solution contains a feature subset and a designed classifier, and sort the solutions according to the training error; Step 5: Extract the selected features obtained through the solution of the entire model and input them into the classifier, and obtain the final diagnosis result or grayscale image classification result through selective neural network integration. The present invention can solve the engineering technology applications with high-dimensional feature samples and improve the classification accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine learning and intelligent computing, and particularly relates to a multi-objective optimization classification method for joint design of feature selection and classifier. Background Art

[0002] With the advent of the big data era and the continuous development of artificial intelligence, high-dimensional data appears in multiple fields such as pattern recognition, machine learning, and data mining. However, high-dimensional data usually contains a large number of redundant and irrelevant features (attributes), and even contains some derogatory features. Its use may bring a series of problems. For example, an identification system using such data usually leads to a decline in performance, makes prediction unnecessarily complex or overfitting, significantly increases memory requirements, and incurs expensive computational costs.

[0003] Therefore, it is necessary to preprocess high-dimensional data sets, delete irrelevant or excessive redundant features, and reduce the number of features.

[0004] Over the years, many feature selection methods have been proposed one after another, which can be roughly divided into three categories: filter methods, wrapper methods, and embedded / integrated methods.

[0005] The filter-based feature selection algorithm and the learning algorithm are independent of each other. Feature selection is the preprocessing process of the latter, and the learning algorithm is the verification process of the former. The whole process is fast and easy to calculate, but not very effective.

[0006] Different from the filter method, the wrapper method combines the feature selection process with the learning algorithm. It regards the learning algorithm as a "black box", takes the classification accuracy as the evaluation criterion for the feature subset, and uses a search strategy to adjust the subset, and finally obtains an approximate optimal subset. Therefore, it is usually more accurate than the filter method, but the calculation speed is relatively slow, the calculation amount is large, and it is not suitable for large data sets.

[0007] Compared with the filter and wrapper feature selection methods, the embedded method embeds the feature selection algorithm in the learning algorithm, and the feature subset can be obtained when the training process of the classification algorithm ends. The embedded feature selection method can select a suitable feature subset at a lower computational cost, so as to obtain the best effect. The search space of high-dimensional data sets is very large. If a data set contains N features, then there are 2 NFor a large number of possible feature combinations, it is impractical and time-consuming to find the optimal combination among all combinations. Therefore, a large number of studies on feature selection have been carried out using meta-heuristic algorithms with good exploration capabilities. However, there are some drawbacks in directly applying meta-heuristic methods to feature selection. For example, these algorithms were originally designed to optimize real-valued parameters, while feature selection is a binary combinatorial optimization problem. At the same time, due to the larger search space of high-dimensional data, it consumes a large amount of time. Summary of the Invention

[0008] In order to overcome the deficiencies of the above-mentioned prior art, the purpose of the present invention is to provide a multi-objective optimization classification method for joint design of feature selection and classifier. Through this method, engineering and technical applications with high-dimensional feature samples can be solved, the classification accuracy is improved, and it can be used in applications such as medical diagnosis and grayscale image classification.

[0009] In order to achieve the above purpose, the technical solution adopted by the present invention is:

[0010] A multi-objective optimization classification method for joint design of feature selection and classifier, comprising the following steps;

[0011] Step 1: Use the stratified random technique to divide the high-dimensional data set for medical diagnosis and grayscale image classification into a training set and a test set, and standardize the features of the training set and the test set. All subsequent operations are carried out on the standardized data set;

[0012] Step 2: Use a hybrid coding scheme to encode individuals (simulating natural selection, and individuals are randomly generated by encoding) and generate an initial population;

[0013] Step 3: Use NSGA-II and hybrid operators to establish and solve the multi-objective feature selection model;

[0014] Step 4: Solve through a multi-objective evolutionary algorithm with hybrid operators to obtain multiple Pareto optimal solutions. Each optimal solution contains a feature subset and a designed classifier, and sort the solutions according to the training error;

[0015] Step 5: Extract the selected features obtained through the solution of the entire model and input them into the classifier, and obtain the final diagnosis result or grayscale image classification result through selective neural network integration.

[0016] The hybrid coding scheme in step 2 simultaneously performs the tasks of feature selection and classifier optimization;

[0017] The coding scheme of each individual includes the following three parts:

[0018] (1) The first part mainly performs binary encoding on feature selection; the first part can be expressed as θ1 = (1,..., 1,..., 0), and the features with a selected value of 1 are input into the network. On the contrary, those not input into the network have corresponding nodes represented in white. The encoding length of this part is N;

[0019] (2) The second part mainly performs binary encoding on the network structure of the classifier. The second part is expressed as θ2 = (1,..., 1,..., 0,..., 1), where a value of 1 indicates that the corresponding hidden neuron is activated, and a value of 0 indicates the contrary. Similarly, the nodes not activated are represented in white. The encoding length of this part is L;

[0020] (3) In the third part, real - number encoding is performed on the input parameters of the classifier. The input parameters include input weights and hidden - layer biases, which are randomly generated within the interval [-1, 1]. The third part is expressed as θ3 = (ω, b) = (0.8,...,-0.3). The encoding length of this part is (N + 1)×L;

[0021] Therefore, each individual can be represented as θ = (θ1, θ2, θ3), with a length of (N + 2)×L+N.

[0022] The individuals obtained in step 2 are the initial individuals in the randomly generated evolutionary process, and these initial individuals will be evolutionarily cycled and optimized under the multi - objective model in step 3 to obtain multiple final optimal solutions.

[0023] In step 3, a multi - objective feature selection model is established to find the trade - off between classification performance and feature subsets, and multiple solutions are obtained for decision - makers to use;

[0024] The first objective function of the model is the error function commonly used in neural network training, namely the root - mean - square error (RMSE);

[0025]

[0026] Regarding the complexity of the network, the structure of the hidden layer is mainly considered and defined using the following formula,

[0027] where θ 2j is the second part of the encoding scheme, and the number of selected features is calculated using a proportional function;

[0028] where N s is the number of selected features, N is the total number of features. According to the previous discussion, these three objectives are contradictory. Therefore, feature selection is regarded as a problem consisting of three objectives to be minimized simultaneously:

[0029] min{E, H, F}; where E is the training error function, H is the network complexity, and F is the selected feature ratio.

[0030] In step 3, the multi-objective feature selection model established by using NSGA-II with a hybrid operator includes hybrid crossover and hybrid mutation operations;

[0031] Crossover is to select two individuals from the population and exchange parts of the two individuals with a certain probability, and perform the following hybrid crossover operations on the individuals:

[0032] (1) At the beginning of crossover, randomly select a group of individuals from the population (this is a process that loops, and each time the selected individuals are selected from the previous generation, that is, the offspring individuals are obtained by genetic operations on the parent generation), and mark them as parent individuals.

[0033] (2) Divide the parent generation into three parts.

[0034] (3) Perform the operations of the following equations on each part to obtain the corresponding crossover individuals.

[0035]

[0036] where X 1j and X 2j are the selected parents, Y 1j and Y 2j are the offspring, and α j is a random number, randomly selected from {0, 1} when performing binary crossover and randomly selected from [0, 1] when performing real number crossover;

[0037] (4) After completing the crossover operations respectively, recombine the three parts to form new individuals;

[0038] Mutation means replacing some gene values in the individual coding string with other gene values according to the mutation probability to form a new individual, and perform the following hybrid mutation operations on the individual:

[0039] (1) Randomly select a paternal individual and divide the individual into three parts according to the coding method;

[0040] (2) Perform the following mutation operations respectively. When facing binary coding, use the following equations;

[0041]

[0042] where X j is the parent individual, and Z jis the offspring, rand(0,1) is a random number, μ is a control parameter. When dealing with real - number coding, the following equation is used;

[0043]

[0044] (3) After completing the mutation operation respectively, recombine the three parts to form a new individual, and this new individual is used to generate a new population R t .

[0045] The new individual finally formed in step 3 is used to generate a new population. Step 3 is looped. After reaching the maximum number of iterations, the final population obtained can be used in step 4. The final population includes multiple solutions, and each solution includes the selected features and the optimized network input parameters. Step 4 is to apply a non - iterative algorithm to the network input parameters of each solution to solve the corresponding output parameters, so as to obtain the designed network classifier.

[0046] In step 4, the classifier uses a single - hidden - layer neural network that can automatically design the structure and input parameters, and uses a non - iterative algorithm to solve the output weights;

[0047] A single - hidden - layer feed - forward neural network (SLFN) is considered for feature selection, where N, L, and K are the numbers of neurons in the input layer, hidden layer, and output layer respectively;

[0048] The Extreme Learning Machine (ELM) is a single - hidden - layer feed - forward neural network algorithm. ELM needs to preset the number of hidden layer nodes and does not need to adjust the input weights and hidden layer biases during the operation. The output of an ELM with m samples can be defined as where ω j and β j are the input and output weight vectors of the j - th hidden neuron respectively, and b j is the bias of the j - th hidden neuron. g(·) is the activation function of the hidden layer, and the sigmoid function is used;

[0049] Let H=(h ij ) m×L =g(ω j ,b j ,x i ) be the output matrix of the hidden layer, then the solution of ELM can be expressed as follows:

[0050] Hβ = T; where T = [t1,t2,...,t m ) T is the target output matrix. So the output weight β can be calculated by the following formula:

[0051] wherein represents the Moore-Penrose (MP) generalized inverse matrix of H, and is used for the weights and structure of the neural network for classification.

[0052] Through step 4, multiple designed network classifiers and corresponding selected feature subsets are obtained. For these multiple solutions, to obtain the final result, step 5 is used to select some of these solutions for ensemble learning to make the final decision and improve the classification accuracy.

[0053] The said step 5 makes the final decision by using the method of selective neural network ensemble, which improves the classification accuracy. The weights adopt weighted average, and a weight value λ is assigned to each network i , satisfying the following conditions

[0054]

[0055] By establishing a multi-objective feature selection model through step 3, that is, it is necessary to minimize the training error, network complexity, and the proportion of selected features, and use NSGA-II with a hybrid coding genetic operator for optimization and solution. Finally, multiple optimal solutions can be obtained. Each solution respectively includes the selected feature subset and the optimized network input parameters. For the multiple solutions obtained, in order to make the final decision and also to improve the classification accuracy;

[0056] First, all the obtained Pareto optimal solutions are first sorted in ascending order of the training error. Then, the first three solutions are selected, the selected feature subsets are extracted from the test set, and the designed classifiers are optimized respectively. Then, these three classifiers are integrated into base classifiers, and their weights are 0.7, 0.2, and 0.1 in sequence. Finally, the final classification accuracy is obtained according to the following equation;

[0057] where x is the test data, and C i (x) = [C i1 (x), C i2 (x), …, C iK (x)] is the classification result of the i-th classifier, and Class(x) = [C1(x), C2(x), …, C K (x)] is the obtained output result. The category of the final x is given by the following formula, Class(x) = argmax 1≤j≤K Class j (x). The finally obtained result is the category of the test data, that is, which category the data x belongs to. For medical diagnosis, it is the diagnosis result. For grayscale image classification, it is which category the picture belongs to.

[0058] Advantages of the present invention:

[0059] Build a multi-objective feature selection model to find the trade-off between classification performance and feature subsets and obtain multiple solutions for decision-makers to choose from.

[0060] A hybrid coding scheme is proposed and different genetic operators are designed, enabling feature selection and classifier optimization to be carried out simultaneously.

[0061] A single-hidden layer neural network that can automatically design the structure and input parameters is used as the classifier, and a non-iterative algorithm is used to solve the output weights.

[0062] The method of selective neural network ensemble is adopted to make the final decision, improving the classification accuracy. Description of the Drawings

[0063] Figure 1 It is a diagram of hybrid coding.

[0064] Figure 2 It is the process framework diagram of the present invention.

[0065] Figure 3 It is the Pareto front of the present invention on different datasets under different iteration times.

[0066] Figure 4 shows the accuracy comparison results of the present invention: (a) comparison with LRSSR; (b) comparison with MSFS; (c) comparison with CUS-SPSO; (d) comparison with bAAAs1; (e) comparison with SLMEA; (f) comparison with SaWDE.

[0067] Figure 5 shows the comparison results of the selected features of the present invention: (a) comparison with CUS-SPSO on low- and medium-dimensional datasets; (b) comparison with CUS-SPSO on high-dimensional datasets; (c) comparison with bAAAs1; (d) comparison with SLMEA; (e) comparison with SaWDE. Detailed Implementation Manner

[0068] The present invention will be further described in detail below with reference to the drawings.

[0069] As Figure 1 As shown at the bottom, the network constructed by the present invention is an N-L-K type network, which includes an input layer, a hidden layer, and an output layer. The input layer contains N neurons corresponding to N attributes of the sample data, the output layer contains K neurons corresponding to K classification labels, and the hidden layer contains L neurons, where L = N + 1. First, the training dataset with m samples after feature selection can be expressed as:

[0070] Where Denotes the Hadamard product of two vectors, which is the original x i Initial training data, Is the new training data after feature selection, and θ1 is the first part of the encoding.

[0071] Secondly, for the activation state θ2 of the hidden neurons in the SLFN and the input parameters θ3 = (ω, b) of the SLFN, calculate the actual input weight vector between the input layer and the j-th hidden neuron And the actual offset vector of the hidden layer

[0072]

[0073] Then the actual output of the SLFN with m samples can be defined as:

[0074]

[0075] Where β j Is the output weight vector between the j-th hidden neuron and the output layer, Is the bias of the j-th hidden neuron. g(·) is the activation function of the hidden layer, and the Sigmoid function is used.

[0076]

[0077] The analytical calculation of the output weight matrix β between the hidden layer and the output layer is as follows:

[0078] Let H = (h ij ) m×L Be the output matrix of the hidden layer nodes. According to the theory, the solution of ELM (i.e., equation (4)) can be expressed as follows:

[0079] Hβ = T (5)

[0080] Where T = [t1, t2,..., t m T Is the target output matrix. So the output weight β can be calculated by the following formula:

[0081]

[0082] Where Represents the Moore-Penrose (MP) generalized inverse matrix of H.

[0083] ​The present invention can perform feature selection and classifier optimization simultaneously, and adopt a hybrid coding scheme and corresponding hybrid operations. From the perspective of multi-objective optimization, multiple Pareto optimal solutions are obtained, that is, multiple selected feature subsets and optimized designed classifiers can be obtained in one run.

[0084] Figure 2 It is the algorithm flow chart of the present invention, and the specific steps are as follows:

[0085] 1: Initialize the parent population P pop with size N t , and set t = 0.

[0086] 2: For each individual in P t , calculate the output weight β using Equation (6), then calculate the objective function, and perform fast non-dominated sorting.

[0087] 3: While t ≤ t max do

[0088] 4: Based on p c , calculate the number of individuals performing the crossover operation "nc", and select nc parents from the population.

[0089] 5: for

[0090] 6: Randomly select two individuals from the nc parents as parents, and then divide them into a binary vector, a binary vector, and a real vector according to the coding method.

[0091] 7: Use Equation (11) to perform the crossover operation on each part of the offspring, and then splice the offspring into a complete individual.

[0092] 8: end for

[0093] 9: Calculate the mutation number "nc" of the offspring according to nc.

[0094] 10: for i = 1 → nm do

[0095] 11: Select a solution from the offspring, and then divide it into a binary vector, a binary vector, and a real vector according to the coding method.

[0096] 12: For the binary vector, use Equation (12) for mutation, and for the real vector, use Equation (13) for mutation, and then the offspring are concatenated into a new individual.

[0097] 13: end for

[0098] 14: Combine the parent population P t and the offspring population Q tCombined into 2N pop to form a new population R t ;

[0099] 15: Perform fast non - dominated sorting and crowding degree calculation on population R t to generate a new parent population P t+1 ;

[0100] 16: t = t + 1.

[0101] 17: end while

[0102] 18: Return population P tmax .

[0103] 19: For each individual in population P tmax , extract the feature subset and input weights of the classifier, and then calculate the output weight β through equation (6).

[0104] 20: All individuals are sorted in ascending order according to the training error, and the top three individuals are selected for ensemble learning.

[0105] The implementation of step 1 mainly includes an encoding scheme:

[0106] The encoding scheme of each individual is as Figure 1 shown, mainly including the following three parts:

[0107] The first part mainly performs binary encoding on feature selection. Figure 1 The green part in can be expressed as θ1=(1,...,1,...,0). When the value is 1, the corresponding feature is selected and input into the network. On the contrary, it is not input into the network, and the corresponding node is represented by white. The encoding length of this part is N.

[0108] The second part mainly performs binary encoding on the network structure of the classifier. Figure 1 The corresponding blue part in represents θ2=(1,...,1,...,0,...,1), where the value of 1 indicates that the corresponding hidden neuron is activated, and the value of 0 indicates the opposite. Similarly, the nodes that are not activated are represented by white. The encoding length of this part is L.

[0109] In the third part, real - number encoding is performed on the input parameters of the classifier. The input parameters mainly include input weights and hidden - layer biases, which are randomly generated within the interval [-1,1]. Corresponding to Figure 1 the yellow part in, it is expressed as θ3=(ω,b)=(0.8,...,-0.3). The encoding length of this part is (N + 1)×L.

[0110] Therefore, each individual can be represented as θ = (θ1, θ2, θ3), with a length of (N + 2) × L + N.

[0111] The implementation of Step 2 mainly involves a multi-objective model:

[0112] The feature selection process is a discrete multi-objective optimization problem with multiple objectives: 1) Maximize the classification accuracy, 2) Make the classifier network structure as compact as possible, and 3) Select as few features as possible. Therefore, we evaluate the quality of the selected feature subsets based on the above objectives. Obviously, the first objective function is the error function commonly used in neural network training, i.e., the root mean square error (RMSE).

[0113]

[0114] Regarding the complexity of the network, the structure of the hidden layer is mainly considered and defined using the following formula

[0115]

[0116] where θ 2j is the second part of the coding scheme, and the number of selected features is calculated using a proportional function.

[0117]

[0118] where N s is the number of selected features, and N is the total number of features. According to the previous discussion, these three objectives are mutually contradictory. Therefore, we can regard feature selection as a problem consisting of three objectives to be minimized simultaneously:

[0119] min{E, H, F} (10) The hybrid crossover and hybrid mutation operations of Step 7 and Step 12 are introduced in detail below:

[0120] For the adopted hybrid coding scheme, traditional genetic operators can no longer solve this problem (Equation (10)). Therefore, NSGA-II with hybrid operators is designed, including hybrid crossover and hybrid mutation operations, to ensure the feasibility of the solutions. The following hybrid crossover operation is performed on the individuals:

[0121] (1) At the beginning of the crossover, a group of individuals is randomly selected from the population and marked as the parent individuals.

[0122] (2) According to Figure 1 the coding scheme in, the parents are divided into three parts.

[0123] (3) The operation of Equation (11) is performed on each part to obtain the corresponding crossover individuals.

[0124]

[0125] where X 1j and X 2j are the selected parents, and Y 1j and Y 2j are the offspring. α j is a random number, randomly selected from {0, 1} when performing binary crossover, and randomly selected from [0, 1] when performing real - valued crossover.

[0126] (4) After completing the crossover operations respectively, recombine the three parts to form a new individual.

[0127] Perform the following hybrid mutation operation on the individual:

[0128] (1) Select a parent individual and divide the individual into three parts according to the coding method.

[0129] (2) Perform the following mutation operations respectively. When facing binary coding, use equation (12).

[0130]

[0131] where X j is the selected parent individual, Z j is the offspring. rand(0, 1) is a random number and μ is a control parameter. When facing real - valued coding, use equation (13).

[0132]

[0133] (3) After completing the mutation operations respectively, recombine the three parts to form a new individual.

[0134] The implementation of step 20 mainly includes a selective ensemble learning:

[0135] To improve the classification performance and make a final decision, selective classifier ensemble learning is introduced. Neural network ensemble is to use a finite number of neural networks to learn the same problem and integrate the outputs under a certain input example, and the output of this input is jointly determined by the outputs of each neural network. This method can significantly improve the generalization ability of the neural network system. The ensemble adopts weighted average, and assigns a weight λ i to each network, satisfying the following conditions,

[0136]

[0137] By performing NSGA-II optimization through hybrid coding, multiple optimal solutions can be obtained, namely multiple optimized feature subsets and the corresponding designed classifiers. To obtain the highest classification accuracy, the smallest feature subset is selected to make the classifier structure as simple as possible. All the obtained Pareto optimal solutions are first sorted in ascending order of the training error. Then, the first three solutions are selected, the selected feature subsets are extracted from the test set, and the designed classifiers are optimized separately. Then, these three classifiers are integrated into the base classifier, and their weights are 0.7, 0.2, and 0.1 in turn. Finally, the final classification accuracy is obtained according to Equation (15).

[0138]

[0139] where x is the test data, and C i (x) = [C i1 (x), C i2 (x), …, C iK (x)] is the classification result of the i-th classifier. Class(x) = [C1(x), C2(x), …, C K (x)] is the obtained output result. The class of the final x is given by Equation (16),

[0140] Class(x) = argmax 1≤j≤K Class j (x) (16)

[0141] Experimental Results and Analysis:

[0142] Many experiments are conducted to evaluate the performance of the proposed method and some state-of-the-art feature selection methods on real datasets. According to previous research work on feature selection, a total of 35 well-known real datasets are selected from the UCI Machine Learning Repository and the Scientific Feature Feature Selection Repository for experiments. These datasets come from multiple fields, including diseases, biology, gene expression, chemistry, handwriting recognition, speech recognition, text recognition, etc. Their detailed information is shown in Table 1, including the number of features, the number of samples, and the number of classes of the datasets. The number of features ranges from 16 to 7129, and the sample size ranges from 32 to 9298. All datasets are sorted in ascending order of the number of features. In the experiment, all data is divided into two groups. The datasets numbered 1 to 22 are mainly low- and medium-dimensional datasets, and the datasets numbered 23 to 35 (i.e., the datasets with thousands of features) are high-dimensional classification problems, and the scalability of the algorithm is studied.

[0143] Table 1 Details of the Datasets Used

[0144]

[0145]

[0146] Normalize each feature in each dataset to [-1, 1] using the following formula, where x i is normalized to x' i .

[0147]

[0148] In the feature selection experiment, the stratified random sampling technique was adopted. In other words, for a given dataset, the samples were first divided into five parts according to the categories, four of which were used as the training set and the rest as the test set. In the training stage, feature selection was performed on the training set to select the optimal feature subset. In the test stage, a subset was extracted from the test data, and then the extracted subset was input into the neural network classifier to calculate the classification accuracy. Finally, after 20 overall runs, the average of the 20 results was taken as the final result of feature selection.

[0149] A single-hidden-layer neural network was used for learning, with the sigmoid function and the linear function as the activation functions of the hidden layer and the output layer respectively. Since the number of neurons in the input layer and the output layer of the network corresponds to the number of features of the data samples and the number of classes respectively, we only need to determine the number of neurons in the hidden layer, and there is no unified standard for this. In the work, considering the characteristics of high-dimensional data, the maximum number of hidden neurons is: L = N + 1, where N is the number of features of the original dataset. After optimization, the final network structure was determined. For NSGA-II, there are mainly four parameters to be determined, as shown in Table 2.

[0150] Table 2: Parameter settings of algorithm JMO-FSCD

[0151]

[0152]

[0153] As mentioned above, the present invention introduces ensemble learning to make a final decision among multiple Pareto solutions. After sorting all the solutions, the top three solutions with the smallest training error are selected, that is, three ELM classifiers are integrated. Table 3 lists the network structures that obtain the best classification results on 9 datasets, where the input layer corresponds to the input subset selected by feature selection, the number of hidden layer nodes is automatically optimized on different datasets by JMO-FSCD, and the output nodes correspond to the number of classes of the dataset. Therefore, the structure of the entire network can be determined. It can be found that on different datasets, the obtained network structures are completely different and are closely related to the characteristics of the datasets.

[0154] Table 3 Structures of three extreme learning machines (ELMs) for ensemble learning

[0155]

[0156] Figure 3 It shows the Pareto front obtained by JMO - FSCD on these 9 datasets. It can be found that as the number of iterations increases, the solution gets closer and closer to the origin of the coordinates, and the Pareto front is a surface, verifying the correctness of the model, that is, these three objectives conflict with each other.

[0157] Comparison with LRSSR, the specific comparison results are shown in Table 4. In terms of accuracy, no matter which classifier LRSSR uses, the algorithm JMO - FSCD of this application has achieved the best performance except for USPS and has had a significant improvement. After the best classification accuracy of JMO - FSCD, the average classification accuracy is also listed. In comparison, the average classification accuracy is also competitive. In the sorting comparison, the smallest sorting is also obtained, which shows the good performance of the algorithm of this application, and Figure 4(a) also provides strong evidence. For the selected features, JMO - FSCD selects fewer features on USPS, obtains the same features on Isolet, and selects more features on Yale. Generally speaking, the features selected by these two methods are of the same order of magnitude. In terms of running time, JMO - FSCD uses more time on all datasets, but the accuracy has been significantly improved. Therefore, generally speaking, the algorithm of this application is feasible and competitive, and the running time is acceptable on the premise of ensuring accuracy.

[0158] Table 4 Best classification accuracy (Best), number of selected features (Feature number), and running time (Time) of JMO - FSCD and LRSSR

[0159]

[0160] Comparison with MSFS, our algorithm is compared in detail with MSFS on 4 high - dimensional datasets, including average classification accuracy, algorithm running time, and average number of selected features.

[0161] Table 5 Average classification accuracy (Accuracy), number of selected features (Feature number), and running time (Time) of JMO - FSCD and MSFS

[0162]

[0163]

[0164] Comparison with CUS-SPSO. JMO-FSCD was compared with CUS-SPSO on 15 low- and medium-dimensional datasets and 7 high-dimensional datasets. Table 6 shows the comparison results in terms of average classification results, the number of selected features, and the running time of the algorithm. In terms of average classification accuracy, JMO-FSCD achieved better accuracy on more than half of the datasets, that is, 12 datasets, and better performance on high-dimensional datasets. At the same time, the performance of the Zoo, Segment(210), Movement, HillValley, Isolet, Prostate_GE, and Leukemia datasets was significantly improved. In summary, its ranking was 1.4545. Figure 4(c) also shows the classification ability of JMO-FSCD. In terms of the number of selected features, although JMO-FSCD only reached the minimum number of features on 8 datasets, generally speaking, the number of selected features was the same. As can be seen from the observations in Figure 5, the difference between the two was not significant.

[0165] Table 6 Average classification accuracy (accuracy), number of selected features (number of features), and running time (time) of JMO-FSCD and CUS-SPSO

[0166]

[0167]

[0168] In terms of running time, as can be seen from Table 2, the maximum evaluation time of CUS-SPSO was 10,000, and the maximum evaluation time of JMO-FSCD was 1,500, with a difference of one order of magnitude, so a specific comparison could not be made. However, it can be seen that JMO-FSCD had a shorter usage time on low- and medium-dimensional datasets. However, high-dimensional datasets required more time because the individual coding brought by high-dimensional data was too long, resulting in time-consuming calculations. Generally speaking, JMO-FSCD was competitive on low- and medium-dimensional datasets, but it significantly improved the classification accuracy.

[0169] Comparison with bAAAs1. JMO-FSCD was compared with bAAA on 10 datasets in Table 1.

[0170] To obtain the best classification accuracy, the algorithm achieved optimal results on 7 datasets, and the best accuracy on most datasets was significantly improved. In terms of average accuracy, fsmhcelm achieved better results on 6 datasets, with an improvement of more than 5 percentage points on HillValley and Sonar, which can also be seen from Figure 4(d). Although the standard deviation obtained by the algorithm of this application is larger, the overall accuracy has been improved. From the rankings, it can be seen that the performance of JMO-FSCD is better than that of bAAA in both the best accuracy and average accuracy. For the average number of selected features, JMO-FSCD selected a smaller feature subset on all datasets, which can also be verified by observing Figure 5. Therefore, it can be said that JMO-FSCD is an effective feature selection method, and a small subset of features can be obtained in the case of significant differences in classification performance. Table 7 Best classification accuracy (Best), average classification accuracy (Accuracy), and number of selected features (Number of Features) of JMO-FSCD and bAAAs1

[0171]

[0172]

[0173] Comparison with SLMEA. To explore the relationship between data, features, and labels during data structure mining, an algorithm based on sparse low-dimensional representation and maximum entropy adaptive graph was proposed to improve the performance of feature selection. Through extensive evaluation of multiple mainstream datasets, the experimental results show that the SLMEA algorithm has better feature selection effect than other comparison algorithms. Table 8 shows the details of the average accuracy and the number of selected features of the two algorithms. In terms of average accuracy, the proposed JMO-FSCD performs well on all datasets, with a significant improvement in accuracy, as also demonstrated by Figure 4(e). For the number of selected features, SLMEA gives the number of selected features when obtaining the best accuracy, while JMO-FSCD provides the average number of features during 20 independent runs. Although JMO-FSCD selects more features, it has an obvious advantage in accuracy. Considering that classification accuracy takes precedence over the number of selected features, JMO-FSCD is more competitive.

[0174] Table 8 Average classification accuracy (Accuracy) and number of selected features (Number of Features) of JMO-FSCD and SLMEA

[0175]

[0176]

[0177] Comparison with SaWDE, many evolutionary algorithms have been used to solve the FS problem, but most of the research mainly focuses on low-dimensional problems. When dealing with large-scale problems, they are prone to premature convergence and instability. Therefore, SaWDE, a new weighted differential evolution algorithm based on an adaptive mechanism, is proposed to solve large-scale FS problems. The detailed comparison between SaWDE and JMO-FSCD is shown in Table 9 and also in Figures 4(f) and 5(e). It can be found that, except for the effective datasets, JMO-FSCD achieves the best classification accuracy, which can also be observed from Figure 4(f). In terms of the number of selected features, JMO-FSCD exceeds SaWDE, but JMO-FSCD still deletes many features. Generally speaking, considering that the classification accuracy is better than the number of selected features, JMO-FSCD performs better.

[0178] Table 9 Average classification accuracy (accuracy rate) and the number of selected features (number of features) of JMO-FSCD and SaWDE

[0179]

[0180] The present invention proposes a new method for multi-objective feature selection based on a single-hidden-layer feedforward neural network. By designing a new hybrid coding scheme, feature selection and classifier optimization are carried out simultaneously. The Pareto optimal solutions are obtained using NSGA-II with hybrid genetic operations. Each solution contains the selected feature subset and the optimized classifier structure and input parameters. For the output parameters of the classifier, a non-iterative learning algorithm is introduced to ensure good learning performance and improve the learning speed. The final result is calculated using the method of ensemble learning, further improving the performance of the algorithm. The new method, namely JMO-FSCD, is compared with six state-of-the-art algorithms on 35 benchmark classification datasets with the number of features ranging from 16 to 7129. The results show that JMO-FSCD is superior to or equal to all competing algorithms in terms of classification accuracy, successfully selects nearly half of the features, and achieves a shorter running time on medium and small datasets. In addition, the JMO-FSCD algorithm also obtains better diversity in the solution.

Claims

1. A multi-objective optimization grayscale image classification method for joint design of feature selection and classifier, characterized in that Including the following steps; Step 1: Use the stratified random technique to divide the high-dimensional dataset for grayscale image classification into a training set and a test set, and standardize the features of the training set and the test set; Step 2: Adopt a hybrid coding scheme to encode individuals and generate an initial population; Step 3: Use NSGA-II and a hybrid operator to establish and solve a multi-objective feature selection model; Step 4: Obtain multiple Pareto optimal solutions through a multi-objective evolutionary algorithm with a hybrid operator. Each optimal solution contains a feature subset and a designed classifier, and sort the solutions according to the training error; Step 5: Extract the selected features obtained through the solution of the entire model and input them into the classifier, and obtain the final grayscale image classification result through selective neural network integration; In the hybrid coding scheme in Step 2, the feature selection and classifier optimization tasks are carried out simultaneously; The coding scheme of each individual includes the following three parts: (1) The first part mainly performs binary coding for feature selection; expressed as θ1=(1,...,1,...,0), select the features with a value of 1 and input them into the network. On the contrary, it is not input into the network, and the corresponding nodes are represented in white. The coding length of this part is N; (2) The second part mainly performs binary coding for the network structure of the classifier, expressed as θ2=(1,...,1,...,0,...,1), where a value of 1 indicates that the corresponding hidden neuron is activated, and a value of 0 indicates the opposite. Similarly, the coding length of this part for the nodes that are not activated is L; (3) In the third part, real number coding is performed on the input parameters of the classifier. The input parameters include input weights and hidden layer biases, which are randomly generated within the interval [-1,1], expressed as θ3=(ω, b). The coding length of this part is (N + 1)×L; Therefore, each individual is represented as θ=(θ1, θ2, θ3), with a length of (N + 2)×L + N.

2. The multi-objective optimization grayscale image classification method for joint design of feature selection and classifier according to claim 1, characterized in that The individuals obtained in Step 2 are the initial individuals in the randomly generated evolutionary process, and these initial individuals will be evolutionarily cycled and optimized in the multi-objective model in Step 3 to obtain the final multiple optimal solutions.

3. A multi-objective optimization grayscale image classification method for joint design of feature selection and classifier according to claim 1, characterized in that, In Step 3, a multi-objective feature selection model is established to find the trade-off between classification performance and feature subsets, and obtain multiple solutions for decision-makers to use; The first objective function of the model is the error function commonly used in neural network training, namely the root mean square error; Regarding the complexity of the network, the structure of the hidden layer is mainly considered and defined using the following formula: where θ 2j is the second part of the coding scheme, and the number of selected features is calculated using a proportional function; where N s is the number of selected features, N is the total number of features, and according to the previous discussion, these three objectives are mutually contradictory. Therefore, feature selection is regarded as a problem consisting of three objectives to be minimized simultaneously: min{E, H, F}; where E is the training error function, H is the network complexity, and F is the proportion of selected features.

4. A multi-objective optimization grayscale image classification method with combined design of feature selection and classifier according to claim 1, characterized in that, In Step 3, NSGA-II with a hybrid operator is used to solve the established multi-objective feature selection model, including hybrid crossover and hybrid mutation operations; Crossover is to select two individuals from the population and exchange parts of the two individuals with a certain probability, and perform the following hybrid crossover operation on the individuals: (1) At the beginning of crossover, randomly select a group of individuals from the population and mark them as parent individuals; (2) Divide the parents into three parts; (3) Perform the operations of the following equations on each part to obtain the corresponding crossover individuals; Where X 1j and X 2j are the selected parents, Y 1j and Y 2j are the offspring, α j is a random number, randomly selected from {0, 1} when performing binary crossover and randomly selected from [0, 1] when performing real-valued crossover; (4) After completing the crossover operation respectively, recombine the three parts to form a new individual.

5. A multi-objective optimization gray image classification method with combined design of feature selection and classifier according to claim 4, characterized in that, Mutation means replacing some gene values in the individual coding string with other gene values according to the mutation probability to form a new individual. The following hybrid mutation operations are performed on the individual: (1) Randomly select a paternal individual and divide the individual into three parts according to the coding method; (2) Perform the following mutation operations respectively. When facing binary coding, use the following equations; Where X j is the parental individual, Z j is the offspring, rand(0,1) is a random number, and μ is a control parameter. When dealing with real-coded representation, the following equation is used; (3) After completing the mutation operations separately, recombine the three parts to form a new individual, and this new individual is used to generate a new population R t .

6. The multi-objective optimization grayscale image classification method for joint design of feature selection and classifier according to claim 1, characterized in that The new individuals finally formed in step 3 are used to generate a new population. Step 3 is looped. After reaching the maximum number of iterations, the final population obtained can be used in step 4. The final population includes multiple solutions. Each solution includes the selected features and the optimized network input parameters. Step 4 is to apply a non-iterative algorithm to the network input parameters of each solution to solve the corresponding output parameters, thereby obtaining the designed network classifier.

7. A multi-objective optimization grayscale image classification method with combined design of feature selection and classifier according to claim 1, characterized in that In step 4, the classifier adopts a single-hidden-layer neural network with an automatically designed structure and input parameters, and uses a non-iterative algorithm to solve the output weights; A single-hidden-layer feedforward neural network is considered for feature selection; Extreme learning machine ELM is a single-hidden-layer feedforward neural network algorithm. ELM needs to preset the number of hidden layer nodes and does not need to adjust the input weights and hidden layer biases during operation. The output of ELM with m samples is defined as i = 1, 2, ..., m.; where ω j and β j are the input and output weight vectors of the j-th hidden neuron respectively, b j is the bias of the j-th hidden neuron, and g(·) is the activation function of the hidden layer, using the sigmoid function; Let \(H=(h ij ) m×L = g(\omega j , b j , x i ) be the output matrix of the hidden layer. Then the solution of ELM is expressed as follows: \(H\beta = T\); where \(T = [t_1, t_2, \ldots, t m ) T is the target output matrix. Therefore, the output weight \(\beta\) is calculated by the following formula: where represents the Moore - Penrose (MP) generalized inverse matrix of \(H\), which is used for the weights and structure of the neural network for classification; Through step 4, multiple designed network classifiers and the corresponding selected feature subsets are obtained. For these multiple solutions, to obtain the final result, step 5 is used to select some of the solutions for ensemble learning to make a final decision to improve the classification accuracy.

8. A multi-objective optimization grayscale image classification method with combined design of feature selection and classifier according to claim 1, characterized in that, In step 5, a selective neural network ensemble method is used to make the final decision to improve the classification accuracy. The weights are weighted averages, and a weight value λ is assigned to each network. i , satisfying the following conditions; A multi-objective feature selection model is established through step 3, that is, it is necessary to minimize the training error, network complexity, and the proportion of selected features. NSGA-II with a hybrid-coded genetic operator is used for optimization and solution. Finally, multiple optimal solutions are obtained. Each solution includes the selected feature subset and the optimized network input parameters respectively. For the multiple solutions obtained, in order to make the final decision and also to improve the classification accuracy; First, arrange all the obtained Pareto optimal solutions in ascending order of the training error. Then, select the first three solutions, extract their selected feature subsets from the test set, and optimize the designed classifiers respectively. Then, integrate these three classifiers into base classifiers, and their weights are 0.7, 0.2, and 0.1 in turn. Finally, obtain the final classification accuracy according to the following equation; j = 1, 2, ..., K. Where x is the test data, C i (x) = [C i1 (x), C i2 (x), …, C iK (x)] is the classification result of the i-th classifier, Class(x) = [C1(x), C2(x), …, C K (x)] is the obtained output result, and the final class of x is given by the following formula, Class(x) = argmax 1≤j≤K Class j (x).

Citation Information

Patent Citations

  • High-dimensional data classification method based on two-stage mixed feature selection

    CN113780334A

  • Communication signal identification method and system based on adaptive feedforward neural network

    CN115049006A