Small sample image classification method based on hierarchical learning genetic programming algorithm
Through a hierarchical learning framework and an integrated strategy based on individual differences values, the problem of insufficient effectiveness of existing genetic programming algorithms in low-quality and small-sample image classification is solved, and higher classification accuracy and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510014560.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-06
AI Technical Summary
The existing image classification method based on genetic programming algorithms is not effective when processing low-quality and small sample images, and the large solution space makes the algorithm easily fall into local optimality and has weak generalization ability.
Using a hierarchical learning framework, the first layer of parallel exploration genetic programming (PEGP) module is responsible for image preprocessing and feature extraction, and the second layer of development of integrated genetic programming (DEGP) module to build an integrated solution based on feature storage tables, and optimize the classification effect through an integration strategy based on individual differences.
It significantly reduces the search space, improves the accuracy and generalization ability of image classification, and maintains high performance especially when there are fewer training samples.
Smart Images

Figure CN119942198A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of image classification, and in particular to a small sample image classification method based on a hierarchical learning genetic programming algorithm. Background Art
[0002] Image classification tasks involve automatically classifying images into predefined categories. This task can usually be accomplished efficiently and accurately when the image quality is high and the samples are sufficient. However, in practical applications, the common problem faced is the low quality and insufficient number of samples, which makes the accurate classification of low-quality and few-sample images a challenging task. Genetic programming GP algorithms can model the solutions to various application problems as individuals in a population and automatically evolve suitable solutions. Nowadays, genetic programming algorithms have been proven to be effective in solving image classification problems. However, current genetic programming-based methods have limitations in program structure, especially when dealing with low-quality and few-sample images. The solution space of genetic programming algorithms is large, which makes the algorithm prone to local optimality, but this problem is often not fully explored and studied. In addition, the optimal individuals generated by genetic programming algorithms may over-fit the training data, resulting in poor performance on new or unknown data.
[0003] Disadvantages of the existing technology: The existing image classification method based on genetic programming algorithm integrates all image processing steps into its program structure, which leads to a complex program structure and a huge search space. Since the search ability of genetic programming is weak, this seriously affects the performance of the algorithm. At the same time, the existing research is mainly designed for the case of sufficient samples, so when the training data is small, the performance of the existing algorithm is usually insufficient. Summary of the invention
[0004] The present invention provides a small sample image classification method based on a hierarchical learning genetic programming algorithm, which can achieve good classification effect when the number of samples is limited.
[0005] To achieve the above object, the present invention provides a small sample image classification method based on a hierarchical learning genetic programming algorithm, the key of which is to include the following steps:
[0006] Step 1: construct a small sample image classification system based on a hierarchical learning genetic programming algorithm, wherein the small sample image classification system is provided with an image acquisition module, wherein the image acquisition module is connected to a hierarchical evolutionary learning framework, wherein the hierarchical evolutionary learning framework is provided with a first-layer parallel exploration genetic programming PEGP module and a second-layer development integrated genetic programming DEGP module;
[0007] Step 2: The image acquisition module acquires an image data set and divides the image data set into a training set and a test set;
[0008] Step 3: The parallel exploration genetic programming PEGP module obtains training set data, and performs image preprocessing and feature extraction operations on the training set data to construct a feature storage table, and then passes the feature storage table to the development integrated genetic programming DEGP module;
[0009] Step 4: The developed integrated genetic programming DEGP module uses the features in the feature storage table as terminal input, further constructs an integrated solution, and finally optimizes the final classification effect through an integrated strategy based on individual difference values, thereby outputting a high-performance image classification solution;
[0010] Step 5: Use the test set as input to the image classification solution, then output the predicted class labels of the test set, and finally evaluate the performance of the image classification solution based on the actual labels of the test set;
[0011] Step 6: The image acquisition module acquires the image data to be classified, uses the image data to be classified as the input of the image classification solution, and outputs the image classification result.
[0012] Through the above design, the present invention adopts a hierarchical evolutionary learning framework, aiming to narrow the search space and efficiently find classification solutions.
[0013] In the first layer, a parallel exploration genetic programming PEGP is defined, which focuses on image preprocessing and feature extraction to build a feature storage table. By exploring diverse and effective feature building blocks in parallel, this layer can improve the combined generation efficiency of preprocessing and feature extraction in the GP algorithm, and finally output diverse feature blocks as input for the second stage.
[0014] The second layer is to develop an integrated genetic programming DEGP, which combines feature blocks in the search space and generates an integrated solution based on the feature storage table under the guidance of the newly defined program structure. The hierarchical learning-based framework significantly reduces the search space of the GP algorithm, thereby ensuring that the final solution can capture key feature information and maintain high performance even with fewer training samples.
[0015] In addition, given that the GP method has weak generalization ability when dealing with few-sample image classification tasks, an integration strategy based on individual difference values is proposed to improve the classification performance.
[0016] Preferably: in step 1, the parallel exploration genetic programming PEGP module is provided with three different types of feature exploration blocks in parallel, the three feature exploration blocks are respectively a first feature exploration block, a second feature exploration block and a third feature exploration block, the program structure, function set and terminal set of the first feature exploration block, the second feature exploration block and the third feature exploration block are respectively set, and the first feature exploration block, the second feature exploration block and the third feature exploration block generate feature building blocks of three different evolutionary directions in parallel.
[0017] In order to ensure that the generated feature building blocks are effective and diverse, the first-level PEGP algorithm must be configured with three feature building blocks with different evolutionary directions, each with a new program structure, function set and terminal set.
[0018] The parallel exploration genetic programming PEGP module uses three different program structures to generate populations, and these different structures generate individuals representing different types of feature construction methods. At the end of the evolutionary learning process, the algorithm outputs multiple high-quality individuals, which are defined as different feature building blocks according to the types of features they represent.
[0019] Preferably, the first feature exploration block adopts a linear genetic programming (LGP) algorithm focusing on local features, and the program structure of the first feature exploration block includes a region extraction layer, a first image filtering layer, a first feature extraction layer, and a first feature series connection layer;
[0020] The second feature exploration block adopts a grammar-guided genetic programming (GGP) algorithm focusing on global features, and the program structure of the second feature exploration block includes a second image filtering layer and a second feature extraction layer;
[0021] The third feature exploration block adopts a Cartesian genetic programming (CGP) algorithm that focuses on serial features. The program structure of the third feature exploration block includes a third image filtering layer, a maximum pooling layer, a third feature extraction layer, and a second feature serial layer.
[0022] The parallel exploration genetic programming PEGP module adopts the LGP algorithm focusing on local features, the GGP algorithm focusing on global features, and the CGP algorithm focusing on serial features to evolve different types of feature building blocks.
[0023] Considering the effectiveness of obtaining local features, the design of the LGP algorithm includes a region extraction layer, a filtering layer, a feature extraction layer and a feature concatenation layer to extract effective local features. The function of the region extraction layer selects a valid region, and the function of the image filtering layer performs filtering on the extracted region. The function of the feature extraction layer selects an existing feature extraction method to extract features from the image. The function of the feature concatenation layer can concatenate the output of the feature extraction layer. The program structure designed for the LGP algorithm will make the individuals generated by its evolution focus on different features of a single region or the same features of multiple regions.
[0024] The GGP algorithm includes an image filtering layer and a feature extraction layer. The image filtering layer places multiple types of filter functions, and its input and output are consistent. The generated individuals may undergo multiple filtering processes. The functions of the feature extraction layer can extract features from the filtered image or the original image. The program structure of the GGP algorithm guides individuals to tend to combine multiple filtering processes and global feature searches.
[0025] For the CGP algorithm, it adds a maximum pooling layer and a feature construction layer on the basis of the GGP algorithm. The design of the pooling layer helps reduce the time spent in subsequent steps. The function of the feature concatenation layer is designed as a flexible layer, which will concatenate different types of global features. The program structure designed for the CGP algorithm will make the individual learning biased towards concatenating different global features.
[0026] Preferably, in step 3, the parallel exploration genetic programming PEGP module constructs a feature storage table, comprising the following steps:
[0027] Step A1: Population initialization: The feature exploration block obtains training set data and initializes the population according to a predetermined program structure, function set and terminal set; each individual in the population can be regarded as a feature extraction algorithm.
[0028] Step A2: Evaluate individual fitness: Each individual in the population extracts features from the training set, and then inputs the extracted features into a support vector machine (SVM). The support vector machine (SVM) outputs a predicted class label, and then the individual is evaluated using classification accuracy to obtain a fitness value of the corresponding individual.
[0029] The calculation expression of classification accuracy is as follows:
[0030]
[0031] Among them, N correct is the number of correctly predicted instances, N total is the total number of instances, and Fitness represents the fitness value of the individual;
[0032] In this process, in order to reduce the possibility of overfitting, the k-fold cross-validation method was used. In addition, in order to ensure the adequacy of cross-validation, the k value was set to the smaller of the number of training samples n for each category and 10.
[0033] Step A3: Elite operation: Use the elite strategy to select the best individuals in the population and directly copy them to the next generation population;
[0034] Step A4: Selection operation: A certain number of individuals are selected from the population using the tournament selection method, and each individual has the same probability of being selected; according to the fitness value of each individual, the individual with the best fitness value is selected to perform crossover and mutation operations to generate new individuals;
[0035] Step A5: Repeat steps A2-A4 until the maximum number of iterations is reached, then proceed to step A6;
[0036] Step A6: Select the top 50% individuals in the last generation of the population as feature building blocks;
[0037] Step A7: Use the feature building block to extract corresponding features from the training set, number them, and store them in a feature storage table as input for the second layer learning.
[0038] The parallel exploratory genetic programming PEGP module takes the training data set as input and generates initial populations in parallel according to three different program structures, function sets and terminal sets. Subsequently, these populations evolve in parallel, and each evolved population outputs building blocks with specific characteristics, which are then stored and used as input for the second layer of learning, ensuring that the development integrated genetic programming DEGP module has sufficient diverse features for selection and optimization.
[0039] As a preference: the development integrated genetic programming DEGP module is set with a new program structure, function set and terminal set, and the program structure of the development integrated genetic programming DEGP module includes a feature construction layer, a classification layer and a combination layer;
[0040] The feature construction layer is used to take at least two features in the feature storage table as input and return a concatenated feature, or construct a new feature according to parameters;
[0041] The classification layer is used to take the output features of the feature construction layer as input and output a predicted class label;
[0042] The combination layer is used to take at least two groups of predicted class labels output by the classification layer as input, and perform voting or weighting to output a new predicted class label.
[0043] The program structure, function set and terminal set of the genetic programming DEGP module of the development integration have been completely redesigned to optimize the integration process; under the guidance of the newly defined program structure, terminal set and function set, the final solution to the image classification task is developed through an evolutionary learning process.
[0044] Preferably, in step 4, the integrated genetic programming DEGP module is further developed to construct an integrated solution, comprising the following steps:
[0045] (1) Population evolutionary learning
[0046] Step B1: Population initialization: The developed integrated genetic programming DEGP module obtains the features in the feature storage table and initializes the population according to the new program structure, function set and terminal set; each individual in the population is mapped to a classification scheme and output as a predicted class label.
[0047] Step B2: Evaluate individual fitness: Each individual in the population extracts features from the feature storage table, outputs a predicted class label, and then evaluates the fitness value of the individual;
[0048] Step B3: Elite operation: Use the elite strategy to select the best individuals in the population and directly copy them to the next generation population;
[0049] Step B4: Selection operation: A certain number of individuals are selected from the population using the tournament selection method. According to the fitness value of each individual, the individual with the best fitness value is selected to perform crossover and mutation operations to generate new individuals.
[0050] Step B5: Repeat steps B2-B4 until the preset upper limit of iteration is reached, then proceed to step B6;
[0051] (2) Integration strategy based on individual difference values
[0052] Step B6: Select the best performing individual in the last generation of the population as the benchmark, calculate the difference value of other individuals in the population, and evaluate the characteristic difference between the best individual and other individuals in the population based on the calculated difference value;
[0053] The expression for calculating the difference value is as follows:
[0054] D(best,i)=|S best ∪S i ∣-∣S best ∩S i ∣
[0055] Among them, S best represents the number of feature labels of the best individual in the population, D represents the difference value, S iIndicates the number of feature labels of other individuals in the population;
[0056] Step B7: Select the difference value S best The largest seven individuals are voted together to obtain the image classification solution.
[0057] The developed integrated genetic programming DEGP module follows the guidance of the new program structure, function set and terminal set, selects features from the feature storage table as the terminal input of DEGP, and generates the initial population accordingly. The population undergoes fine evolutionary learning through genetic operators until the set termination conditions are met. Through this process, DEGP not only improves the effectiveness of individual features, but also enhances the wide applicability and robustness of the solution. Finally, the algorithm selects and integrates multiple high-quality individuals based on the difference values between individuals, and optimizes the final classification effect through this difference-based integration strategy, thereby outputting a high-performance image classification solution. This process ensures that the classification model can effectively handle various complex and changing image scenes in practical applications.
[0058] Preferably, the image classification solution includes seven individuals in a voting ensemble. In step 6, the seven individuals respectively perform image classification prediction on the image data to be classified, and respectively output predicted class labels. The predicted class label with the largest cumulative number is selected and output as the image classification result.
[0059] In the genetic programming DEGP module of the development integration, the population undergoes an evolutionary learning process to continuously optimize each individual until a preset iteration limit is reached, and finally returns a set of individuals, each of which maps a classification scheme and outputs a predicted class label. Ensembling diverse and effective individuals can improve classification accuracy and reduce the possibility of overfitting.
[0060] Therefore, the present invention proposes an integration strategy based on individual difference values, aiming to integrate the outputs of each individual to achieve a better classification effect; this method selects the seven individuals with the largest difference values for voting integration, aggregates the outputs of multiple individuals, and ensures that the final prediction does not only rely on a single scheme, thereby effectively improving the classification accuracy and generalization ability.
[0061] Beneficial effects of the present invention: The present invention adopts a hierarchical evolutionary learning framework, aiming to narrow the search space and efficiently find classification solutions.
[0062] In the first layer, a parallel exploration genetic programming PEGP is defined, which focuses on image preprocessing and feature extraction to build a feature storage table. By exploring diverse and effective feature building blocks in parallel, this layer can improve the combined generation efficiency of preprocessing and feature extraction in the GP algorithm, and finally output diverse feature blocks as input for the second stage.
[0063] The second layer is to develop an integrated genetic programming DEGP, which combines feature blocks in the search space and generates an integrated solution based on the feature storage table under the guidance of the newly defined program structure. The hierarchical learning-based framework significantly reduces the search space of the GP algorithm, thereby ensuring that the final solution can capture key feature information and maintain high performance even with fewer training samples.
[0064] In addition, given that the GP method has weak generalization ability when dealing with few-sample image classification tasks, an integration strategy based on individual difference values is proposed to improve the classification performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 This is a flowchart of the HLGP algorithm applied to image classification;
[0066] Figure 2 Schematic diagram of the process of generating feature building blocks for the PEGP algorithm;
[0067] Figure 3 Schematic diagram of the process of generating classification solutions for the DEGP algorithm;
[0068] Figure 4 Individual example graph generated by the DEGP algorithm;
[0069] Figure 5 The comparison chart of classification accuracy between HLGP algorithm and three sub-algorithms. DETAILED DESCRIPTION
[0070] The present invention is further described in detail below in conjunction with the accompanying drawings and specific examples. The following examples or drawings are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0071] like Figure 1 As shown: A small sample image classification method based on a hierarchical learning genetic programming algorithm includes the following steps:
[0072] Step 1: construct a small sample image classification system based on a hierarchical learning genetic programming algorithm, wherein the small sample image classification system is provided with an image acquisition module, wherein the image acquisition module is connected to a hierarchical evolutionary learning framework, wherein the hierarchical evolutionary learning framework is provided with a first-layer parallel exploration genetic programming PEGP module and a second-layer development integrated genetic programming DEGP module;
[0073] Step 2: The image acquisition module acquires an image data set and divides the image data set into a training set and a test set;
[0074] Step 3: The parallel exploration genetic programming PEGP module obtains training set data, and performs image preprocessing and feature extraction operations on the training set data to construct a feature storage table, and then passes the feature storage table to the development integrated genetic programming DEGP module;
[0075] The feature storage table stores local features and global features corresponding to the image data.
[0076] Step 4: The developed integrated genetic programming DEGP module uses the features in the feature storage table as terminal input, further constructs an integrated solution, and finally optimizes the final classification effect through an integrated strategy based on individual difference values, thereby outputting a high-performance image classification solution;
[0077] Step 5: Use the test set as input to the image classification solution, then output the predicted class labels of the test set, and finally evaluate the performance of the image classification solution based on the actual labels of the test set;
[0078] Step 6: The image acquisition module acquires the image data to be classified, uses the image data to be classified as the input of the image classification solution, and outputs the image classification result.
[0079] In step 1, the parallel exploration genetic programming PEGP module is provided with three different types of feature exploration blocks in parallel, and the three feature exploration blocks are respectively a first feature exploration block, a second feature exploration block and a third feature exploration block, and the program structure, function set and terminal set of the first feature exploration block, the second feature exploration block and the third feature exploration block are respectively set, and the first feature exploration block, the second feature exploration block and the third feature exploration block generate feature building blocks of three different evolutionary directions in parallel.
[0080] The first feature exploration block adopts a linear genetic programming LGP algorithm focusing on local features, and the program structure of the first feature exploration block includes a region extraction layer, a first image filtering layer, a first feature extraction layer and a first feature series connection layer;
[0081] The function set of the first feature exploration block is shown in Table 1:
[0082] Table 1
[0083]
[0084] The second feature exploration block adopts a grammar-guided genetic programming (GGP) algorithm focusing on global features, and the program structure of the second feature exploration block includes a second image filtering layer and a second feature extraction layer;
[0085] The function set of the second feature exploration block is shown in Table 2:
[0086] Table 2
[0087]
[0088] The third feature exploration block adopts a Cartesian genetic programming (CGP) algorithm that focuses on serial features. The program structure of the third feature exploration block includes a third image filtering layer, a maximum pooling layer, a third feature extraction layer, and a second feature serial layer.
[0089] The function set of the third feature exploration block is shown in Table 3:
[0090] Table 3
[0091]
[0092]
[0093] The terminal set of the parallel exploration genetic programming PEGP module includes a training set, as well as the parameters required by the functions in the linear genetic programming LGP algorithm, the grammar-guided genetic programming GGP algorithm, and the Cartesian genetic programming CGP algorithm. The specific required parameters are shown in Table 4:
[0094] Table 4
[0095]
[0096] like Figure 2 As shown: In step 3, the parallel exploration genetic programming PEGP module constructs a feature storage table, including the following steps:
[0097] Step A1: Population initialization: The feature exploration block obtains training set data and initializes the population according to a predetermined program structure, function set and terminal set; each individual in the population can be regarded as a feature extraction algorithm.
[0098] Step A2: Evaluate individual fitness: Each individual in the population extracts features from the training set, and then inputs the extracted features into a support vector machine (SVM). The support vector machine (SVM) outputs a predicted class label, and then the individual is evaluated using classification accuracy to obtain a fitness value of the corresponding individual.
[0099] The calculation expression of classification accuracy is as follows:
[0100]
[0101] Among them, N correct is the number of correctly predicted instances, N total is the total number of instances, and Fitness represents the fitness value of the individual;
[0102] In this process, in order to reduce the possibility of overfitting, the k-fold cross-validation method was used. In addition, in order to ensure the adequacy of cross-validation, the k value was set to the smaller of the number of training samples n for each category and 10.
[0103] Step A3: Elite operation: Use the elite strategy to select the best individuals in the population and directly copy them to the next generation population;
[0104] Step A4: Selection operation: A certain number of individuals are selected from the population using the tournament selection method, and each individual has the same probability of being selected; according to the fitness value of each individual, the individual with the best fitness value is selected to perform crossover and mutation operations to generate new individuals;
[0105] Step A5: Repeat steps A2-A4 until the maximum number of iterations is reached, then proceed to step A6;
[0106] Step A6: Select the top 50% individuals in the last generation of the population as feature building blocks;
[0107] Step A7: Use the feature building block to extract corresponding features from the training set, number them, and store them in a feature storage table as input for the second layer learning.
[0108] The development-integrated genetic programming DEGP module is set with a new program structure, function set and terminal set, and the program structure of the development-integrated genetic programming DEGP module includes a feature construction layer, a classification layer and a combination layer;
[0109] The feature construction layer is used to take at least two features in the feature storage table as input and return a concatenated feature, or construct a new feature according to parameters;
[0110] The classification layer is used to take the output features of the feature construction layer as input and output a predicted class label;
[0111] The combination layer is used to take at least two groups of predicted class labels output by the classification layer as input, and perform voting or weighting to output a new predicted class label.
[0112] The function set of the developed integrated genetic programming DEGP module is shown in Table 5:
[0113] Table 5
[0114]
[0115] like Figure 3 As shown: In step 4, the developed integrated genetic programming DEGP module further constructs an integrated solution, including the following steps:
[0116] (1) Population evolutionary learning
[0117] Step B1: Population initialization: The developed integrated genetic programming DEGP module obtains the features in the feature storage table and initializes the population according to the new program structure, function set and terminal set; each individual in the population is mapped to a classification scheme and output as a predicted class label.
[0118] Step B2: Evaluate individual fitness: Each individual in the population extracts features from the feature storage table, outputs a predicted class label, and then evaluates the fitness value of the individual;
[0119] Step B3: Elite operation: Use the elite strategy to select the best individuals in the population and directly copy them to the next generation population;
[0120] Step B4: Selection operation: A certain number of individuals are selected from the population using the tournament selection method. According to the fitness value of each individual, the individual with the best fitness value is selected to perform crossover and mutation operations to generate new individuals.
[0121] Step B5: Repeat steps B2-B4 until the preset upper limit of iteration is reached, then proceed to step B6;
[0122] (2) Integration strategy based on individual difference values
[0123] Step B6: Select the best performing individual in the last generation of the population as the benchmark, calculate the difference value of other individuals in the population, and evaluate the characteristic difference between the best individual and other individuals in the population based on the calculated difference value;
[0124] The difference value calculation method used is to count whether the feature types of the terminal nodes of these individuals are the same. These terminal nodes represent different features selected from the feature storage table. The individual expression can be represented as a series of strings, where the terminal contains different feature vectors in the feature storage table. Traverse the individual expression and count the number of occurrences of different feature vectors S i The expression for calculating the difference value is as follows:
[0125] D(best,i)=|S best ∪S i ∣-∣S best ∩S i ∣
[0126] Among them, S best represents the number of feature labels of the best individual in the population, D represents the difference value, S i Indicates the number of feature labels of other individuals in the population;
[0127] The feature difference between two individuals is determined by calculating the size of the symmetric difference between the two sets, that is, the number of elements that exist only in one set. Such a calculation directly reflects the difference in feature selection between the two individuals. If the two individuals have exactly the same features, the difference value will be zero; if they have no common features, the difference value will be equal to the sum of the number of unique elements in the two sets.
[0128] Step B7: Select the difference value S best The largest seven individuals are voted together to obtain the image classification solution.
[0129] The image classification solution includes seven individuals in a voting ensemble. In step 6, the seven individuals respectively perform image classification prediction on the image data to be classified, and respectively output predicted class labels. The predicted class label with the largest cumulative number is selected and output as the image classification result.
[0130] Figure 4 shows examples of individuals generated by the DEGP algorithm, from Figure 4 As can be seen in Figure 1, the DEGP algorithm randomly selects features from the feature storage table as its input for the second-level evolutionary learning. For the DEGP algorithm, the output of the generated individual is the predicted class label generated by the combination layer function, and the fitness function is also the classification accuracy calculated using k-fold cross validation on the training set.
[0131] Next, the classification performance of the present invention is further verified through specific experiments.
[0132] 1. Dataset
[0133] This embodiment evaluates the performance of a small sample image classification method based on a hierarchical learning genetic programming algorithm on four different image datasets. Among them, CIFAR10 is a widely used object classification dataset, containing 50,000 32×32 training images and 10,000 test images in 10 categories. Fashion_MNIST, referred to as FMNIST, is an image classification task divided into 10 fashion categories. The dataset includes 60,000 28×28 grayscale training images and 10,000 test images. SVHN is a digital classification dataset, including 10 categories, consisting of 73,257 32×32 color training images and 26,032 test images. ORL is a face recognition dataset, which contains 40 different individuals, each with 10 different images, and the image size is 92x112 pixels. In this embodiment, the image size is set to 46×56.
[0134] For the few-sample image classification problem, 10, 20, 40 and 80 images of each category are randomly selected from CIFAR10, FMNIST and SVHN datasets as training data. 2, 3, 4 and 5 training images are used for each category in ORL dataset respectively. This design aims to explore the performance of the present invention when the training samples are extremely limited.
[0135] 2. Benchmark Methodology
[0136] In order to verify the effectiveness of the hierarchical learning genetic programming HLGP algorithm on the few-sample image classification problem, this embodiment compares it with a variety of benchmark methods. The compared methods include the current advanced genetic programming GP algorithm and the most advanced deep learning method based on the benchmark dataset.
[0137] (1) Image classification methods based on genetic programming (GP): Three GP-based methods are used to compare all datasets to show the effectiveness of the proposed method. These methods are FGP, a genetic programming image classification method based on image-related operations and flexible program structure, FLGP, a genetic programming feature learning method based on image description, and BERGP, a genetic programming image classification method based on building block evolution and reuse. These three methods use different individual representations to automatically learn different types of features and / or evolve effective image classification ensembles, and have achieved satisfactory results on different image datasets. At the same time, in order to prove the effectiveness of hierarchical learning, this method is compared with three genetic programming GP algorithms for first-layer learning.
[0138] (2) Deep learning methods: On CIFAR10, FMNIST, and SVHN, the state-of-the-art methods are based on convolutional neural networks (CNNs) and the deep residual network ResNet20 in the study. The adopted CNN models are divided into three categories according to the complexity of the network architecture: low complexity CNN (CNN-lc), medium complexity CNN (CNN-mc), and high complexity CNN (CNN-hc). The difference in these complexities lies mainly in the number of filters in the convolutional layer and the depth of the network layer. In addition, the performance of these three types of models under different dropout rates (0, 0.4, 0.7) is also studied to explore their effectiveness in preventing overfitting in the case of few samples.
[0139] 3. Parameter settings
[0140] To ensure the comparability of the experimental results, the settings of all algorithms in the comparison method based on genetic programming GP are as follows: the maximum number of generations is set to 50, the population size is 100, the elite rate is 0.01, the mutation rate is 0.19, and the crossover rate is 0.8. The selection method adopted is tournament selection with a size of 5. The minimum and maximum depths of the trees of FGP, FLGP and BERGP are set to 2 and 8 respectively. For the HLGP algorithm, the population size of its first-layer evolutionary learning, including LGP, GGP and CGP algorithms, is set to 250 and the maximum number of generations is 10, aiming to explore a diverse solution set; the population size of the DEGP algorithm of the second-layer evolutionary learning is 100 and the maximum number of generations is 25 to speed up the convergence speed and keep the same number of evaluations as other GP algorithms. Each method uses different random seeds and runs independently 30 times on each training set to evaluate its stability. The experimental results will report the average accuracy and its standard deviation on the test set.
[0141] 4. Classification performance analysis
[0142] The HLGP algorithm is compared with CNNs and ResNet-20 with different complexity and dropout rates, as well as other advanced GP-based image classification methods.
[0143] (1) Comparison with CNNs, ResNet-20 and BERGP algorithms: Considering that the BERGP method only provides average accuracy data and lacks detailed iteration data, this variant cannot be included in the subsequent rank sum test. Therefore, it is reported together with the CNNs method. Tables 6 to 8 show the performance comparison results of the baseline method and the HLGP algorithm.
[0144] Table 6 Performance comparison on CIFAR-10 dataset
[0145]
[0146] On the CIFAR-10 dataset in Table 6, 44 comparisons were made with the baseline methods. HLGP outperformed the baseline methods in 41 comparisons and was inferior to the comparison algorithms in three comparisons of CIFAR10-10. As a diverse small image dataset, the resolution and number of samples of CIFAR10 pose great challenges to the performance of the classification model, especially when the number of samples is small. HLGP effectively addresses this challenge through a hierarchical learning strategy. This strategy enables the algorithm to find more effective discriminative features in less data by exploring and combining features at different levels, thereby improving the classification accuracy. For example, in the smallest sample case of CIFAR10-10, HLGP achieved an accuracy of 30.7%, slightly higher than BERGP's 30.6%, and significantly better than ResNet-20's 23.3%. In the larger sample configuration CIFAR10-80, HLGP's performance is further improved to 51.1%, higher than BERGP's 49.7%. This performance improvement demonstrates the ability of HLGP to effectively extract key features through a hierarchical learning strategy when dealing with a small number of samples. Especially under low-sample conditions, compared with advanced CNNs models and GP algorithm-based image classification methods, HLGP can better adapt to the limitations of sample number and achieve higher classification accuracy and better generalization performance.
[0147] Table 7 Performance comparison on FMNIST dataset
[0148]
[0149] On the FMNIST dataset in Table 7, the HLGP algorithm was compared with various benchmark methods 44 times and outperformed the benchmark methods in all 44 comparisons. Fashion-MNIST, as a grayscale image dataset involving the classification of fashion items, poses considerable challenges to classification algorithms, especially when the number of samples is limited. HLGP effectively addresses these challenges through its unique hierarchical learning strategy, which explores and integrates features at different levels, allowing the algorithm to find highly discriminative features even when the amount of data is small, thereby improving classification accuracy. For example, in the smallest sample configuration FMNIST-10, HLGP achieved an accuracy of 75.8%, which not only exceeds the performance of most CNN models, but also significantly outperforms BERGP's 72.1%. In the larger sample configuration FMNIST-80, HLGP's performance further improved to 85.8%, continuing to lead BERGP (83.4%). This series of results not only demonstrates the advantages of HLGP in dealing with few-sample problems, but also highlights its performance improvement brought by the hierarchical learning framework on diverse image datasets.
[0150] Table 8 Performance comparison on SVHN dataset
[0151]
[0152] On the SVHN dataset in Table 8, the HLGP algorithm is compared with a variety of baseline methods and shows significant performance advantages. The SVHN dataset is well-known for its wide range of street-level digital images. The diversity of digits in the images and background noise pose a high challenge to the classification algorithm. The HLGP algorithm effectively improves the processing ability of such complex images through feature exploration and combinatorial optimization. In the smallest sample SVHN-10, HLGP's performance is slightly lower than BERGP, achieving an accuracy of 59.2%, while BERGP is 60.2%. However, as the number of samples increases, HLGP's performance begins to exceed BERGP and other CNNs-based methods. In the settings of SVHN-20, SVHN-40, and SVHN-80, HLGP achieves accuracies of 69.9%, 77.3%, and 78.2%, respectively, showing that its classification performance increases synchronously with the increase in the number of samples. When compared with the CNNs model, HLGP shows its superiority. For example, in SVHN-80, the highest accuracy of the CNNs model is 74.6%, while HLGP reaches 78.2%. This performance improvement is attributed to the fact that HLGP adopts an ensemble strategy based on individual difference values, which significantly improves the excellent generalization ability on diverse data.
[0153] (2) Comparison with GP-based methods: In the comparison with GP-based methods, the Wilcoxon rank sum test with a 5% significance level was used to show the significance of performance improvement. In this section, we focus on those GP method variants that provide complete data (including the results of each cycle). The rank sum test will help statistically verify the performance differences between the variants and provide a scientific basis for the final method selection. The comparison with the GP-based method is shown in Table 9.
[0154] Table 9 Comparison results with GP-based methods on various data sets
[0155]
[0156]
[0157] The average classification performance indicators of HLGP and GP-based image classification methods were statistically compared at the 5% significance level using the Wilcoxon rank sum test, and the results are summarized in Table 10. As can be seen from the table, out of 80 performance comparisons, HLGP significantly outperformed the comparison method in 78 comparisons and performed equally well with the comparison method in 2 comparisons. This result clearly shows that when dealing with image classification problems, HLGP outperforms traditional methods on multiple datasets, demonstrating the effectiveness and stability of its method. This advantage is not only reflected in improving classification accuracy, but also in having good generalization capabilities for datasets of different types and sizes.
[0158] Table 10 Statistical test on average classification accuracy compared with GP-based methods
[0159]
[0160] Compared with the FGP algorithm, HLGP showed better performance in 14 comparisons and only performed worse than FGP in 2 comparisons. HLGP outperformed FGP in particular when processing complex CIFAR10, FMNIST, and SVHN datasets. This is mainly due to HLGP's two-layer learning framework, which not only increases the flexibility of the algorithm but also improves the overall solution quality. In the first-layer learning stage of HLGP, the feature information in the data is effectively captured by constructing a variety of feature building blocks. Then, in the second-layer learning, these features are used to build an integrated solution, further optimizing the classification results. The solutions generated based on this are significantly better than those generated by the single-layer learning method, which is also the main reason why HLGP surpasses FGP on these datasets.
[0161] Compared with the FLGP algorithm, the HLGP algorithm showed better performance in 16 comparisons. The program structure of FLGP focuses on capturing local and global features of the image at the same time, while HLGP effectively narrows the search space and optimizes the feature exploration process by using the LGP and GGP algorithms in parallel in the first layer of learning. This strategy allows HLGP to more accurately identify useful features when dealing with complex image classification tasks, thus significantly outperforming FLGP on various data sets, demonstrating its efficient feature learning and solution building capabilities.
[0162] HLGP significantly outperforms LGP, GGP, and CGP in 48 comparisons. In the first layer of learning, LGP, GGP, and CGP independently extract a class of features for image classification. This single feature extraction method performs poorly in most cases, especially in image classification tasks that require complex feature fusion. Specifically, on most datasets, LGP, GGP, and CGP perform worse than FGP or FLGP. Through the second layer of evolutionary learning of HLGP, these primary extracted features are further combined and optimized, significantly improving the final classification effect. This result clearly shows that the hierarchical learning structure adopted by HLGP is effective in improving algorithm performance.
[0163] In order to further illustrate the effectiveness of the hierarchical learning designed by HLGP, it is necessary to compare the three sub-GP algorithms designed in the first layer of learning with the overall HLGP algorithm. This comparison can clearly show the improvement brought by hierarchical learning, as shown in the following figure. Figure 5 shown.
[0164] from Figure 5 From the statistics in , we can see that the three sub-GP algorithms designed in the first stage have achieved good results and each showed different classification results. Nevertheless, HLGP with a hierarchical learning architecture significantly outperformed the three sub-algorithms in overall performance. The only exception occurred on the ORL dataset, where the performance improvement of HLGP was not as significant as in other datasets. This may be because the classification task of the ORL dataset is relatively simple, while HLGP is designed mainly to cope with image classification tasks with high complexity and a limited number of samples. In summary, the hierarchical learning framework can significantly improve performance when dealing with complex and challenging classification tasks, such as few-sample image classification problems.
[0165] This paper focuses on solving the problem of few-sample image classification and proposes a hierarchical learning genetic programming algorithm HLGP. This method is divided into two levels: the first level focuses on image preprocessing and feature extraction, with the goal of building an efficient feature storage table; the second level selects high-quality features that have been optimized through learning, and develops integrated solutions under the guidance of the program structure. The hierarchical learning framework successfully narrows the search space of the GP algorithm and generates excellent integrated solutions. The proposed integration strategy based on individual difference values performs secondary integration of multiple integrated solutions by evaluating the differences between individuals. This method not only utilizes the uniqueness of each individual in the algorithm, but also enhances the final classification performance by combining diverse solutions. Finally, after comparing with multiple advanced comparative algorithms, the experimental results show that the HLGP algorithm outperforms all comparative algorithms when the number of samples is limited. These results show that HLGP is an efficient method for dealing with the problem of few-sample image classification, and it has the ability to achieve accurate classification in complex classification tasks.
[0166] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A small sample image classification method based on hierarchical learning genetic programming algorithm, characterized in that: The following steps are involved: Step 1: construct a small sample image classification system based on a hierarchical learning genetic programming algorithm, wherein the small sample image classification system is provided with an image acquisition module, wherein the image acquisition module is connected to a hierarchical evolutionary learning framework, wherein the hierarchical evolutionary learning framework is provided with a first-layer parallel exploration genetic programming PEGP module and a second-layer development integrated genetic programming DEGP module; Step 2: The image acquisition module acquires an image data set and divides the image data set into a training set and a test set; Step 3: The parallel exploration genetic programming PEGP module obtains training set data, and performs image preprocessing and feature extraction operations on the training set data to construct a feature storage table, and then passes the feature storage table to the development integrated genetic programming DEGP module; Step 4: The developed integrated genetic programming DEGP module uses the features in the feature storage table as terminal input, further constructs an integrated solution, and finally optimizes the final classification effect through an integrated strategy based on individual difference values, thereby outputting a high-performance image classification solution; Step 5: Use the test set as input to the image classification solution, then output the predicted class labels of the test set, and finally evaluate the performance of the image classification solution based on the actual labels of the test set; Step 6: The image acquisition module acquires the image data to be classified, uses the image data to be classified as the input of the image classification solution, and outputs the image classification result.
2. The small sample image classification method based on hierarchical learning genetic programming algorithm according to claim 1 is characterized in that: In step 1, the parallel exploration genetic programming PEGP module is provided with three different types of feature exploration blocks in parallel, and the three feature exploration blocks are respectively a first feature exploration block, a second feature exploration block and a third feature exploration block, and the program structure, function set and terminal set of the first feature exploration block, the second feature exploration block and the third feature exploration block are respectively set, and the first feature exploration block, the second feature exploration block and the third feature exploration block generate feature building blocks of three different evolutionary directions in parallel.
3. The small sample image classification method based on hierarchical learning genetic programming algorithm according to claim 2 is characterized by: The first feature exploration block adopts a linear genetic programming LGP algorithm focusing on local features, and the program structure of the first feature exploration block includes a region extraction layer, a first image filtering layer, a first feature extraction layer and a first feature series connection layer; The second feature exploration block adopts a grammar-guided genetic programming (GGP) algorithm focusing on global features, and the program structure of the second feature exploration block includes a second image filtering layer and a second feature extraction layer; The third feature exploration block adopts a Cartesian genetic programming (CGP) algorithm that focuses on serial features. The program structure of the third feature exploration block includes a third image filtering layer, a maximum pooling layer, a third feature extraction layer, and a second feature serial layer.
4. The small sample image classification method based on hierarchical learning genetic programming algorithm according to claim 3 is characterized by: The function set of the first feature exploration block is shown in Table 1: Table 1 The function set of the second feature exploration block is shown in Table 2: Table 2 The function set of the third feature exploration block is shown in Table 3: Table 3 5. The small sample image classification method based on hierarchical learning genetic programming algorithm according to claim 4 is characterized in that: The terminal set of the parallel exploration genetic programming PEGP module includes a training set, as well as the parameters required by the functions in the linear genetic programming LGP algorithm, the grammar-guided genetic programming GGP algorithm, and the Cartesian genetic programming CGP algorithm. The specific required parameters are shown in Table 4: Table 4 6. The small sample image classification method based on hierarchical learning genetic programming algorithm according to claim 2 is characterized by: In step 3, the parallel exploratory genetic programming PEGP module constructs a feature storage table, including the following steps: Step A1: Population initialization: The feature exploration block obtains training set data and initializes the population according to a predetermined program structure, function set and terminal set; Step A2: Evaluate individual fitness: Each individual in the population extracts features from the training set, and then inputs the extracted features into a support vector machine (SVM). The support vector machine (SVM) outputs a predicted class label, and then the individual is evaluated using classification accuracy to obtain a fitness value of the corresponding individual. The calculation expression of classification accuracy is as follows: Among them, N correct is the number of correctly predicted instances, N total is the total number of instances, and Fitness represents the fitness value of the individual; Step A3: Elite operation: Use the elite strategy to select the best individuals in the population and directly copy them to the next generation population; Step A4: Selection operation: A certain number of individuals are selected from the population using the tournament selection method, and each individual has the same probability of being selected; according to the fitness value of each individual, the individual with the best fitness value is selected to perform crossover and mutation operations to generate new individuals; Step A5: Repeat steps A2-A4 until the maximum number of iterations is reached, then proceed to step A6; Step A6: Select the top 50% individuals in the last generation of the population as feature building blocks; Step A7: Use the feature building block to extract corresponding features from the training set, number them, and store them in a feature storage table as input for the second layer learning.
7. The small sample image classification method based on hierarchical learning genetic programming algorithm according to claim 1 is characterized by: The development-integrated genetic programming DEGP module is set with a new program structure, function set and terminal set, and the program structure of the development-integrated genetic programming DEGP module includes a feature construction layer, a classification layer and a combination layer; The feature construction layer is used to take at least two features in the feature storage table as input and return a concatenated feature, or construct a new feature according to parameters; The classification layer is used to take the output features of the feature construction layer as input and output a predicted class label; The combination layer is used to take at least two groups of predicted class labels output by the classification layer as input, and perform voting or weighting to output a new predicted class label.
8. The small sample image classification method based on hierarchical learning genetic programming algorithm according to claim 7 is characterized by: The function set of the developed integrated genetic programming DEGP module is shown in Table 5: Table 5 9. The small sample image classification method based on hierarchical learning genetic programming algorithm according to claim 7 is characterized by: In step 4, the development integrated genetic programming DEGP module further constructs an integrated solution, including the following steps: (1) Population evolutionary learning Step B1: Population initialization: the developed integrated genetic programming DEGP module obtains the features in the feature storage table and initializes the population according to the new program structure, function set and terminal set; Step B2: Evaluate individual fitness: Each individual in the population extracts features from the feature storage table, outputs a predicted class label, and then evaluates the fitness value of the individual; Step B3: Elite operation: Use the elite strategy to select the best individuals in the population and directly copy them to the next generation population; Step B4: Selection operation: A certain number of individuals are selected from the population using the tournament selection method. According to the fitness value of each individual, the individual with the best fitness value is selected to perform crossover and mutation operations to generate new individuals. Step B5: Repeat steps B2-B4 until the preset upper limit of iteration is reached, then proceed to step B6; (2) Integration strategy based on individual difference values Step B6: Select the best performing individual in the last generation of the population as the benchmark, calculate the difference value of other individuals in the population, and evaluate the characteristic difference between the best individual and other individuals in the population based on the calculated difference value; The expression for calculating the difference value is as follows: D(best,i)=∣S best ∪S i ∣-∣S best ∩S i ∣ Among them, S best represents the number of feature labels of the best individual in the population, D represents the difference value, S i Indicates the number of feature labels of other individuals in the population; Step B7: Select the difference value S best The largest seven individuals are voted together to obtain the image classification solution.
10. The small sample image classification method based on hierarchical learning genetic programming algorithm according to claim 9 is characterized in that: The image classification solution includes seven individuals in a voting ensemble. In step 6, the seven individuals respectively perform image classification prediction on the image data to be classified, and respectively output predicted class labels. The predicted class label with the largest cumulative number is selected and output as the image classification result.
Citation Information
Patent Citations
Information processing device, information processing program, and information processing method
CN110914864A
Image classification method based on genetic algorithm and data set division
CN114662593A
Image classification method based on new genetic programming structure and gene modification
CN117274692A
Small sample image classification integration method based on genetic programming
CN118552768A
Information processing device, information processing method, and computer-readable recording medium recording information processing program
US20210110215A1